Martha Palmer

dblp:p/MarthaStonePalmer · also Martha Stone, Martha Stone Palmer · DBLP profile ↗
← Back
112ranked-venue papers
14as first author
15since 2021 · last 2026
0000-0001-9864-6974ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 102 · 14 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 CLEVR-3D-DeRef
Mary Lynn Martin, Martha Palmer, Maria Leonor Pacheco
LREC2
2026 Building Bridges between Student and Curricular Language: Creating a Corpus of Abstract Meaning Representations for the Classroom
Kristin Wright-Bettner, Zheng Cai, Zekun Zhao, James H. Martin, Jeffrey Flanigan, Martha Palmer
LREC6
2025 Speech Is Not Enough: Interpreting Nonverbal Indicators of Common Knowledge and Engagement
abstract
Our goal is to develop an AI Partner that can provide support for group problem solving and social dynamics. In multi-party working group environments, multimodal analytics is crucial for identifying non-verbal interactions of group members. In conjunction with their verbal participation, this creates an holistic understanding of collaboration and engagement that provides necessary context for the AI Partner. In this demo, we illustrate our present capabilities at detecting and tracking nonverbal behavior in student task-oriented interactions in the classroom, and the implications for tracking common ground and engagement.
Derek Palmer, Yifan Zhu 0014, Kenneth Lai, Hannah VanderHoeven, Mariah Bradford, Ibrahim Khebour, Carlos Mabrey, Jack Fitzgerald, Nikhil Krishnaswamy, Martha Palmer, James Pustejovsky
AAAI10
2024 Linear Cross-document Event Coreference Resolution with X-AMR
abstract
Event Coreference Resolution (ECR) as a pairwise mention classification task is expensive both for automated systems and manual annotations. The task’s quadratic difficulty is exacerbated when using Large Language Models (LLMs), making prompt engineering for ECR prohibitively costly. In this work, we propose a graphical representation of events, X-AMR, anchored around individual mentions using a cross-document version of Abstract Meaning Representation. We then linearize the ECR with a novel multi-hop coreference algorithm over the event graphs. The event graphs simplify ECR, making it a) LLM cost-effective, b) compositional and interpretable, and c) easily annotated. For a fair assessment, we first enrich an existing ECR benchmark dataset with these event graphs using an annotator-friendly tool we introduce. Then, we employ GPT-4, the newest LLM by OpenAI, for these annotations. Finally, using the ECR algorithm, we assess GPT-4 against humans and analyze its limitations. Through this research, we aim to advance the state-of-the-art for efficient ECR and shed light on the potential shortcomings of current LLMs at this task. Code and annotations: https://github.com/ahmeshaf/gpt_coref
Shafiuddin Rehan Ahmed, George Arthur Baker, Evi Judge, Michael Regan, Kristin Wright-Bettner, Martha Palmer, James H. Martin
LREC/COLING6
2024 ReCAP: Semantic Role Enhanced Caption Generation
abstract
Even though current vision language (V+L) models have achieved success in generating image captions, they often lack specificity and overlook various aspects of the image. Additionally, the attention learned through weak supervision operates opaquely and is difficult to control. To address these limitations, we propose the use of semantic roles as control signals in caption generation. Our hypothesis is that, by incorporating semantic roles as signals, the generated captions can be guided to follow specific predicate argument structures. To validate the effectiveness of our approach, we conducted experiments using data and compared the results with a baseline model VL-BART(CITATION). The experiments showed a significant improvement, with a gain of 45% in Smatch score (Standard NLP evaluation metric for semantic representations), demonstrating the efficacy of our approach. By focusing on specific objects and their associated semantic roles instead of providing a general description, our framework produces captions that exhibit enhanced quality, diversity, and controllability.
Abhidip Bhattacharyya, Martha Palmer, Christoffer R. Heckman
LREC/COLING2
2024 Building a Broad Infrastructure for Uniform Meaning Representations
abstract
This paper reports the first release of the UMR (Uniform Meaning Representation) data set. UMR is a graph-based meaning representation formalism consisting of a sentence-level graph and a document-level graph. The sentence-level graph represents predicate-argument structures, named entities, word senses, aspectuality of events, as well as person and number information for entities. The document-level graph represents coreferential, temporal, and modal relations that go beyond sentence boundaries. UMR is designed to capture the commonalities and variations across languages and this is done through the use of a common set of abstract concepts, relations, and attributes as well as concrete concepts derived from words from invidual languages. This UMR release includes annotations for six languages (Arapaho, Chinese, English, Kukama, Navajo, Sanapana) that vary greatly in terms of their linguistic properties and resource availability. We also describe on-going efforts to enlarge this data set and extend it to other genres and modalities. We also briefly describe the available infrastructure (UMR annotation guidelines and tools) that others can use to create similar data sets.
Julia Bonn, Matthew J. Buchholz, Jayeol Chun, Andrew Cowell, William Croft 0001, Lukas Denk, Sijia Ge, Jan Hajic 0001, Kenneth Lai, James H. Martin, Skatje Myers, Alexis Palmer, Martha Palmer, Claire Benet Post, James Pustejovsky, Kristine Stenzel, Haibo Sun, Zdenka Uresová, Rosa Vallejos, Jens E. L. Van Gysel, Meagan Vigus, Nianwen Xue, Jin Zhao 0009
LREC/COLING13
2024 GLAMR: Augmenting AMR with GL-VerbNet Event Structure
abstract
This paper introduces GLAMR, an Abstract Meaning Representation (AMR) interpretation of Generative Lexicon (GL) semantic components. It includes a structured subeventual interpretation of linguistic predicates, and encoding of the opposition structure of property changes of event arguments. Both of these features are recently encoded in VerbNet (VN), and form the scaffolding for the semantic form associated with VN frame files. We develop a new syntax, concepts, and roles for subevent structure based on VN for connecting subevents to atomic predicates. Our proposed extension is compatible with current AMR specification. We also present an approach to automatically augment AMR graphs by inserting subevent structure of the predicates and identifying the subevent arguments from the semantic roles. A pilot annotation of GLAMR graphs of 65 documents (486 sentences), based on procedural texts as a source, is presented as a public dataset. The annotation includes subevents, argument property change, and document-level anaphoric links. Finally, we provide baseline models for converting text to GLAMR and vice versa, along with the application of GLAMR for generating enriched paraphrases with details on subevent transformation and arguments that are not present in the surface form of the texts.
Jingxuan Tu, Timothy Obiso, Bingyang Ye, Kyeongmin Rim, Keer Xu, Liulu Yue, Susan Windisch Brown, Martha Palmer, James Pustejovsky
LREC/COLING8
2024 Prompting as Panacea? A Case Study of In-Context Learning Performance for Qualitative Coding of Classroom Dialog
Ananya Ganesh, Chelsea Chandler, Sidney K. D'Mello, Martha Palmer, Katharina Kann
EDM4
2024 My Big, Fat 50-Year Journey
abstract
Abstract My most heartfelt thanks to ACL for this tremendous honor. I’m completely thrilled. I cannot tell you how surprised I was when I got Iryna’s email. It is amazing that my first ACL conference since 2019 in Florence includes this award. What a wonderful way to be back with all of my friends and family here at ACL. I’m going to tell you about my big fat 50-year journey. What have I been doing for the last 50 years? Well, finding meaning, quite literally in words. Or in other words, exploring how computational lexical semantics can support natural language understanding. This is going to be quick. Hold onto your hats, here we go.
Martha Palmer
Comput. Linguistics1
2023 Navigating Wanderland: Highlighting Off-Task Discussions in Classrooms
Ananya Ganesh, Michael Alan Chang, Rachel Dickler, Michael Regan, Jon Z. Cai, Kristin Wright-Bettner, James Pustejovsky, James H. Martin, Jeffrey Flanigan, Martha Palmer, Katharina Kann
AIED10
2023 GLEN: General-Purpose Event Detection for Thousands of Types
abstract
The progress of event extraction research has been hindered by the absence of wide-coverage, large-scale datasets.To make event extraction systems more accessible, we build a generalpurpose event detection dataset GLEN which covers 205K event mentions with 3,465 different types, making it more than 20x larger in ontology than today's largest event dataset.GLEN is created by utilizing the DWD Overlay, which provides a mapping between Wikidata Qnodes and PropBank rolesets.This enables us to use the abundant existing annotation for PropBank as distant supervision.In addition, we also propose a new multi-stage event detection model CEDAR specifically designed to handle the large ontology size in GLEN.We show that our model exhibits superior performance compared to a range of baselines including InstructGPT.Finally, we perform error analysis and show that label noise is still the largest challenge for improving performance for this new dataset. 1
Qiusi Zhan, Kathryn Conger, Martha Palmer, Heng Ji 0001, Jiawei Han 0001
EMNLP4
2023 A Comparative Analysis of Automatic Speech Recognition Errors in Small Group Classroom Discourse
abstract
In collaborative learning environments, effective intelligent learning systems need to accurately analyze and understand the collaborative discourse between learners (i.e., group modeling) to provide adaptive support. We investigate how automatic speech recognition (ASR) errors influence discourse models of small group collaboration in noisy real-world classrooms. Our dataset consisted of 30 students recorded by consumer off-the-shelf microphones (Yeti Blue) while engaging in dyadic- and triadic- collaborative learning in a multi-day STEM curriculum unit. We found that two state-of-the-art ASR systems (Google Speech and OpenAI Whisper) yielded very high word error rates (0.822, 0.847) but very different profiles of error with Google being more conservative, rejecting 38% of utterances instead of 12% for Whisper. Next, we examined how these ASR errors influenced down-stream small group modeling based on pre-trained large language models for three tasks: Abstract Meaning Representation parsing (AMRParsing), on-task/off-task detection (OnTask), and Accountable Productive Talk prediction (TalkMove). As expected, models trained on clean human transcripts yielded degraded performance on all three tasks, measured by the transfer ratio (TR). However, the TR of the specific sentence-level AMRParsing task (.39 - .62) was much lower than that of the abstract discourse-level OnTask (.63- .94) and TalkMove tasks (.64-.72). Furthermore, different training strategies that incorporated ASR transcripts alone or as augmentations of human transcripts increased accuracy for the discourse-level tasks (OnTask and TalkMove) but not AMRParsing. Simulation experiments suggested that the models were tolerant of missing utterances in the dialog context, and that jointly improving ASR accuracy on important word classes (e.g., verbs and nouns) can improve performance across all tasks. Overall, our results provide insights into how different types of NLP-based tasks might be tolerant of ASR errors under extremely noisy conditions and provide suggestions for how to improve accuracy in small group modeling settings for a more equitable, engaging, and adaptive collaborative learning environment.
Jie Cao 0010, Ananya Ganesh, Jon Z. Cai, Rosy Southwell, Margaret Perkoff, Michael Regan, Katharina Kann, James H. Martin, Martha Palmer, Sidney K. D'Mello
UMAP9
2022 NewsClaims: A New Benchmark for Claim Detection from News with Attribute Knowledge
abstract
Revanth Gangi Reddy, Sai Chetan Chinthakindi, Zhenhailong Wang, Yi Fung, Kathryn Conger, Ahmed ELsayed, Martha Palmer, Preslav Nakov, Eduard Hovy, Kevin Small, Heng Ji. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Revanth Gangi Reddy, Sai Chetan Chinthakindi, Zhenhailong Wang, Yi R. Fung 0001, Kathryn Conger, Ahmed Elsayed, Martha Palmer, Preslav Nakov, Eduard H. Hovy, Kevin Small, Heng Ji 0001
EMNLP7
2022 Aligning Images and Text with Semantic Role Labels for Fine-Grained Cross-Modal Understanding
abstract
As vision processing and natural language processing continue to advance, there is increasing interest in multimodal applications, such as image retrieval, caption generation, and human-robot interaction. These tasks require close alignment between the information in the images and text. In this paper, we present a new multimodal dataset that combines state of the art semantic annotation for language with the bounding boxes of corresponding images. This richer multimodal labeling supports cross-modal inference for applications in which such alignment is useful. Our semantic representations, developed in the natural language processing community, abstract away from the surface structure of the sentence, focusing on specific actions and the roles of their participants, a level that is equally relevant to images. We then utilize these representations in the form of semantic role labels in the captions and the images and demonstrate improvements in standard tasks such as image retrieval. The potential contributions of these additional labels is evaluated using a role-aware retrieval system based on graph convolutional and recurrent neural networks. The addition of semantic roles into this system provides a significant increase in capability and greater flexibility for these tasks, and could be extended to state-of-the-art techniques relying on transformers with larger amounts of annotated data.
Abhidip Bhattacharyya, Cecilia Mauceri, Martha Palmer, Christoffer R. Heckman
LREC3
2021 Fine-grained Information Extraction from Biomedical Literature based on Knowledge-enriched Abstract Meaning Representation
abstract
Zixuan Zhang, Nikolaus Parulian, Heng Ji, Ahmed Elsayed, Skatje Myers, Martha Palmer. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Nikolaus Nova Parulian, Heng Ji 0001, Ahmed Elsayed, Skatje Myers, Martha Palmer
ACL/IJCNLP (1)6
2020 Verb Class Induction with Partial Supervision
Daniel W. Peterson, Susan Windisch Brown, Martha Palmer
AAAI3
2020 Structured Tuning for Semantic Role Labeling
abstract
Recent neural network-driven semantic role labeling (SRL) systems have shown impressive improvements in F1 scores.These improvements are due to expressive input representations, which, at least at the surface, are orthogonal to knowledge-rich constrained decoding mechanisms that helped linear SRL models.Introducing the benefits of structure to inform neural models presents a methodological challenge.In this paper, we present a structured tuning framework to improve models using softened constraints only at training time.Our framework leverages the expressiveness of neural networks and provides supervision with structured loss components.We start with a strong baseline (RoBERTa) to validate the impact of our approach, and show that our framework outperforms the baseline by learning to comply with declarative constraints.Additionally, our experiments with smaller training sizes show that we can achieve consistent improvements under low-resource scenarios.
Tao Li 0039, Parth Anand Jawale, Martha Palmer, Vivek Srikumar
ACL3
2020 Spatial AMR: Expanded Spatial Annotation in the Context of a Grounded Minecraft Corpus
abstract
This paper presents an expansion to the Abstract Meaning Representation (AMR) annotation schema that captures fine-grained semantically and pragmatically derived spatial information in grounded corpora. We describe a new lexical category conceptualization and set of spatial annotation tools built in the context of a multimodal corpus consisting of 170 3D structure-building dialogues between a human architect and human builder in Minecraft. Minecraft provides a particularly beneficial spatial relation-elicitation environment because it automatically tracks locations and orientations of objects and avatars in the space according to an absolute Cartesian coordinate system. Through a two-step process of sentence-level and document-level annotation designed to capture implicit information, we leverage these coordinates and bearings in the AMRs in combination with spatial framework annotation to ground the spatial language in the dialogues to absolute space.
Julia Bonn, Martha Palmer, Zheng Cai, Kristin Wright-Bettner
LREC2
2020 From Spatial Relations to Spatial Configurations
abstract
Spatial Reasoning from language is essential for natural language understanding. Supporting it requires a representation scheme that can capture spatial phenomena encountered in language as well as in images and videos. Existing spatial representations are not sufficient for describing spatial configurations used in complex tasks. This paper extends the capabilities of existing spatial representation languages and increases coverage of the semantic aspects that are needed to ground spatial meaning of natural language text in the world. Our spatial relation language is able to represent a large, comprehensive set of spatial concepts crucial for reasoning and is designed to support composition of static and dynamic spatial configurations. We integrate this language with the Abstract Meaning Representation (AMR) annotation schema and present a corpus annotated by this extended AMR. To exhibit the applicability of our representation scheme, we annotate text taken from diverse datasets and show how we extend the capabilities of existing spatial representation languages with fine-grained decomposition of semantics and blend it seamlessly with AMRs of sentences and discourse representations as a whole.
Soham Dan, Parisa Kordjamshidi, Julia Bonn, Archna Bhatia, Zheng Cai, Martha Palmer, Dan Roth 0001
LREC6
2020 The Russian PropBank
abstract
This paper presents a proposition bank for Russian (RuPB), a resource for semantic role labeling (SRL). The motivating goal for this resource is to automatically project semantic role labels from English to Russian. This paper describes frame creation strategies, coverage, and the process of sense disambiguation. It discusses language-specific issues that complicated the process of building the PropBank and how these challenges were exploited as language-internal guidance for consistency and coherence.
Sarah R. Moeller, Irina Wagner, Martha Palmer, Kathryn Conger, Skatje Myers
LREC3
2019 Linguistic Analysis Improves Neural Metaphor Detection
abstract
In the field of metaphor detection, deep learning systems are the ubiquitous and achieve strong performance on many tasks.However, due to the complicated procedures for manually identifying metaphors, the datasets available are relatively small and fraught with complications.We show that using syntactic features and lexical resources can automatically provide additional high-quality training data for metaphoric language, and this data can cover gaps and inconsistencies in metaphor annotation, improving state-of-the-art word-level metaphor identification.This novel application of automatically improving training data improves classification across numerous tasks, and reconfirms the necessity of high-quality data for deep learning frameworks.
Kevin Stowe, Sarah R. Moeller, Laura A. Michaelis, Martha Palmer
CoNLL4
2018 Bayesian Verb Sense Clustering
abstract
This work performs verb sense induction and clustering based on observed syntactic distributions in a large corpus. VerbNet is a hierarchical clustering of verbs and a useful semantic resource. We address the main drawbacks of VerbNet, by proposing a Bayesian model to build VerbNet-like clusters automatically and with full coverage. Relative to the prior state of the art, we improve accuracy on verb sense induction by over 20% absolute F1. We then propose a new model, inspired by the positive pointwise mutual information (PPMI). Our PPMI-based mixture model permits an extremely efficient sampler, while improving performance. Our best model shows a 4.5% absolute F1 improvement over the best non-PPMI model, with over an order of magnitude less computation time. Though this model is inspired by clustering verb senses, it may be applicable in other situations where multiple items are being sampled as a group.
Daniel W. Peterson, Martha Palmer
AAAI2
2018 Automatically Extracting Qualia Relations for the Rich Event Ontology
abstract
Commonsense, real-world knowledge about the events that entities or “things in the world” are typically involved in, as well as part-whole relationships, is valuable for allowing computational systems to draw everyday inferences about the world. Here, we focus on automatically extracting information about (1) the events that typically bring about certain entities (origins), (2) the events that are the typical functions of entities, and (3) part-whole relationships in entities. These correspond to the agentive, telic and constitutive qualia central to the Generative Lexicon. We describe our motivations and methods for extracting these qualia relations from the Suggested Upper Merged Ontology (SUMO) and show that human annotators overwhelmingly find the information extracted to be reasonable. Because ontologies provide a way of structuring this information and making it accessible to agents and computational systems generally, efforts are underway to incorporate the extracted information to an ontology hub of Natural Language Processing semantic role labeling resources, the Rich Event Ontology.
Ghazaleh Kazeminejad, Claire Bonial, Susan Windisch Brown, Martha Palmer
COLING4
2018 AMR Beyond the Sentence: the Multi-sentence AMR corpus
abstract
There are few corpora that endeavor to represent the semantic content of entire documents. We present a corpus that accomplishes one way of capturing document level semantics, by annotating coreference and similar phenomena (bridging and implicit roles) on top of gold Abstract Meaning Representations of sentence-level semantics. We present a new corpus of this annotation, with analysis of its quality, alongside a plausible baseline for comparison. It is hoped that this Multi-Sentence AMR corpus (MS-AMR) may become a feasible method for developing rich representations of document meaning, useful for tasks such as information extraction and question answering.
Tim O'Gorman, Michael Regan, Kira Griffitt, Ulf Hermjakob, Kevin Knight, Martha Palmer
COLING6
2018 Abstract Meaning Representation of Constructions: The More We Include, the Better the Representation
Claire Bonial, Bianca Badarau, Kira Griffitt, Ulf Hermjakob, Kevin Knight, Tim O'Gorman, Martha Palmer, Nathan Schneider 0001
LREC7
2018 Integrating Generative Lexicon Event Structures into VerbNet
Susan Windisch Brown, James Pustejovsky, Annie Zaenen, Martha Palmer
LREC4
2018 The New Propbank: Aligning Propbank with AMR through POS Unification
Tim O'Gorman, Sameer Pradhan, Martha Palmer, Julia Bonn, Kathryn Conger, James Gung
LREC3
2017 Unsupervised AMR-Dependency Parse Alignment
abstract
In this paper, we introduce an Abstract Meaning Representation (AMR) to Dependency Parse aligner.Alignment is a preliminary step for AMR parsing, and our aligner improves current AMR parser performance.Our aligner involves several different features, including named entity tags and semantic role labels, and uses Expectation-Maximization training.Results show that our aligner reaches an 87.1% F-Score score with the experimental data, and enhances AMR parsing.
Wei-Te Chen, Martha Palmer
EACL (1)2
2017 Coreference annotation and resolution in the Colorado Richly Annotated Full Text (CRAFT) corpus of biomedical journal articles
abstract
BACKGROUND: Coreference resolution is the task of finding strings in text that have the same referent as other strings. Failures of coreference resolution are a common cause of false negatives in information extraction from the scientific literature. In order to better understand the nature of the phenomenon of coreference in biomedical publications and to increase performance on the task, we annotated the Colorado Richly Annotated Full Text (CRAFT) corpus with coreference relations. RESULTS: The corpus was manually annotated with coreference relations, including identity and appositives for all coreferring base noun phrases. The OntoNotes annotation guidelines, with minor adaptations, were used. Interannotator agreement ranges from 0.480 (entity-based CEAF) to 0.858 (Class-B3), depending on the metric that is used to assess it. The resulting corpus adds nearly 30,000 annotations to the previous release of the CRAFT corpus. Differences from related projects include a much broader definition of markables, connection to extensive annotation of several domain-relevant semantic classes, and connection to complete syntactic annotation. Tool performance was benchmarked on the data. A publicly available out-of-the-box, general-domain coreference resolution system achieved an F-measure of 0.14 (B3), while a simple domain-adapted rule-based system achieved an F-measure of 0.42. An ensemble of the two reached F of 0.46. Following the IDENTITY chains in the data would add 106,263 additional named entities in the full 97-paper corpus, for an increase of 76% percent in the semantic classes of the eight ontologies that have been annotated in earlier versions of the CRAFT corpus. CONCLUSIONS: The project produced a large data set for further investigation of coreference and coreference resolution in the scientific literature. The work raised issues in the phenomenon of reference in this domain and genre, and the paper proposes that many mentions that would be considered generic in the general domain are not generic in the biomedical domain due to their referents to specific classes in domain-specific ontologies. The comparison of the performance of a publicly available and well-understood coreference resolution system with a domain-adapted system produced results that are consistent with the notion that the requirements for successful coreference resolution in this genre are quite different from those of the general domain, and also suggest that the baseline performance difference is quite large.
Kevin Cohen 0001, Arrick Lanfranchi, Miji Joo-young Choi, Michael Bada, William A. Baumgartner Jr., Natalya Panteleyeva, Karin Verspoor, Martha Palmer, Lawrence Hunter
BMC Bioinform.8
2016 Adam Kilgarriff's Legacy to Computational Linguistics and Beyond
Roger Evans, Alexander F. Gelbukh, Gregory Grefenstette, Patrick Hanks, Milos Jakubícek, Diana McCarthy, Martha Palmer, Ted Pedersen, Michael Rundell, Pavel Rychlý, Serge Sharoff, David Tugwell
CICLing (1)7
2016 Linguistic features for Hindi light verb construction identification
abstract
Light verb constructions (LVC) in Hindi are highly productive. If we can distinguish a case such as nirnay lenaa ‘decision take; decide’ from an ordinary verb-argument combination kaagaz lenaa ‘paper take; take (a) paper’,it has been shown to aid NLP applications such as parsing (Begum et al., 2011) and machine translation (Pal et al., 2011). In this paper, we propose an LVC identification system using language specific features for Hindi which shows an improvement over previous work(Begum et al., 2011). To build our system, we carry out a linguistic analysis of Hindi LVCs using Hindi Treebank annotations and propose two new features that are aimed at capturing the diversity of Hindi LVCs in the corpus. We find that our model performs robustly across a diverse range of LVCs and our results underscore the importance of semantic features, which is in keeping with the findings for English. Our error analysis also demonstrates that our classifier can be used to further refine LVC annotations in the Hindi Treebank and make them more consistent across the board.
Ashwini Vaidya, Sumeet Agarwal, Martha Palmer
COLING3
2016 A Proposition Bank of Urdu
Maaz Anwar, Riyaz A. Bhat, Dipti Misra Sharma, Ashwini Vaidya, Martha Palmer, Tafseer Ahmed
LREC5
2016 Comprehensive and Consistent PropBank Light Verb Annotation
Claire Bonial, Martha Palmer
LREC2
2016 Large Multi-lingual, Multi-level and Multi-genre Annotation Corpus
Xuansong Li, Martha Palmer, Nianwen Xue, Lance A. Ramshaw, Mohamed Maamouri, Ann Bies, Kathryn Conger, Stephen Grimes, Stephanie M. Strassel
LREC2
2015 English Light Verb Construction Identification Using Lexical Knowledge
abstract
This research describes the development of a supervised classifier of English light verb constructions, for example, "take a walk" and "make a speech." This classifier relies on features from dependency parses, OntoNotes sense tags, WordNet hypernyms and WordNet lexical file information. Evaluation shows that this system achieves an 89% F1 score (four points above the state of the art) on the BNC test set used by Tu & Roth (2011), and an F1 score of 80.68 on the OntoNotes test set, which is significantly more challenging. We attribute the superior F1 score to the use of our rich linguistic features, including the use of WordNet synset and hypernym relations for the detection of previously unattested light verb constructions. We describe the classifier and its features, as well as the characteristics of the OntoNotes light verb construction test set, which relies on linguistically motivated PropBank annotation.
Wei-Te Chen, Claire Bonial, Martha Palmer
AAAI3
2014 A Step-wise Usage-based Method for Inducing Polysemy-aware Verb Classes
abstract
We present an unsupervised method for in-ducing verb classes from verb uses in giga-word corpora. Our method consists of two clustering steps: verb-specific seman-tic frames are first induced by clustering verb uses in a corpus and then verb classes are induced by clustering these frames. By taking this step-wise approach, we can not only generate verb classes based on a massive amount of verb uses in a scalable manner, but also deal with verb polysemy, which is bypassed by most of the previous studies on verb clustering. In our exper-iments, we acquire semantic frames and verb classes from two giga-word corpora, the larger comprising 20 billion words. The effectiveness of our approach is veri-fied through quantitative evaluations based on polysemy-aware gold-standard data. 1
Daisuke Kawahara, Daniel W. Peterson, Martha Palmer
ACL (1)3
2014 Verb Clustering for Brazilian Portuguese
Carolina Scarton, Lin Sun 0003, Karin Kipper Schuler, Magali Sanches Duran, Martha Palmer, Anna Korhonen
CICLing (1)5
2014 Inducing Example-based Semantic Frames from a Massive Amount of Verb Uses
abstract
We present an unsupervised method for inducing semantic frames from verb uses in giga-word corpora.Our semantic frames are verb-specific example-based frames that are distinguished according to their senses.We use the Chinese Restaurant Process to automatically induce these frames from a massive amount of verb instances.In our experiments, we acquire broad-coverage semantic frames from two giga-word corpora, the larger comprising 20 billion words.Our experimental results indicate the effectiveness of our approach.
Daisuke Kawahara, Daniel W. Peterson, Octavian Popescu, Martha Palmer
EACL4
2014 PropBank: Semantics of New Predicate Types
Claire Bonial, Julia Bonn, Kathryn Conger, Jena D. Hwang, Martha Palmer
LREC5
2014 Criteria for Identifying and Annotating Caused Motion Constructions in Corpus Data
Jena D. Hwang, Annie Zaenen, Martha Palmer
LREC3
2014 Single Classifier Approach for Verb Sense Disambiguation based on Generalized Features
Daisuke Kawahara, Martha Palmer
LREC2
2014 Focusing Annotation for Semantic Role Labeling
Daniel W. Peterson, Martha Palmer, Shumin Wu
LREC2
2014 Mapping CPA Patterns onto OntoNotes Senses
Octavian Popescu, Martha Palmer, Patrick Hanks
LREC2
2014 Not an Interlingua, But Close: Comparison of English AMRs to Chinese and Czech
Nianwen Xue, Ondrej Bojar, Jan Hajic 0001, Martha Palmer, Zdenka Uresová, Xiuhong Zhang
LREC4
2014 Temporal Annotation in the Clinical Domain
abstract
This article discusses the requirements of a formal specification for the annotation of temporal information in clinical narratives. We discuss the implementation and extension of ISO-TimeML for annotating a corpus of clinical notes, known as the THYME corpus. To reflect the information task and the heavily inference-based reasoning demands in the domain, a new annotation guideline has been developed, "the THYME Guidelines to ISO-TimeML (THYME-TimeML)". To clarify what relations merit annotation, we distinguish between linguistically-derived and inferentially-derived temporal orderings in the text. We also apply a top performing TempEval 2013 system against this new resource to measure the difficulty of adapting systems to the clinical domain. The corpus is available to the community and has been proposed for use in a SemEval 2015 task.
William F. Styler IV, Steven Bethard, Sean Finan, Martha Palmer, Sameer Pradhan, Piet C. de Groen, Bradley James Erickson, Timothy A. Miller, Chen Lin 0002, Guergana K. Savova, James Pustejovsky
Trans. Assoc. Comput. Linguistics4
2013 Panel: Shared Resources, Shared Code, and Shared Activities in Clinical Natural Language Processing
Guergana K. Savova, Wendy W. Chapman, Noémie Elhadad, Martha Palmer
AMIA4
2013 The VerbCorner Project: Toward an Empirically-Based Semantic Decomposition of Verbs
abstract
This research describes efforts to use crowdsourcing to improve the validity of the semantic predicates in VerbNet, a lexicon of about 6300 English verbs.The current semantic predicates can be thought of semantic primitives, into which the concepts denoted by a verb can be decomposed.For example, the verb spray (of the Spray class), involves the predicates MOTION, NOT, and LOCATION, where the event can be decomposed into an AGENT causing a THEME that was originally not in a particular location to now be in that location.Although VerbNet's predicates are theoretically well-motivated, systematic empirical data is scarce.This paper describes a recently-launched attempt to address this issue with a series of human judgment tasks, posed to subjects in the form of games.
Joshua K. Hartshorne, Claire Bonial, Martha Palmer
EMNLP3
2013 Semantic Role Labeling
Martha Palmer, Ivan Titov 0001, Shumin Wu
HLT-NAACL1
2013 Towards comprehensive syntactic and semantic annotations of the clinical narrative
abstract
OBJECTIVE: To create annotated clinical narratives with layers of syntactic and semantic labels to facilitate advances in clinical natural language processing (NLP). To develop NLP algorithms and open source components. METHODS: Manual annotation of a clinical narrative corpus of 127 606 tokens following the Treebank schema for syntactic information, PropBank schema for predicate-argument structures, and the Unified Medical Language System (UMLS) schema for semantic information. NLP components were developed. RESULTS: The final corpus consists of 13 091 sentences containing 1772 distinct predicate lemmas. Of the 766 newly created PropBank frames, 74 are verbs. There are 28 539 named entity (NE) annotations spread over 15 UMLS semantic groups, one UMLS semantic type, and the Person semantic category. The most frequent annotations belong to the UMLS semantic groups of Procedures (15.71%), Disorders (14.74%), Concepts and Ideas (15.10%), Anatomy (12.80%), Chemicals and Drugs (7.49%), and the UMLS semantic type of Sign or Symptom (12.46%). Inter-annotator agreement results: Treebank (0.926), PropBank (0.891-0.931), NE (0.697-0.750). The part-of-speech tagger, constituency parser, dependency parser, and semantic role labeler are built from the corpus and released open source. A significant limitation uncovered by this project is the need for the NLP community to develop a widely agreed-upon schema for the annotation of clinical concepts and their relations. CONCLUSIONS: This project takes a foundational step towards bringing the field of clinical NLP up to par with NLP in the general domain. The corpus creation and NLP components provide a resource for research and application development that would have been previously impossible.
Daniel Albright, Arrick Lanfranchi, Anwen Fredriksen, William F. Styler IV, Colin Warner, Jena D. Hwang, Jinho D. Choi, Dmitriy Dligach, Rodney D. Nielsen, James H. Martin, Wayne H. Ward, Martha Palmer, Guergana K. Savova
J. Am. Medical Informatics Assoc.12
2012 Verb Classification using Distributional Similarity in Syntactic and Semantic Structures
Danilo Croce, Alessandro Moschitti, Roberto Basili 0001, Martha Palmer
ACL (1)4
2012 Learning to Tutor Like a Tutor: Ranking Questions in Context
Lee Becker, Martha Palmer, Sarel van Vuuren, Wayne H. Ward
ITS2
2012 Foundations of a Multilayer Annotation Framework for Twitter Communications During Crisis Events
William J. Corvey, Sudha Verma, Sarah Vieweg, Martha Palmer, James H. Martin
LREC4
2012 Empty Argument Insertion in the Hindi PropBank
Ashwini Vaidya, Jinho D. Choi, Martha Palmer, Bhuvana Narasimhan
LREC3
2012 A corpus of full-text journal articles is a robust evaluation tool for revealing differences in performance of biomedical natural language processing tools
abstract
BACKGROUND: We introduce the linguistic annotation of a corpus of 97 full-text biomedical publications, known as the Colorado Richly Annotated Full Text (CRAFT) corpus. We further assess the performance of existing tools for performing sentence splitting, tokenization, syntactic parsing, and named entity recognition on this corpus. RESULTS: Many biomedical natural language processing systems demonstrated large differences between their previously published results and their performance on the CRAFT corpus when tested with the publicly available models or rule sets. Trainable systems differed widely with respect to their ability to build high-performing models based on this data. CONCLUSIONS: The finding that some systems were able to train high-performing models based on this corpus is additional evidence, beyond high inter-annotator agreement, that the quality of the CRAFT corpus is high. The overall poor performance of various systems indicates that considerable work needs to be done to enable natural language processing systems to work well when the input is full-text journal articles. The CRAFT corpus provides a valuable resource to the biomedical natural language processing community for evaluation and training of new models for biomedical full text publications.
Karin Verspoor, Kevin Cohen 0001, Arrick Lanfranchi, Colin Warner, Helen L. Johnson 0001, Christophe Roeder, Jinho D. Choi, Christopher S. Funk, Yuriy Malenkiy, Miriam Eckert, Nianwen Xue, William A. Baumgartner Jr., Michael Bada, Martha Palmer, Lawrence Hunter
BMC Bioinform.14
2011 Natural Language Processing to the Rescue? Extracting "Situational Awareness" Tweets During Mass Emergency
Sudha Verma, Sarah Vieweg, William J. Corvey, Leysia Palen, James H. Martin, Martha Palmer, Aaron Schram, Kenneth M. Anderson
ICWSM6
2010 Empty Categories in a Hindi Treebank
Archna Bhatia, Rajesh Bhatt, Bhuvana Narasimhan, Martha Palmer, Owen Rambow, Dipti Misra Sharma, Michael Tepper, Ashwini Vaidya, Fei Xia 0004
LREC4
2010 Number or Nuance: Which Factors Restrict Reliable Word Sense Annotation?
Susan Windisch Brown, Travis Rood, Martha Palmer
LREC3
2010 Propbank Instance Annotation Guidelines Using a Dedicated Editor, Jubilee
Jinho D. Choi, Claire Bonial, Martha Palmer
LREC3
2010 Propbank Frameset Annotation Guidelines Using a Dedicated Editor, Cornerstone
Jinho D. Choi, Claire Bonial, Martha Palmer
LREC3
2010 A Road Map for Interoperable Language Resource Metadata
Christopher Cieri, Khalid Choukri, Nicoletta Calzolari, D. Terence Langendoen, Johannes Leveling, Martha Palmer, Nancy Ide, James Pustejovsky
LREC6
2009 Towards Temporal Relation Discovery from the Clinical Narrative
Guergana K. Savova, Steven Bethard, William F. Styler IV, James H. Martin, Martha Palmer, James J. Masanz, Wayne H. Ward
AMIA5
2009 Adding semantic roles to the Chinese Treebank
abstract
Abstract We report work on adding semantic role labels to the Chinese Treebank, a corpus already annotated with phrase structures. The work involves locating all verbs and their nominalizations in the corpus, and semi-automatically adding semantic role labels to their arguments, which are constituents in a parse tree. Although the same procedure is followed, different issues arise in the annotation of verbs and nominalized predicates. For verbs, identifying their arguments is generally straightforward given their syntactic structure in the Chinese Treebank as they tend to occupy well-defined syntactic positions. Our discussion focuses on the syntactic variations in the realization of the arguments as well as our approach to annotating dislocated and discontinuous arguments. In comparison, identifying the arguments for nominalized predicates is more challenging and we discuss criteria and procedures for distinguishing arguments from non-arguments. In particular we focus on the role of support verbs as well as the relevance of event/result distinctions in the annotation of the predicate-argument structure of nominalized predicates. We also present our approach to taking advantage of the syntactic structure in the Chinese Treebank to bootstrap the predicate-argument structure annotation of verbs. Finally, we discuss the creation of a lexical database of frame files and its role in guiding predicate-argument annotation. Procedures for ensuring annotation consistency and inter-annotator agreement evaluation results are also presented.
Nianwen Xue, Martha Palmer
Nat. Lang. Eng.2
2008 Annotating Students' Understanding of Science Concepts
Rodney D. Nielsen, Wayne H. Ward, James H. Martin, Martha Palmer
LREC4
2008 A Pilot Arabic Propbank
Martha Palmer, Olga Babko-Malaya, Ann Bies, Mona T. Diab, Mohamed Maamouri, Aous Mansouri, Wajdi Zaghouani
LREC1
2007 Can Semantic Roles Generalize Across Genres?
Szu-ting Yi, Edward Loper, Martha Palmer
HLT-NAACL3
2007 Making fine-grained and coarse-grained sense distinctions, both manually and automatically
abstract
In this paper we discuss a persistent problem arising from polysemy: namely the difficulty of finding consistent criteria for making fine-grained sense distinctions, either manually or automatically. We investigate sources of human annotator disagreements stemming from the tagging for the English Verb Lexical Sample Task in the SENSEVAL-2 exercise in automatic Word Sense Disambiguation. We also examine errors made by a high-performing maximum entropy Word Sense Disambiguation system we developed. Both sets of errors are at least partially reconciled by a more coarse-grained view of the senses, and we present the groupings we use for quantitative coarse-grained evaluation as well as the process by which they were created. We compare the system's performance with our human annotator performance in light of both fine-grained and coarse-grained sense distinctions and show that well-defined sense groups can be of value in improving word sense disambiguation by both humans and machines.
Martha Palmer, Hoa Trang Dang, Christiane Fellbaum
Nat. Lang. Eng.1
2006 Aligning Features with Sense Distinction Dimensions
Nianwen Xue, Jinying Chen, Martha Palmer
ACL3
2006 Extending VerbNet with Novel Verb Classes
Karin Kipper Schuler, Anna Korhonen, Neville Ryant, Martha Palmer
LREC4
2006 An Empirical Study of the Behavior of Active Learning for Word Sense Disambiguation
Jinying Chen, Andrew I. Schein, Lyle H. Ungar, Martha Palmer
HLT-NAACL4
2006 OntoNotes: The 90% Solution
Eduard H. Hovy, Mitchell P. Marcus, Martha Palmer, Lance A. Ramshaw, Ralph M. Weischedel
HLT-NAACL3
2005 The Role of Semantic Roles in Disambiguating Verb Senses
abstract
We describe an automatic Word Sense Disambiguation (WSD) system that disambiguates verb senses using syntactic and semantic features that encode information about predicate arguments and semantic classes. Our system performs at the best published accuracy on the English verbs of Senseval-2. We also experiment with using the gold-standard predicate-argument labels from PropBank for disambiguating fine-grained WordNet senses and course-grained PropBank framesets, and show that disambiguation of verb senses can be further improved with better extraction of semantic roles.
Hoa Trang Dang, Martha Palmer
ACL2
2005 Machine Translation Using Probabilistic Synchronous Dependency Insertion Grammars
abstract
Syntax-based statistical machine translation (MT) aims at applying statistical models to structured data. In this paper, we present a syntax-based statistical machine translation system based on a probabilistic synchronous dependency insertion grammar. Synchronous dependency insertion grammars are a version of synchronous grammars defined on dependency trees. We first introduce our approach to inducing such a grammar from parallel corpora. Second, we describe the graphical model for the machine translation task, which can also be viewed as a stochastic tree-to-tree transducer. We introduce a polynomial time decoding algorithm for the model. We evaluate the outputs of our MT system using the NIST and Bleu automatic MT evaluation software. The result shows that our system outperforms the baseline system based on the IBM models in both translation speed and quality.
Yuan Ding 0005, Martha Palmer
ACL2
2005 The Integration of Syntactic Parsing and Semantic Role Labeling
Szu-ting Yi, Martha Palmer
CoNLL2
2005 Automatic Semantic Role Labeling for Chinese Verbs
Nianwen Xue, Martha Palmer
IJCAI2
2005 Towards Robust High Performance Word Sense Disambiguation of English Verbs Using Rich Linguistic Features
Jinying Chen, Martha Palmer
IJCNLP2
2005 Automatically Generating Tree Adjoining Grammars from Abstract Specifications
abstract
The paper describes a system that can automatically generate tree adjoining grammars from abstract specifications. Our system is based on the use of tree descriptions to specify a grammar by separately defining pieces of tree structure that encode independent syntactic principles. Various individual specifications are then combined to form the elementary trees of the grammar. The system enables efficient development and maintenance of a grammar, and also allows underlying linguistic constructions (such as wh-movement) to be expressed explicitly. We have carefully designed our system to be as language independent as possible and tested its performance by constructing both English and Chinese grammars, with significant reductions in grammar development time. Provably consistent abstract specifications for different languages also offer unique opportunities for investigating how languages relate to themselves and to each other. For instance, the impact of a linguistic structure such as wh-movement can be traced from its specification to the descriptions that it combines with, to its actual realization in trees. By focusing on syntactic properties at a higher level, our approach allowed a unique comparison of our English and Chinese grammars.
Fei Xia 0004, Martha Palmer, K. Vijay-Shanker
Comput. Intell.2
2005 The Proposition Bank: An Annotated Corpus of Semantic Roles
abstract
The Proposition Bank project takes a practical approach to semantic representation, adding a layer of predicate-argument information, or semantic role labels, to the syntactic structures of the Penn Treebank. The resulting resource can be thought of as shallow, in that it does not represent coreference, quantification, and many other higher-order phenomena, but also broad, in that it covers every instance of every verb in the corpus and allows representative statistics to be calculated. We discuss the criteria used to define the sets of semantic roles used in the annotation process and to analyze the frequency of syntactic/semantic alternations in the corpus. We describe an automatic system for semantic role tagging trained on the corpus and discuss the effect on its performance of various types of information, including a comparison of full syntactic parsing with a flat representation and the contribution of the empty “trace” categories of the treebank.
Martha Palmer, Paul R. Kingsbury, Daniel Gildea
Comput. Linguistics1
2005 The Penn Chinese TreeBank: Phrase structure annotation of a large corpus
abstract
With growing interest in Chinese Language Processing, numerous NLP tools (e.g., word segmenters, part-of-speech taggers, and parsers) for Chinese have been developed all over the world. However, since no large-scale bracketed corpora are available to the public, these tools are trained on corpora with different segmentation criteria, part-of-speech tagsets and bracketing guidelines, and therefore, comparisons are difficult. As a first step towards addressing this issue, we have been preparing a large bracketed corpus since late 1998. The first two installments of the corpus, 250 thousand words of data, fully segmented, POS-tagged and syntactically bracketed, have been released to the public via LDC ( www.ldc.upenn.edu ). In this paper, we discuss several Chinese linguistic issues and their implications for our treebanking efforts and how we address these issues when developing our annotation guidelines. We also describe our engineering strategies to improve speed while ensuring annotation quality.
Nianwen Xue, Fei Xia 0004, Fu-Dong Chiou, Martha Palmer
Nat. Lang. Eng.4
2004 Chinese Verb Sense Discrimination Using an EM Clustering Model with Rich Linguistic Features
abstract
This paper discusses the application of the Expectation-Maximization (EM) clustering algorithm to the task of Chinese verb sense discrimination. The model utilized rich linguistic features that capture predicate-argument structure information of the target verbs. A semantic taxonomy for Chinese nouns, which was built semi-automatically based on two electronic Chinese semantic dictionaries, was used to provide semantic features for the model. Purity and normalized mutual information were used to evaluate the clustering performance on 12 Chinese verbs. The experimental results show that the EM clustering model can learn sense or sense group distinctions for most of the verbs successfully. We further enhanced the model with certain fine-grained semantic categories called lexical sets. Our results indicate that these lexical sets improve the model's performance for the three most challenging verbs chosen from the first set of experiments.
Jinying Chen, Martha Palmer
ACL2
2004 Putting Meaning into Your Trees
Martha Palmer
CoNLL1
2004 Calibrating Features for Semantic Role Labeling
Nianwen Xue, Martha Palmer
EMNLP2
2004 Using a Smoothing Maximum Entropy Model for Chinese Nominal Entity Tagging
Jinying Chen, Nianwen Xue, Martha Palmer
IJCNLP3
2004 Automatic Learning of Parallel Dependency Treelet Pairs
Yuan Ding 0005, Martha Palmer
IJCNLP2
2004 Extending a Verb-lexicon Using a Semantically Annotated Corpus
Karin Kipper Schuler, Benjamin Snyder, Martha Palmer
LREC3
2004 A Morphological Tagger for Korean: Statistical Tagging Combined with Corpus-Based Morphological Rule Application
Chung-hye Han, Martha Palmer
Mach. Transl.2
2003 An algorithm for word-level alignment of parallel dependency trees
abstract
Structural divergence presents a challenge to the use of syntax in statistical machine translation. We address this problem with a new algorithm for alignment of loosely matched non-isomorphic dependency trees. The algorithm selectively relaxes the constraints of the two tree structures while keeping computational complexity polynomial in the length of the sentences. Experimentation with a large Chinese-English corpus shows an improvement in alignment results over the unstructured models of (Brown et al., 1993).
Yuan Ding 0005, Daniel Gildea, Martha Palmer
MTSummit3
2003 Microplanning with Communicative Intentions: The SPUD System
abstract
The process of microplanning in natural language generation (NLG) encompasses a range of problems in which a generator must bridge underlying domain‐specific representations and general linguistic representations. These problems include constructing linguistic referring expressions to identify domain objects, selecting lexical items to express domain concepts, and using complex linguistic constructions to concisely convey related domain facts. In this paper, we argue that such problems are best solved through a uniform, comprehensive, declarative process. In our approach, the generator directly explores a search space for utterances described by a linguistic grammar. At each stage of search, the generator uses a model of interpretation, which characterizes the potential links between the utterance and the domain and context, to assess its progress in conveying domain‐specific representations. We further address the challenges for implementation and knowledge representation in this approach. We show how to implement this approach effectively by using the lexicalized tree‐adjoining grammar (LTAG) formalism to connect structure to meaning and using modal logic programming to connect meaning to context. We articulate a detailed methodology for designing grammatical and conceptual resources which the generator can use to achieve desired microplanning behavior in a specified domain. In describing our approach to microplanning, we emphasize that we are in fact realizing a deliberative process of goal‐directed activity. As we formulate it, interpretation offers a declarative representation of a generator's communicative intent. It associates the concrete linguistic structure planned by the generator with inferences that show how the meaning of that structure communicates needed information about some application domain in the current discourse context. Thus, interpretations areplansthat the microplanner constructs and outputs. At the same time, communicative intent representations provide arich and uniform resourcefor theprocessof NLG. Using representations of communicative intent, a generator can augment the syntax, semantics, and pragmatics of an incomplete sentence simultaneously, and can work incrementally toward solutions for the various problems of microplanning.
Matthew Stone, Christine Doran, Bonnie L. Webber, Tonia Bleam, Martha Palmer
Comput. Intell.5
2002 The Necessity of Parsing for Predicate Argument Recognition
abstract
Broad-coverage corpora annotated with semantic role, or argument structure, information are becoming available for the first time. Statistical systems have been trained to automatically label semantic roles from the output of statistical parsers on unannotated text. In this paper, we quantify the effect of parser accuracy on these systems' performance, and examine the question of whether a flatter "chunked" representation of the input can be as effective for the purposes of semantic role identification.
Daniel Gildea, Martha Palmer
ACL2
2002 Simple Features for Chinese Word Sense Disambiguation
Hoa Trang Dang, Ching-yi Chia, Martha Palmer, Fu-Dong Chiou
COLING3
2002 Building a Large-Scale Annotated Chinese Corpus
Nianwen Xue, Fu-Dong Chiou, Martha Palmer
COLING3
2002 From Resources to Applications. Designing the Multilingual ISLE Lexical Entry
Sue Atkins, Núria Bel, Francesca Bertagna, Pierrette Bouillon, Nicoletta Calzolari, Christiane Fellbaum, Ralph Grishman, Alessandro Lenci, Catherine Macleod, Martha Palmer, Gregor Thurmair, Marta Villegas, Antonio Zampolli
LREC10
2002 Standards & best practice for multilingual computational lexicons: ISLE MILE and more
Nicoletta Calzolari, Ralph Grishman, Martha Palmer
LREC3
2002 Development and Evaluation of a Korean Treebank and its Application to NLP
Chung-hye Han, Na-Rare Han, Eon-Suk Ko, Martha Palmer
LREC4
2002 From TreeBank to PropBank
Paul R. Kingsbury, Martha Palmer
LREC2
2002 Penn Korean Treebank : Development and Evaluation
Chung-hye Han, Na-Rare Han, Eon-Suk Ko, Martha Palmer, Heejong Yi
PACLIC4
2001 Automatically Extracting and Comparing Lexicalized Grammars for Different Languages
Fei Xia 0004, Chung-hye Han, Martha Palmer, Aravind K. Joshi
IJCAI3
2000 Integrating compositional semantics into a verb lexicon
Hoa Trang Dang, Karin Kipper Schuler, Martha Palmer
COLING3
2000 A Uniform Method of Grammar Extraction and Its Applications
abstract
Grammars are core elements of many NLP applications. In this paper, we present a system that automatically extracts lexicalized grammars from annotated corpora. The data produced by this system have been used in several tasks, such as training NLP tools (such as Supertaggers) and estimating the coverage of hand-crafted grammars. We report experimental results on two of those tasks and compare our approaches with related work.
Fei Xia 0004, Martha Palmer, Aravind K. Joshi
EMNLP2
2000 Semantic Tagging for the Penn Treebank
Martha Palmer, Hoa Trang Dang, Joseph Rosenzweig
LREC1
2000 Developing Guidelines and Ensuring Consistency for Chinese Text Annotation
Fei Xia 0004, Martha Palmer, Nianwen Xue, Mary Ellen Okurowski, John Kovarik, Fu-Dong Chiou, Shizhe Huang, Tony Kroch, Mitchell P. Marcus
LREC2
1996 A Statistically Emergent Approach for Language Processing: Application to Modeling Context Effects in Ambiguous Chinese Word Boundary Perception
Kok-Wee Gan, Martha Palmer, Kim-Teng Lua
Comput. Linguistics2
1995 Verb semantics for English-Chinese translation
Martha Palmer, Zhibiao Wu
Mach. Transl.1
1994 Verb Semantics and Lexical Selection
abstract
This paper will focus on the semantic representation of verbs in computer systems and its impact on lexical selection problems in machine translation (MT). Two groups of English and Chinese verbs are examined to show that lexical selection must be based on interpretation of the sentences as well as selection restrictions placed on the verb arguments. A novel representation scheme is suggested, and is compared to representations with selection restrictions used in transfer-based MT. We see our approach as closely aligned with knowledge-based MT approaches (KBMT), and as a separate component that could be incorporated into existing systems. Examples and experimental results will show that, using this scheme, inexact matches can achieve correct lexical selection.
Zhibiao Wu, Martha Palmer
ACL2
1993 The KERNEL Text Understanding System
Martha Palmer, Rebecca J. Passonneau, Carl Weir, Tim Finin
Artif. Intell.1
1990 Integrating Natural Language Processing and Knowledge Based Processing
Rebecca J. Passonneau, Carl Weir, Tim Finin, Martha Palmer
AAAI4
1990 Workshop on the Evaluation of Natural Language Processing Systems
Martha Palmer, Tim Finin
Comput. Linguistics1
1990 Customizing verb definitions for specific semantic domains
Martha Palmer
Mach. Transl.1
1987 Nominalizations in PUNDIT
abstract
This paper describes the treatment of nominalizations in the PUNDIT text processing system. A single semantic definition is used for both nominalizations and the verbs to which they are related, with the same semantic roles, decompositions, and selectional restrictions on the semantic roles. However, because syntactically nominalizations are noun phrases, the processing which produces the semantic representation is different in several respects from that used for clauses. (1) The rules relating the syntactic positions of the constituents to the roles that they can fill are different. (2) The fact that nominalizations are untensed while clauses normally are tensed means that an alternative treatment of time is required for nominalizations. (3) Because none of the arguments of a nominalization is syntactically obligatory, some differences in the control of the filling of roles are required, in particular, roles can be filled as part of reference resolution for the nominalization. The differences in processing are captured by allowing the semantic interpreter to operate in two different modes, one for clauses, and one for nominalizations. Because many nominalizations are noun-noun compounds, this approach also addresses this problem, by suggesting a way of dealing with one relatively tractable subset of noun-noun compounds.
Deborah A. Dahl, Martha Palmer, Rebecca J. Passonneau
ACL2
1986 Recovering Implicit Information
abstract
This paper describes the SDC PUNDIT, (Prolog UNDerstands Integrated Text), system for processing natural language messages.1 PUNDIT, written in Prolog, is a highly modular system consisting of distinct syntactic, semantic and pragmatics components.Each component draws on one or more sets of data, including a lexicon, a broad-coverage grammar of EngLish, semantic verb decompositions, rules mapping between syntactic and semantic constituents, and a domain model.This paper discusses the communication between the syntactic, semantic and pragmatic modules that is necessary for making implicit linguistic information explicit.The key is letting syntax and semantics recognize missing linguistic entities as implicit entities, so that they can be labelled as such, and referenee resolution can be directed to find specific referents for the entities.In this way the task of making implicit linguistic information explicit becomes a subset of the tasks performed by reference resolution.The success of this approach is dependent on marking missing syntactic constituents as elided and missing semantic roles as ESSENTIAL so that reference resolution can know when to look for referents.
Martha Palmer, Deborah A. Dahl, Rebecca J. Schiffman, Lynette Hirschman, Marcia C. Linebarger, John Dowding
ACL1
1983 Inference-Driven Semantic Analysis
Martha Palmer
AAAI1
1981 A Case for Rule-Driven Semantic Processing
abstract
INTRODUCTION The primary cask of semantic procasein S is to provide an appropriate mapping between the eyntactic conetltuents of a parsed eentence and the arSumente of the semantic predicates implied by the verb. known as the Alignment Problem.[Levin] Section One of thle paper overview of a generally accepted approach to semantic procasein S that soee through eeveral levele of repreeentation to achieve this mappin S . Although somewhat inflexible and cumbersome, the different levels succeed in preeervin S the context sensitive information provided by verb eemantice. Section Two preeente the author'e rule-driven approach which is more uniform and flezible yet etill accommodates context eeneitive conattaints. This approach is baeed on sanetel underlyin S principles for eyntactic methods of introducing semantic arguments and has interesting implications for linguistic theories about case. These implications are dlcueeed in Section Three. A system that implements this approach has been d
Martha Palmer
ACL1
1981 The Design Of A System For Designing Knowledge Representation Systems
James L. Weiner, Martha Palmer
IJCAI2