VLDB 2026 Research / reviewers in the wild / expert
James H. Martin
dblp:56/3331
· DBLP profile ↗
44ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-9654-3864ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 since 2021Software engineering, systems software and programming languages · 2Databases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Building Bridges between Student and Curricular Language: Creating a Corpus of Abstract Meaning Representations for the Classroom
Kristin Wright-Bettner, Zheng Cai, Zekun Zhao, James H. Martin, Jeffrey Flanigan, Martha Palmer |
LREC | 4 |
| 2025 | Enhancing Talk Moves Analysis in Mathematics Tutoring through Classroom Teaching DiscourseabstractHuman tutoring interventions play a crucial role in supporting student learning, improving academic performance, and promoting personal growth. This paper focuses on analyzing mathematics tutoring discourse using talk moves—a framework of dialogue acts grounded in Accountable Talk theory. However, scaling the collection, annotation, and analysis of extensive tutoring dialogues to develop machine learning models is a challenging and resource-intensive task. To address this, we present SAGA22, a compact dataset, and explore various modeling strategies, including dialogue context, speaker information, pretraining datasets, and further fine-tuning. By leveraging existing datasets and models designed for classroom teaching, our results demonstrate that supplementary pretraining on classroom data enhances model performance in tutoring settings, particularly when incorporating longer context and speaker information. Additionally, we conduct extensive ablation studies to underscore the challenges in talk move modeling. Jie Cao 0010, Abhijit Suresh, Jennifer Jacobs 0002, Charis Clevenger, Amanda Howard, Chelsea Brown, Brent Milne, Tom Fischaber, Tamara Sumner, James H. Martin |
COLING | 10 |
| 2025 | Towards Actionable Pedagogical Feedback: A Multi-Perspective Analysis of Mathematics Teaching and Tutoring Dialogue
Jannatun Naim, Jie Cao 0010, Fareen Tasneem, Jennifer Jacobs 0002, Brent Milne, James H. Martin, Tamara Sumner |
EDM | 6 |
| 2024 | Linear Cross-document Event Coreference Resolution with X-AMRabstractEvent Coreference Resolution (ECR) as a pairwise mention classification task is expensive both for automated systems and manual annotations. The task’s quadratic difficulty is exacerbated when using Large Language Models (LLMs), making prompt engineering for ECR prohibitively costly. In this work, we propose a graphical representation of events, X-AMR, anchored around individual mentions using a cross-document version of Abstract Meaning Representation. We then linearize the ECR with a novel multi-hop coreference algorithm over the event graphs. The event graphs simplify ECR, making it a) LLM cost-effective, b) compositional and interpretable, and c) easily annotated. For a fair assessment, we first enrich an existing ECR benchmark dataset with these event graphs using an annotator-friendly tool we introduce. Then, we employ GPT-4, the newest LLM by OpenAI, for these annotations. Finally, using the ECR algorithm, we assess GPT-4 against humans and analyze its limitations. Through this research, we aim to advance the state-of-the-art for efficient ECR and shed light on the potential shortcomings of current LLMs at this task. Code and annotations: https://github.com/ahmeshaf/gpt_coref Shafiuddin Rehan Ahmed, George Arthur Baker, Evi Judge, Michael Regan, Kristin Wright-Bettner, Martha Palmer, James H. Martin |
LREC/COLING | 7 |
| 2024 | Building a Broad Infrastructure for Uniform Meaning RepresentationsabstractThis paper reports the first release of the UMR (Uniform Meaning Representation) data set. UMR is a graph-based meaning representation formalism consisting of a sentence-level graph and a document-level graph. The sentence-level graph represents predicate-argument structures, named entities, word senses, aspectuality of events, as well as person and number information for entities. The document-level graph represents coreferential, temporal, and modal relations that go beyond sentence boundaries. UMR is designed to capture the commonalities and variations across languages and this is done through the use of a common set of abstract concepts, relations, and attributes as well as concrete concepts derived from words from invidual languages. This UMR release includes annotations for six languages (Arapaho, Chinese, English, Kukama, Navajo, Sanapana) that vary greatly in terms of their linguistic properties and resource availability. We also describe on-going efforts to enlarge this data set and extend it to other genres and modalities. We also briefly describe the available infrastructure (UMR annotation guidelines and tools) that others can use to create similar data sets. Julia Bonn, Matthew J. Buchholz, Jayeol Chun, Andrew Cowell, William Croft 0001, Lukas Denk, Sijia Ge, Jan Hajic 0001, Kenneth Lai, James H. Martin, Skatje Myers, Alexis Palmer, Martha Palmer, Claire Benet Post, James Pustejovsky, Kristine Stenzel, Haibo Sun, Zdenka Uresová, Rosa Vallejos, Jens E. L. Van Gysel, Meagan Vigus, Nianwen Xue, Jin Zhao 0009 |
LREC/COLING | 10 |
| 2024 | Multimodal Cross-Document Event Coreference Resolution Using Linear Semantic Transfer and Mixed-Modality EnsemblesabstractEvent coreference resolution (ECR) is the task of determining whether distinct mentions of events within a multi-document corpus are actually linked to the same underlying occurrence. Images of the events can help facilitate resolution when language is ambiguous. Here, we propose a multimodal cross-document event coreference resolution method that integrates visual and textual cues with a simple linear map between vision and language models. As existing ECR benchmark datasets rarely provide images for all event mentions, we augment the popular ECB+ dataset with event-centric images scraped from the internet and generated using image diffusion models. We establish three methods that incorporate images and text for coreference: 1) a standard fused model with finetuning, 2) a novel linear mapping method without finetuning and 3) an ensembling approach based on splitting mention pairs by semantic and discourse-level difficulty. We evaluate on 2 datasets: the augmented ECB+, and AIDA Phase 1. Our ensemble systems using cross-modal linear mapping establish an upper limit (91.9 CoNLL F1) on ECB+ ECR performance given the preprocessing assumptions used, and establish a novel baseline on AIDA Phase 1. Our results demonstrate the utility of multimodal information in ECR for certain challenging coreference problems, and highlight a need for more multimodal resources in the coreference resolution space. Abhijnan Nath, Huma Jamil, Shafiuddin Rehan Ahmed, George Arthur Baker, Rahul Ghosh, James H. Martin, Nathaniel Blanchard, Nikhil Krishnaswamy |
LREC/COLING | 6 |
| 2024 | "Keep up the good work!": Using Constraints in Zero Shot Prompting to Generate Supportive Teacher ResponsesabstractEducational dialogue systems have been used to support students and teachers for decades.Such systems rely on explicit pedagogicallymotivated dialogue rules.With the ease of integrating large language models (LLMs) into dialogue systems, applications have been arising that directly use model responses without the use of human-written rules, raising concerns about their use in classroom settings.Here, we explore how to constrain LLM outputs to generate appropriate and supportive teacher-like responses.We present results comparing the effectiveness of different constraint variations in a zero-shot prompting setting on a large mathematics classroom corpus.Generated outputs are evaluated with human annotation for Fluency, Relevance, Helpfulness, and Adherence to the provided constraints.Including all constraints in the prompt led to the highest values for Fluency and Helpfulness, and the second highest value for Relevance.The annotation results also demonstrate that the prompts that result in the highest adherence to constraints do not necessarily indicate higher perceived scores for Fluency, Relevance, or Helpfulness.In a direct comparison, all of the non-baseline LLM responses were ranked higher than the actual teacher responses in the corpus over 50% of the time. Margaret Perkoff, Angela Maria Ramirez, Sean von Bayern, Marilyn A. Walker, James H. Martin |
SIGDIAL | 5 |
| 2023 | Navigating Wanderland: Highlighting Off-Task Discussions in Classrooms
Ananya Ganesh, Michael Alan Chang, Rachel Dickler, Michael Regan, Jon Z. Cai, Kristin Wright-Bettner, James Pustejovsky, James H. Martin, Jeffrey Flanigan, Martha Palmer, Katharina Kann |
AIED | 8 |
| 2023 | A Comparative Analysis of Automatic Speech Recognition Errors in Small Group Classroom DiscourseabstractIn collaborative learning environments, effective intelligent learning systems need to accurately analyze and understand the collaborative discourse between learners (i.e., group modeling) to provide adaptive support. We investigate how automatic speech recognition (ASR) errors influence discourse models of small group collaboration in noisy real-world classrooms. Our dataset consisted of 30 students recorded by consumer off-the-shelf microphones (Yeti Blue) while engaging in dyadic- and triadic- collaborative learning in a multi-day STEM curriculum unit. We found that two state-of-the-art ASR systems (Google Speech and OpenAI Whisper) yielded very high word error rates (0.822, 0.847) but very different profiles of error with Google being more conservative, rejecting 38% of utterances instead of 12% for Whisper. Next, we examined how these ASR errors influenced down-stream small group modeling based on pre-trained large language models for three tasks: Abstract Meaning Representation parsing (AMRParsing), on-task/off-task detection (OnTask), and Accountable Productive Talk prediction (TalkMove). As expected, models trained on clean human transcripts yielded degraded performance on all three tasks, measured by the transfer ratio (TR). However, the TR of the specific sentence-level AMRParsing task (.39 - .62) was much lower than that of the abstract discourse-level OnTask (.63- .94) and TalkMove tasks (.64-.72). Furthermore, different training strategies that incorporated ASR transcripts alone or as augmentations of human transcripts increased accuracy for the discourse-level tasks (OnTask and TalkMove) but not AMRParsing. Simulation experiments suggested that the models were tolerant of missing utterances in the dialog context, and that jointly improving ASR accuracy on important word classes (e.g., verbs and nouns) can improve performance across all tasks. Overall, our results provide insights into how different types of NLP-based tasks might be tolerant of ASR errors under extremely noisy conditions and provide suggestions for how to improve accuracy in small group modeling settings for a more equitable, engaging, and adaptive collaborative learning environment. Jie Cao 0010, Ananya Ganesh, Jon Z. Cai, Rosy Southwell, Margaret Perkoff, Michael Regan, Katharina Kann, James H. Martin, Martha Palmer, Sidney K. D'Mello |
UMAP | 8 |
| 2022 | The TalkMoves Dataset: K-12 Mathematics Lesson Transcripts Annotated for Teacher and Student Discursive MovesabstractTranscripts of teaching episodes can be effective tools to understand discourse patterns in classroom instruction. According to most educational experts, sustained classroom discourse is a critical component of equitable, engaging, and rich learning environments for students. This paper describes the TalkMoves dataset, composed of 567 human-annotated K-12 mathematics lesson transcripts (including entire lessons or portions of lessons) derived from video recordings. The set of transcripts primarily includes in-person lessons with whole-class discussions and/or small group work, as well as some online lessons. All of the transcripts are human-transcribed, segmented by the speaker (teacher or student), and annotated at the sentence level for ten discursive moves based on accountable talk theory. In addition, the transcripts include utterance-level information in the form of dialogue act labels based on the Switchboard Dialog Act Corpus. The dataset can be used by educators, policymakers, and researchers to understand the nature of teacher and student discourse in K-12 math classrooms. Portions of this dataset have been used to develop the TalkMoves application, which provides teachers with automated, immediate, and actionable feedback about their mathematics instruction. Abhijit Suresh, Jennifer Jacobs 0002, Charis Harty, Margaret Perkoff, James H. Martin, Tamara Sumner |
LREC | 5 |
| 2021 | Using AI to Promote Equitable Classroom Discussions: The TalkMoves Application
Abhijit Suresh, Jennifer Jacobs 0002, Charis Clevenger, Vivian Lai, Chenhao Tan, James H. Martin, Tamara Sumner |
AIED (2) | 6 |
| 2017 | Abstract Meaning Representation Parsing using LSTM Recurrent Neural NetworksabstractWe present a system which parses sentences into Abstract Meaning Representations, improving state-of-the-art results for this task by more than 5%.AMR graphs represent semantic content using linguistic properties such as semantic roles, coreference, negation, and more.The AMR parser does not rely on a syntactic preparse, or heavily engineered features, and uses five recurrent neural networks as the key architectural components for inferring AMR graphs. William Foland, James H. Martin |
ACL (1) | 2 |
| 2016 | A Tangled Web: The Faint Signals of Deception in Text - Boulder Lies and Truth Corpus (BLT-C)
Franco Salvetti, John B. Lowe, James H. Martin |
LREC | 3 |
| 2013 | Towards comprehensive syntactic and semantic annotations of the clinical narrativeabstractOBJECTIVE: To create annotated clinical narratives with layers of syntactic and semantic labels to facilitate advances in clinical natural language processing (NLP). To develop NLP algorithms and open source components. METHODS: Manual annotation of a clinical narrative corpus of 127 606 tokens following the Treebank schema for syntactic information, PropBank schema for predicate-argument structures, and the Unified Medical Language System (UMLS) schema for semantic information. NLP components were developed. RESULTS: The final corpus consists of 13 091 sentences containing 1772 distinct predicate lemmas. Of the 766 newly created PropBank frames, 74 are verbs. There are 28 539 named entity (NE) annotations spread over 15 UMLS semantic groups, one UMLS semantic type, and the Person semantic category. The most frequent annotations belong to the UMLS semantic groups of Procedures (15.71%), Disorders (14.74%), Concepts and Ideas (15.10%), Anatomy (12.80%), Chemicals and Drugs (7.49%), and the UMLS semantic type of Sign or Symptom (12.46%). Inter-annotator agreement results: Treebank (0.926), PropBank (0.891-0.931), NE (0.697-0.750). The part-of-speech tagger, constituency parser, dependency parser, and semantic role labeler are built from the corpus and released open source. A significant limitation uncovered by this project is the need for the NLP community to develop a widely agreed-upon schema for the annotation of clinical concepts and their relations. CONCLUSIONS: This project takes a foundational step towards bringing the field of clinical NLP up to par with NLP in the general domain. The corpus creation and NLP components provide a resource for research and application development that would have been previously impossible. Daniel Albright, Arrick Lanfranchi, Anwen Fredriksen, William F. Styler IV, Colin Warner, Jena D. Hwang, Jinho D. Choi, Dmitriy Dligach, Rodney D. Nielsen, James H. Martin, Wayne H. Ward, Martha Palmer, Guergana K. Savova |
J. Am. Medical Informatics Assoc. | 10 |
| 2013 | Characterizing and Predicting the Multifaceted Nature of Quality in Educational Web ResourcesabstractEfficient learning from Web resources can depend on accurately assessing the quality of each resource. We present a methodology for developing computational models of quality that can assist users in assessing Web resources. The methodology consists of four steps: 1) a meta-analysis of previous studies to decompose quality into high-level dimensions and low-level indicators, 2) an expert study to identify the key low-level indicators of quality in the target domain, 3) human annotation to provide a collection of example resources where the presence or absence of quality indicators has been tagged, and 4) training of a machine learning model to predict quality indicators based on content and link features of Web resources. We find that quality is a multifaceted construct, with different aspects that may be important to different users at different times. We show that machine learning models can predict this multifaceted nature of quality, both in the context of aiding curators as they evaluate resources submitted to digital libraries, and in the context of aiding teachers as they develop online educational resources. Finally, we demonstrate how computational models of quality can be provided as a service, and embedded into applications such as Web search. Philipp G. Wetzler, Steven Bethard, Heather Leary, Kirsten R. Butcher, Soheil Danesh Bahreini, James H. Martin, Tamara Sumner |
ACM Trans. Interact. Intell. Syst. | 7 |
| 2012 | Blogs as a collective war diaryabstractDisaster-related research in human-centered computing has typically focused on the shorter-term, emergency period of a disaster event, whereas effects of some crises are long-term, lasting years. Social media archived on the Internet provides researchers the opportunity to examine societal reactions to a disaster over time. In this paper we examine how blogs written during a protracted conflict might reflect a collective view of the event. The sheer amount of data originating from the Internet about a significant event poses a challenge to researchers; we employ topic modeling and pronoun analysis as methods to analyze such large-scale data. First, we discovered that blog war topics temporally tracked the actual, measurable violence in the society suggesting that blog content can be an indicator of the health or state of the affected population. We also found that people exhibited a collective identity when they blogged about war, as evidenced by a higher use of first-person plural pronouns compared to blogging on other topics. Blogging about daily life decreased as violence in the society increased; when violence waned, there was a resurgence of daily life topics, potentially illustrating how a society returns to normalcy. Gloria Mark, Mossaab Bagdouri, Leysia Palen, James H. Martin, Ban Al-Ani, Kenneth M. Anderson |
CSCW | 4 |
| 2012 | Foundations of a Multilayer Annotation Framework for Twitter Communications During Crisis Events
William J. Corvey, Sudha Verma, Sarah Vieweg, Martha Palmer, James H. Martin |
LREC | 5 |
| 2011 | Natural Language Processing to the Rescue? Extracting "Situational Awareness" Tweets During Mass Emergency
Sudha Verma, Sarah Vieweg, William J. Corvey, Leysia Palen, James H. Martin, Martha Palmer, Aaron Schram, Kenneth M. Anderson |
ICWSM | 5 |
| 2009 | Towards Temporal Relation Discovery from the Clinical Narrative
Guergana K. Savova, Steven Bethard, William F. Styler IV, James H. Martin, Martha Palmer, James J. Masanz, Wayne H. Ward |
AMIA | 4 |
| 2009 | Recognizing entailment in intelligent tutoring systemsabstractAbstract This paper describes a new method for recognizing whether a student's response to an automated tutor's question entails that they understand the concepts being taught. We demonstrate the need for a finer-grained analysis of answers than is supported by current tutoring systems or entailment databases and describe a new representation for reference answers that addresses these issues, breaking them into detailed facets and annotating their entailment relationships to the student's answer more precisely. Human annotation at this detailed level still results in substantial interannotator agreement (86.2%), with a kappa statistic of 0.728. We also present our current efforts to automatically assess student answers, which involves training machine learning classifiers on features extracted from dependency parses of the reference answer and student's response and features derived from domain-independent lexical statistics. Our system's performance, as high as 75.5% accuracy within domain and 68.8% out of domain, is very encouraging and confirms the approach is feasible. Another significant contribution of this work is that it represents a significant step in the direction of providing domain-independent semantic assessment of answers. No prior work in the area of tutoring or educational assessment has attempted to build such domain-independent systems. They have virtually all required hundreds of examples of learner answers for each new question in order to train aspects of their systems or to hand-craft information extraction templates. Rodney D. Nielsen, Wayne H. Ward, James H. Martin |
Nat. Lang. Eng. | 3 |
| 2008 | Pedagogically Useful Extractive Summaries for Science Education
Sebastian de la Chica, Faisal Ahmad, James H. Martin, Tamara Sumner |
COLING | 3 |
| 2008 | Automatic Generation of Fine-Grained Representations of Learner Response Semantics
Rodney D. Nielsen, Wayne H. Ward, James H. Martin |
Intelligent Tutoring Systems | 3 |
| 2008 | Building a Corpus of Temporal-Causal Structure
Steven Bethard, William J. Corvey, Sara Klingenstein, James H. Martin |
LREC | 4 |
| 2008 | Annotating Students' Understanding of Science Concepts
Rodney D. Nielsen, Wayne H. Ward, James H. Martin, Martha Palmer |
LREC | 3 |
| 2008 | Semantic role labeling for protein transport predicatesabstractBACKGROUND: Automatic semantic role labeling (SRL) is a natural language processing (NLP) technique that maps sentences to semantic representations. This technique has been widely studied in the recent years, but mostly with data in newswire domains. Here, we report on a SRL model for identifying the semantic roles of biomedical predicates describing protein transport in GeneRIFs - manually curated sentences focusing on gene functions. To avoid the computational cost of syntactic parsing, and because the boundaries of our protein transport roles often did not match up with syntactic phrase boundaries, we approached this problem with a word-chunking paradigm and trained support vector machine classifiers to classify words as being at the beginning, inside or outside of a protein transport role. RESULTS: We collected a set of 837 GeneRIFs describing movements of proteins between cellular components, whose predicates were annotated for the semantic roles AGENT, PATIENT, ORIGIN and DESTINATION. We trained these models with the features of previous word-chunking models, features adapted from phrase-chunking models, and features derived from an analysis of our data. Our models were able to label protein transport semantic roles with 87.6% precision and 79.0% recall when using manually annotated protein boundaries, and 87.0% precision and 74.5% recall when using automatically identified ones. CONCLUSION: We successfully adapted the word-chunking classification paradigm to semantic role labeling, applying it to a new domain with predicates completely absent from any previous studies. By combining the traditional word and phrasal role labeling features with biomedical features like protein boundaries and MEDPOST part of speech tags, we were able to address the challenges posed by the new domain data and subsequently build robust models that achieved F-measures as high as 83.1. This system for extracting protein transport information from GeneRIFs performs well even with proteins identified automatically, and is therefore more robust than the rule-based methods previously used to extract protein transport roles. Steven Bethard, Zhiyong Lu, James H. Martin, Lawrence Hunter |
BMC Bioinform. | 3 |
| 2008 | Towards Robust Semantic Role LabelingabstractMost semantic role labeling (SRL) research has been focused on training and evaluating on the same corpus. This strategy, although appropriate for initiating research, can lead to overtraining to the particular corpus. This article describes the operation of assert, a state-of-the art SRL system, and analyzes the robustness of the system when trained on one genre of data and used to label a different genre. As a starting point, results are first presented for training and testing the system on the PropBank corpus, which is annotated Wall Street Journal (WSJ) data. Experiments are then presented to evaluate the portability of the system to another source of data. These experiments are based on comparisons of performance using PropBanked WSJ data and PropBanked Brown Corpus data. The results indicate that whereas syntactic parses and argument identification transfer relatively well to a new corpus, argument classification does not. An analysis of the reasons for this is presented and these generally point to the nature of the more lexical/semantic features dominating the classification task where more general structural features are dominant in the argument identification task. Sameer Pradhan, Wayne H. Ward, James H. Martin |
Comput. Linguistics | 3 |
| 2007 | Towards Robust Unsupervised Personal Name Disambiguation
James H. Martin |
EMNLP-CoNLL | 2 |
| 2007 | Towards Robust Semantic Role Labeling
Sameer Pradhan, Wayne H. Ward, James H. Martin |
HLT-NAACL | 3 |
| 2006 | Identification of Event Mentions and their Semantic Class
Steven Bethard, James H. Martin |
EMNLP | 2 |
| 2005 | Semantic Role Labeling Using Different Syntactic ViewsabstractSemantic role labeling is the process of annotating the predicate-argument structure in text with semantic labels. In this paper we present a state-of-the-art baseline semantic role labeling system based on Support Vector Machine classifiers. We show improvements on this system by: i) adding new features including features extracted from dependency parses, ii) performing feature selection and calibration and iii) combining parses obtained from semantic parsers trained using different syntactic views. Error analysis of the baseline system showed that approximately half of the argument identification errors resulted from parse errors in which there was no syntactic constituent that aligned with the correct argument. In order to address this problem, we combined semantic parses from a Minipar syntactic parse and from a chunked syntactic representation with our original baseline system which was based on Charniak parses. All of the reported techniques resulted in performance improvements. Sameer Pradhan, Wayne H. Ward, Kadri Hacioglu, James H. Martin, Daniel Jurafsky |
ACL | 4 |
| 2005 | Semantic Role Chunking Combining Complementary Syntactic Views
Sameer Pradhan, Kadri Hacioglu, Wayne H. Ward, James H. Martin, Daniel Jurafsky |
CoNLL | 4 |
| 2005 | Support Vector Learning for Semantic Argument Classification
Sameer Pradhan, Kadri Hacioglu, Valerie Krugler, Wayne H. Ward, James H. Martin, Daniel Jurafsky |
Mach. Learn. | 5 |
| 2004 | Semantic Role Labeling by Tagging Syntactic Chunks
Kadri Hacioglu, Sameer Pradhan, Wayne H. Ward, James H. Martin, Daniel Jurafsky |
CoNLL | 4 |
| 2004 | Shallow Semantic Parsing using Support Vector Machines
Sameer Pradhan, Wayne H. Ward, Kadri Hacioglu, James H. Martin, Daniel Jurafsky |
HLT-NAACL | 4 |
| 2003 | Semantic Role Parsing: Adding Semantic Structure to Unstructured TextabstractThere is an ever-growing need to add structure in the form of semantic markup to the huge amounts of unstructured text data now available. We present the technique of shallow semantic parsing, the process of assigning a simple WHO did WHAT to WHOM, etc., structure to sentences in text, as a useful tool in achieving this goal. We formulate the semantic parsing problem as a classification problem using support vector machines. Using a hand-labeled training set and a set of features drawn from earlier work together with some feature enhancements, we demonstrate a system that performs better than all other published results on shallow semantic parsing. Sameer Pradhan, Kadri Hacioglu, Wayne H. Ward, James H. Martin, Daniel Jurafsky |
ICDM | 4 |
| 1997 | Evidence-Based Static Branch Prediction Using Machine LearningabstractCorrectly predicting the direction that branches will take is increasingly important in today's wide-issue computer architectures. The name program-based branch prediction is given to static branch prediction techniques that base their prediction on a program's structure. In this article, we investigate a new approach to program-based branch prediction that uses a body of existing programs to predict the branch behavior in a new program. We call this approach to program-based branch prediction evidence-based static prediction , or ESP. The main idea of ESP is that the behavior of a corpus of programs can be used to infer the behavior of new programs. In this article, we use neural networks and decision trees to map static features associated with each branch to a prediction that the branch will be taken. ESP shows significant advantages over other prediction mechanisms. Specifically, it is a program-based technique; it is effective across a range of programming languages and programming styles; and it does not rely on the use of expert-defined heuristics. In this article, we describe the application of ESP to the problem of static branch prediction and compare our results to existing program-based branch predictors. We also investigate the applicability of ESP across computer architectures, programming languages, compilers, and run-time systems. We provide results showing how sensitive ESP is to the number and type of static features and programs included in the ESP training sets, and we compare the efficacy of static branch prediction for subroutine libraries. Averaging over a body of 43 C and Fortran programs, ESP branch prediction results in a miss rate of 20%, as compared with the 25% miss rate obtained using the best existing program-based heuristics. Brad Calder, Dirk Grunwald, Michael P. Jones, Donald C. Lindsay, James H. Martin, Michael C. Mozer, Benjamin G. Zorn |
ACM Trans. Program. Lang. Syst. | 5 |
| 1995 | Corpus-Based Static Branch PredictionabstractCorrectly predicting the direction that branches will take is increasingly important in today's wide-issue computer architectures. The name program-based branch prediction is given to static branch prediction techniques that base their prediction on a program's structure. In this paper, we investigate a new approach to program-based branch prediction that uses a body of existing programs to predict the branch behavior in a new program. We call this approach to program-based branch prediction, evidence-based static prediction, or ESP. The main idea of ESP is that the behavior of a corpus of programs can be used to infer the behavior of new programs. In this paper, we use a neural network to map static features associated with each branch to the probability that the branch will be taken. ESP shows significant advantages over other prediction mechanisms. Specifically, it is a program-based technique, it is effective across a range of programming languages and programming styles, and it does not rely on the use of expert-defined heuristics. Brad Calder, Dirk Grunwald, Donald C. Lindsay, James H. Martin, Michael C. Mozer, Benjamin G. Zorn |
PLDI | 4 |
| 1995 | Expressing Rhetorical Relations in Instructional Text: A Case Study of the Purposes Relation
Keith Vander Linden, James H. Martin |
Comput. Linguistics | 2 |
| 1994 | MetaBank: A Knowledge-Base of Metaphoric Language ConventionsabstractThe frequent and conventional use of nonliteral language has been a major stumbling block for natural language processing systems since the early machine translation efforts. Metaphor, metonymy, and indirect speech acts are among the most troublesome phenomena. Recent computational efforts addressing these problems have taken an approach that emphasizes the use of systematic knowledge about nonliteral language conventions. We are currently engaged in an effort to supply this knowledge in the case of conventional metaphor. We are constructing MetaBank: an empirically derived and theoretically motivated knowledge‐base of English metaphorical conventions. This article describes our three‐part approach to the construction of MetaBank: the collection of on‐line textual resources and databases of linguistic generalizations, the development of a methodology for analyzing these resources, and the construction of a knowledge‐base based on the preceding analyses. James H. Martin |
Comput. Intell. | 1 |
| 1993 | Norvig's Paradigms of Artificial Intelligence Programming: An Instructor's Perspective
James H. Martin |
Artif. Intell. | 1 |
| 1992 | Introduction to the Special Issue on Non-Literal Language
Dan Fass, James H. Martin, Elizabeth A. Hinkelman |
Comput. Intell. | 2 |
| 1988 | Representing regularities in the metaphoric lexicon
James H. Martin |
COLING | 1 |
| 1988 | The Berkeley UNIX Consultant Project
Robert Wilensky, David N. Chin, Marc Luria, James H. Martin, James Mayfield, Dekai Wu |
Comput. Linguistics | 4 |
| 1987 | Understanding New Metaphors
James H. Martin |
IJCAI | 1 |