Steven Bethard

dblp:52/5246 · also Steven J. Bethard · DBLP profile ↗
← Back
56ranked-venue papers
8as first author
10since 2021 · last 2025
0000-0001-9560-6491ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 39 · 7 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
16 papers
Information extraction and text analysis · 44% Transfer learning and domain adaptation · 19% Language models and text generation · 14%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 96% Machine learning and data management · 4%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 30 heaviest of 35, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
temporal information extraction
2.042025
A Semantic Parsing Framework for End-to-End Time Normalization · NeurIPS 2025
Transformer-Based Temporal Information Extraction and Application: A Review · EMNLP 2025
A Synchronous Context Free Grammar for Time Normalization · EMNLP 2013
Natural language and speech › Language models and text generation › language modeling › language model architecture
transformer language model
0.912025
Transformer-Based Temporal Information Extraction and Application: A Review · EMNLP 2025
Compilers and program optimization
code generation
0.912025
A Semantic Parsing Framework for End-to-End Time Normalization · NeurIPS 2025
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.822022
A Comparison of Strategies for Source-Free Domain Adaptation · ACL (1) 2022
Domain Adaptation for Authorship Attribution: Improved Structural Correspondence Learning · ACL (1) 2016
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering
0.822020
Unsupervised Alignment-based Iterative Evidence Retrieval for Multi-hop Question Answering · ACL 2020
Quick and (not so) Dirty: Unsupervised Selection of Justification Sentences for Multi-hop Question Answering · EMNLP/IJCNLP (1) 2019
Machine learning › Learning theory › statistical learning theory
bias-variance tradeoff
0.712023
Two-Stage Fine-Tuning for Improved Bias and Variance for Large Pretrained Language Models · ACL (1) 2023
Machine learning › Transfer learning and domain adaptation
fine-tuning
0.712023
Two-Stage Fine-Tuning for Improved Bias and Variance for Large Pretrained Language Models · ACL (1) 2023
Machine learning › Transfer learning and domain adaptation › domain adaptation
source-free domain adaptation
0.612022
A Comparison of Strategies for Source-Free Domain Adaptation · ACL (1) 2022
Natural language and speech › Information extraction and text analysis › fact-checking
evidence retrieval
0.412020
Unsupervised Alignment-based Iterative Evidence Retrieval for Multi-hop Question Answering · ACL 2020
Natural language and speech › Information extraction and text analysis › text normalization
medical term normalization
0.412020
A Generate-and-Rank Framework with Semantic Type Regularization for Biomedical Concept Normalization · ACL 2020
Natural language and speech › Language models and text generation › natural language understanding › question answering
neural question answering
0.412020
Having Your Cake and Eating it Too: Training Neural Retrieval for Language Inference without Losing Lexical Match · SIGIR 2020
Natural language and speech › Question answering and dialogue systems
open-domain question answering
0.412020
Having Your Cake and Eating it Too: Training Neural Retrieval for Language Inference without Losing Lexical Match · SIGIR 2020
Machine learning › Trustworthy machine learning › interpretability › neural network interpretation
transformer interpretability
0.412020
How does BERT's attention change when you fine-tune? An analysis methodology and a case study in negation scope · ACL 2020
Information retrieval › retrieval models › probabilistic retrieval model
BM25
0.412020
Having Your Cake and Eating it Too: Training Neural Retrieval for Language Inference without Losing Lexical Match · SIGIR 2020
Information retrieval › similarity measure
lexical matching
0.412020
Having Your Cake and Eating it Too: Training Neural Retrieval for Language Inference without Losing Lexical Match · SIGIR 2020
Information retrieval › retrieval models
neural retrieval
0.412020
Having Your Cake and Eating it Too: Training Neural Retrieval for Language Inference without Losing Lexical Match · SIGIR 2020
Information retrieval
retrieval models
0.412020
Having Your Cake and Eating it Too: Training Neural Retrieval for Language Inference without Losing Lexical Match · SIGIR 2020
Medical and health informatics › mental health informatics
depression detection
0.312018
Measuring the Latency of Depression Detection in Social Media · WSDM 2018
Natural language and speech › Language models and text generation
large language model
0.312025
A Semantic Parsing Framework for End-to-End Time Normalization · NeurIPS 2025
Natural language and speech › Information extraction and text analysis › text mining › authorship analysis
authorship attribution
0.212016
Domain Adaptation for Authorship Attribution: Improved Structural Correspondence Learning · ACL (1) 2016
Natural language and speech › Information extraction and text analysis
coreference resolution
0.212015
Adapting Coreference Resolution for Narrative Processing · EMNLP 2015
Natural language and speech › Information extraction and text analysis
semantic role labeling
0.212015
Domain Adaptation in Semantic Role Labeling Using a Neural Language Model and Linguistic Resources · IEEE ACM Trans. Audio Speech Lang. Process. 2015
Machine learning › Kernel, tree and ensemble methods
ensemble learning
0.212023
Two-Stage Fine-Tuning for Improved Bias and Variance for Large Pretrained Language Models · ACL (1) 2023
Usability and user experience research
usability evaluation
0.212014
Easy does it: more usable CAPTCHAs · CHI 2014
Authentication and access control › human interactive proofs
CAPTCHA
0.212014
Easy does it: more usable CAPTCHAs · CHI 2014
Natural language and speech › Information extraction and text analysis › syntactic parsing › dependency parsing
temporal dependency parsing
0.112012
Extracting Narrative Timelines as Temporal Dependency Structures · ACL (1) 2012
Machine learning › Transfer learning and domain adaptation › domain alignment
unsupervised alignment
0.112020
Unsupervised Alignment-based Iterative Evidence Retrieval for Multi-hop Question Answering · ACL 2020
Usable security › security tool usability
CAPTCHA usability
0.112010
How Good Are Humans at Solving CAPTCHAs? A Large Scale Evaluation · IEEE Symposium on Security and Privacy 2010
Machine learning and data management
transfer learning
0.112016
Domain Adaptation for Authorship Attribution: Improved Structural Correspondence Learning · ACL (1) 2016
Knowledge, reasoning and agents › Knowledge representation and reasoning
narrative processing
0.112015
Adapting Coreference Resolution for Narrative Processing · EMNLP 2015

Methods — techniques the papers use, named apart from their topics

data augmentation · 2.3large language model · 1.7RoBERTa · 1.3literature review · 0.9variance decomposition · 0.7ensemble methods · 0.7self-training · 0.6active learning · 0.6transformer · 0.4list-wise ranking · 0.4embedding-based alignment · 0.4BM25 · 0.4BERT · 0.4survey · 0.4production deployment · 0.4crowdsourced evaluation · 0.4latency-weighted f1 · 0.3ERDE metric analysis · 0.3
YearPublicationVenuePosition
2025 Transformer-Based Temporal Information Extraction and Application: A Review
abstract
Temporal information extraction (IE) aims to extract structured temporal information from unstructured text, thereby uncovering the implicit timelines within.This technique is applied across domains such as healthcare, newswire, and intelligence analysis, aiding models in these areas to perform temporal reasoning and enabling human users to grasp the temporal structure of text.Transformer-based pre-trained language models have produced revolutionary advancements in natural language processing, demonstrating exceptional performance across a multitude of tasks.Despite the achievements garnered by Transformer-based approaches in temporal IE, there is a lack of comprehensive reviews on these endeavors.In this paper, we aim to bridge this gap by systematically summarizing and analyzing the body of work on temporal IE using Transformers while highlighting potential future research directions.
Xin Su 0008, Phillip Howard, Steven Bethard
EMNLP3
2025 A Semantic Parsing Framework for End-to-End Time Normalization
abstract
Time normalization is the task of converting natural language temporal expressions into machine-readable representations. It underpins many downstream applications in information retrieval, question answering, and clinical decision-making. Traditional systems based on the ISO-TimeML schema limit expressivity and struggle with complex constructs such as compositional, event-relative, and multi-span time expressions. In this work, we introduce a novel formulation of time normalization as a code generation task grounded in the SCATE framework, which defines temporal semantics through symbolic and compositional operators. We implement a fully executable SCATE Python library and demonstrate that large language models (LLMs) can generate executable SCATE code. Leveraging this capability, we develop an automatic data augmentation pipeline using LLMs to synthesize large-scale annotated data with code-level validation. Our experiments show that small, locally deployable models trained on this augmented data can achieve strong performance, outperforming even their LLM parents and enabling practical, accurate, and interpretable time normalization.
Xin Su 0008, Sungduk Yu, Phillip Howard, Steven Bethard
NeurIPS4
2025 Identifying task groupings for multi-task learning using pointwise V-usable information
Yingya Li, Timothy A. Miller, Steven Bethard, Guergana K. Savova
J. Biomed. Informatics3
2024 Semi-Structured Chain-of-Thought: Integrating Multiple Sources of Knowledge for Improved Language Model Reasoning
abstract
Xin Su, Tiep Le, Steven Bethard, Phillip Howard. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Xin Su 0008, Tiep Le, Steven Bethard, Phillip Howard
NAACL-HLT3
2023 Two-Stage Fine-Tuning for Improved Bias and Variance for Large Pretrained Language Models
abstract
The bias-variance tradeoff is the idea that learning methods need to balance model complexity with data size to minimize both under-fitting and over-fitting.Recent empirical work and theoretical analyses with over-parameterized neural networks challenge the classic bias-variance trade-off notion suggesting that no such trade-off holds: as the width of the network grows, bias monotonically decreases while variance initially increases followed by a decrease.In this work, we first provide a variance decomposition-based justification criteria to examine whether large pretrained neural models in a fine-tuning setting are generalizable enough to have low bias and variance.We then perform theoretical and empirical analysis using ensemble methods explicitly designed to decrease variance due to optimization.This results in essentially a two-stage fine-tuning algorithm that first ratchets down bias and variance iteratively, and then uses a selected fixed-bias model to further reduce variance due to optimization by ensembling.We also analyze the nature of variance change with the ensemble size in low-and high-resource classes.Empirical results show that this two-stage method obtains strong results on SuperGLUE tasks and clinical information extraction tasks.Code and settings are available: https://github.com/christa60/ bias-var-fine-tuning-plms.git
Lijing Wang 0001, Yingya Li, Timothy A. Miller, Steven Bethard, Guergana K. Savova
ACL (1)4
2023 Addressing structural hurdles for metadata extraction from environmental impact statements
abstract
Abstract Natural language processing techniques can be used to analyze the linguistic content of a document to extract missing pieces of metadata. However, accurate metadata extraction may not depend solely on the linguistics, but also on structural problems such as extremely large documents, unordered multi‐file documents, and inconsistency in manually labeled metadata. In this work, we start from two standard machine learning solutions to extract pieces of metadata from Environmental Impact Statements, environmental policy documents that are regularly produced under the US National Environmental Policy Act of 1969. We present a series of experiments where we evaluate how these standard approaches are affected by different issues derived from real‐world data. We find that metadata extraction can be strongly influenced by nonlinguistic factors such as document length and volume ordering and that the standard machine learning solutions often do not scale well to long documents. We demonstrate how such solutions can be better adapted to these scenarios, and conclude with suggestions for other NLP practitioners cataloging large document collections.
Egoitz Laparra, Alex Binford-Walsh, Kirk Emerson, Marc L. Miller, Laura López-Hoffman, Faiz Currim, Steven Bethard
J. Assoc. Inf. Sci. Technol.7
2022 A Comparison of Strategies for Source-Free Domain Adaptation
abstract
Data sharing restrictions are common in NLP, especially in the clinical domain, but there is limited research on adapting models to new domains without access to the original training data, a setting known as source-free domain adaptation.We take algorithms that traditionally assume access to the source-domain training data-active learning, self-training, and data augmentation-and adapt them for source-free domain adaptation.Then we systematically compare these different strategies across multiple tasks and domains.We find that active learning yields consistent gains across all SemEval 2021 Task 10 tasks and domains, but though the shared task saw successful self-trained and data augmented models, our systematic comparison finds these strategies to be unreliable for source-free domain adaptation.
Xin Su 0008, Yiyun Zhao, Steven Bethard
ACL (1)3
2021 Do pretrained transformers infer telicity like humans?
abstract
Pretrained transformer-based language models achieve state-of-the-art performance in many NLP tasks, but it is an open question whether the knowledge acquired by the models during pretraining resembles the linguistic knowledge of humans.We present both humans and pretrained transformers with descriptions of events, and measure their preference for telic interpretations (the event has a natural endpoint) or atelic interpretations (the event does not have a natural endpoint).To measure these preferences and determine what factors influence them, we design an English test and a novel-word test that include a variety of linguistic cues (noun phrase quantity, resultative structure, contextual information, temporal units) that bias toward certain interpretations.We find that humans' choice of telicity interpretation is reliably influenced by theoretically-motivated cues, transformer models (BERT and RoBERTa) are influenced by some (though not all) of the cues, and transformer models often rely more heavily on temporal units than humans do.
Yiyun Zhao, Jian Gang Ngui, Lucy Hall Hartley, Steven Bethard
CoNLL4
2021 Explainable Multi-hop Verbal Reasoning Through Internal Monologue
abstract
Many state-of-the-art (SOTA) language models have achieved high accuracy on several multi-hop reasoning problems.However, these approaches tend to not be interpretable because they do not make the intermediate reasoning steps explicit.Moreover, models trained on simpler tasks tend to fail when directly tested on more complex problems.We propose the Explainable multi-hop Verbal Reasoner (EVR) to solve these limitations by (a) decomposing multi-hop reasoning problems into several simple ones, and (b) using natural language to guide the intermediate reasoning hops.We implement EVR by extending the classic reasoning paradigm General Problem Solver (GPS) with a SOTA generative language model to generate subgoals and perform inference in natural language at each reasoning step.Evaluation of EVR on Clark et al. (2020)'s synthetic question answering (QA) dataset shows that EVR achieves SOTA performance while being able to generate all reasoning steps in natural language.Furthermore, EVR generalizes better than other strong methods when trained on simpler tasks or less training data (up to 35.7% and 7.7% absolute improvement respectively). 1
Zhengzhong Liang, Steven Bethard, Mihai Surdeanu
NAACL-HLT2
2021 If You Want to Go Far Go Together: Unsupervised Joint Candidate Evidence Retrieval for Multi-hop Question Answering
abstract
Multi-hop reasoning requires aggregation and inference from multiple facts.To retrieve such facts, we propose a simple approach that retrieves and reranks set of evidence facts jointly.Our approach first generates unsupervised clusters of sentences as candidate evidence by accounting links between sentences and coverage with the given query.Then, a RoBERTa-based reranker is trained to bring the most representative evidence cluster to the top.We specifically emphasize on the importance of retrieving evidence jointly by showing several comparative analyses to other methods that retrieve and rerank evidence sentences individually.First, we introduce several attention-and embedding-based analyses, which indicate that jointly retrieving and reranking approaches can learn compositional knowledge required for multi-hop reasoning.Second, our experiments show that jointly retrieving candidate evidence leads to substantially higher evidence retrieval performance when fed to the same supervised reranker.In particular, our joint retrieval and then reranking approach achieves new state-of-the-art evidence retrieval performance on two multi-hop question answering (QA) datasets: 30.5 Recall@2 on QASC, and 67.6% F1 on MultiRC.When the evidence text from our joint retrieval approach is fed to a RoBERTa-based answer selection classifier, we achieve new state-ofthe-art QA performance on MultiRC and second best result on QASC.
Vikas Yadav, Steven Bethard, Mihai Surdeanu
NAACL-HLT2
2020 A Generate-and-Rank Framework with Semantic Type Regularization for Biomedical Concept Normalization
abstract
Concept normalization, the task of linking textual mentions of concepts to concepts in an ontology, is challenging because ontologies are large.In most cases, annotated datasets cover only a small sample of the concepts, yet concept normalizers are expected to predict all concepts in the ontology.In this paper, we propose an architecture consisting of a candidate generator and a list-wise ranker based on BERT.The ranker considers pairings of concept mentions and candidate concepts, allowing it to make predictions for any concept, not just those seen during training.We further enhance this list-wise approach with a semantic type regularizer that allows the model to incorporate semantic type information from the ontology during training.Our proposed concept normalization framework achieves stateof-the-art performance on multiple datasets.
Dongfang Xu, Zeyu Zhang 0002, Steven Bethard
ACL3
2020 Unsupervised Alignment-based Iterative Evidence Retrieval for Multi-hop Question Answering
abstract
Evidence retrieval is a critical stage of question answering (QA), necessary not only to improve performance, but also to explain the decisions of the corresponding QA method.We introduce a simple, fast, and unsupervised iterative evidence retrieval method, which relies on three ideas: (a) an unsupervised alignment approach to soft-align questions and answers with justification sentences using only GloVe embeddings, (b) an iterative process that reformulates queries focusing on terms that are not covered by existing justifications, which (c) a stopping criterion that terminates retrieval when the terms in the given question and candidate answers are covered by the retrieved justifications.Despite its simplicity, our approach outperforms all the previous methods (including supervised methods) on the evidence selection task on two datasets: MultiRC and QASC.When these evidence sentences are fed into a RoBERTa answer classification component, we achieve state-of-the-art QA performance on these two datasets.
Vikas Yadav, Steven Bethard, Mihai Surdeanu
ACL2
2020 How does BERT's attention change when you fine-tune? An analysis methodology and a case study in negation scope
abstract
Large pretrained language models like BERT, after fine-tuning to a downstream task, have achieved high performance on a variety of NLP problems.Yet explaining their decisions is difficult despite recent work probing their internal representations.We propose a procedure and analysis methods that take a hypothesis of how a transformer-based model might encode a linguistic phenomenon, and test the validity of that hypothesis based on a comparison between knowledge-related downstream tasks with downstream control tasks, and measurement of cross-dataset consistency.We apply this methodology to test BERT and RoBERTa on a hypothesis that some attention heads will consistently attend from a word in negation scope to the negation cue.We find that after fine-tuning BERT and RoBERTa on a negation scope task, the average attention head improves its sensitivity to negation and its attention consistency across negation datasets compared to the pre-trained models.However, only the base models (not the large models) improve compared to a control task, indicating there is evidence for a shallow encoding of negation only in the base models.
Yiyun Zhao, Steven Bethard
ACL2
2020 A Dataset and Evaluation Framework for Complex Geographical Description Parsing
abstract
Much previous work on geoparsing has focused on identifying and resolving individual toponyms in text like Adrano, S.Maria di Licodia or Catania.However, geographical locations occur not only as individual toponyms, but also as compositions of reference geolocations joined and modified by connectives, e.g., ". . .between the towns of Adrano and S.Maria di Licodia, 32 kilometres northwest of Catania".Ideally, a geoparser should be able to take such text, and the geographical shapes of the toponyms referenced within it, and parse these into a geographical shape, formed by a set of coordinates, that represents the location described.But creating a dataset for this complex geoparsing task is difficult and, if done manually, would require a huge amount of effort to annotate the geographical shapes of not only the geolocation described but also the reference toponyms.We present an approach that automates most of the process by combining Wikipedia and OpenStreetMap.As a result, we have gathered a collection of 360,187 uncurated complex geolocation descriptions, from which we have manually curated 1,000 examples intended to be used as a test set.To accompany the data, we define a new geoparsing evaluation framework along with a scoring methodology and a set of baselines.
Egoitz Laparra, Steven Bethard
COLING2
2020 Having Your Cake and Eating it Too: Training Neural Retrieval for Language Inference without Losing Lexical Match
abstract
We present a study on the importance of information retrieval (IR) techniques for both the interpretability and the performance of neural question answering (QA) methods. We show that the current state-of-the-art transformer methods (like RoBERTa) encode poorly simple information retrieval (IR) concepts such as lexical overlap between query and the document. To mitigate this limitation, we introduce a supervised RoBERTa QA method that is trained to mimic the behavior of BM25 and the soft-matching idea behind embedding-based alignment methods. We show that fusing the simple lexical-matching IR concepts in transformer techniques results in improvement a) of their (lexical-matching) interpretability, b) retrieval performance, and c) the QA performance on two multi-hop QA datasets. We further highlight the lexical-chasm gap bridging capabilities of transformer methods by analyzing the attention distributions of the supervised RoBERTa classifier over the context versus lexically-matched token pairs.
Vikas Yadav, Steven Bethard, Mihai Surdeanu
SIGIR2
2020 Does BERT need domain adaptation for clinical negation detection?
abstract
INTRODUCTION: Classifying whether concepts in an unstructured clinical text are negated is an important unsolved task. New domain adaptation and transfer learning methods can potentially address this issue. OBJECTIVE: We examine neural unsupervised domain adaptation methods, introducing a novel combination of domain adaptation with transformer-based transfer learning methods to improve negation detection. We also want to better understand the interaction between the widely used bidirectional encoder representations from transformers (BERT) system and domain adaptation methods. MATERIALS AND METHODS: We use 4 clinical text datasets that are annotated with negation status. We evaluate a neural unsupervised domain adaptation algorithm and BERT, a transformer-based model that is pretrained on massive general text datasets. We develop an extension to BERT that uses domain adversarial training, a neural domain adaptation method that adds an objective to the negation task, that the classifier should not be able to distinguish between instances from 2 different domains. RESULTS: The domain adaptation methods we describe show positive results, but, on average, the best performance is obtained by plain BERT (without the extension). We provide evidence that the gains from BERT are likely not additive with the gains from domain adaptation. DISCUSSION: Our results suggest that, at least for the task of clinical negation detection, BERT subsumes domain adaptation, implying that BERT is already learning very general representations of negation phenomena such that fine-tuning even on a specific corpus does not lead to much overfitting. CONCLUSION: Despite being trained on nonclinical text, the large training sets of models like BERT lead to large gains in performance for the clinical negation detection task.
Chen Lin 0002, Steven Bethard, Dmitriy Dligach, Farig Sadeque, Guergana K. Savova, Timothy A. Miller
J. Am. Medical Informatics Assoc.2
2020 Unified Medical Language System resources improve sieve-based generation and Bidirectional Encoder Representations from Transformers (BERT)-based ranking for concept normalization
abstract
OBJECTIVE: Concept normalization, the task of linking phrases in text to concepts in an ontology, is useful for many downstream tasks including relation extraction, information retrieval, etc. We present a generate-and-rank concept normalization system based on our participation in the 2019 National NLP Clinical Challenges Shared Task Track 3 Concept Normalization. MATERIALS AND METHODS: The shared task provided 13 609 concept mentions drawn from 100 discharge summaries. We first design a sieve-based system that uses Lucene indices over the training data, Unified Medical Language System (UMLS) preferred terms, and UMLS synonyms to generate a list of possible concepts for each mention. We then design a listwise classifier based on the BERT (Bidirectional Encoder Representations from Transformers) neural network to rank the candidate concepts, integrating UMLS semantic types through a regularizer. RESULTS: Our generate-and-rank system was third of 33 in the competition, outperforming the candidate generator alone (81.66% vs 79.44%) and the previous state of the art (76.35%). During postevaluation, the model's accuracy was increased to 83.56% via improvements to how training data are generated from UMLS and incorporation of our UMLS semantic type regularizer. DISCUSSION: Analysis of the model shows that prioritizing UMLS preferred terms yields better performance, that the UMLS semantic type regularizer results in qualitatively better concept predictions, and that the model performs well even on concepts not seen during training. CONCLUSIONS: Our generate-and-rank framework for UMLS concept normalization integrates key UMLS features like preferred terms and semantic types with a neural network-based ranking model to accurately link phrases in text to UMLS concepts.
Dongfang Xu, Manoj Gopale, Kris Brown, Edmon Begoli, Steven Bethard
J. Am. Medical Informatics Assoc.6
2019 Quick and (not so) Dirty: Unsupervised Selection of Justification Sentences for Multi-hop Question Answering
abstract
Vikas Yadav, Steven Bethard, Mihai Surdeanu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Vikas Yadav, Steven Bethard, Mihai Surdeanu
EMNLP/IJCNLP (1)2
2018 A Survey on Recent Advances in Named Entity Recognition from Deep Learning models
abstract
Named Entity Recognition (NER) is a key component in NLP systems for question answering, information retrieval, relation extraction, etc. NER systems have been studied and developed widely for decades, but accurate systems using deep neural networks (NN) have only been introduced in the last few years. We present a comprehensive survey of deep neural network architectures for NER, and contrast them with previous approaches to NER based on feature engineering and other supervised or semi-supervised learning algorithms. Our results highlight the improvements achieved by neural networks, and show how incorporating some of the lessons learned from past work on feature-based NER systems can yield further improvements.
Vikas Yadav, Steven Bethard
COLING2
2018 Measuring the Latency of Depression Detection in Social Media
abstract
Detecting depression is a key public health challenge, as almost 12% of all disabilities can be attributed to depression. Computational models for depression detection must prove not only that can they detect depression, but that they can do it early enough for an intervention to be plausible. However, current evaluations of depression detection are poor at measuring model latency. We identify several issues with the currently popular ERDE metric, and propose a latency-weighted F1 metric that addresses these concerns. We then apply this evaluation to several models from the recent eRisk 2017 shared task on depression detection, and show how our proposed measure can better capture system differences.
Farig Sadeque, Dongfang Xu, Steven Bethard
WSDM3
2018 From Characters to Time Intervals: New Paradigms for Evaluation and Neural Parsing of Time Normalizations
abstract
This paper presents the first model for time normalization trained on the SCATE corpus. In the SCATE schema, time expressions are annotated as a semantic composition of time entities. This novel schema favors machine learning approaches, as it can be viewed as a semantic parsing task. In this work, we propose a character level multi-output neural network that outperforms previous state-of-the-art built on the TimeML schema. To compare predictions of systems that follow both SCATE and TimeML, we present a new scoring metric for time intervals. We also apply this new metric to carry out a comparative analysis of the annotations of both schemes in the same corpus.
Egoitz Laparra, Dongfang Xu, Steven Bethard
Trans. Assoc. Comput. Linguistics3
2017 Recurrent Neural Network Architectures for Event Extraction from Italian Medical Reports
Natalia Viani, Timothy A. Miller, Dmitriy Dligach, Steven Bethard, Carlo Napolitano, Silvia G. Priori, Riccardo Bellazzi, Lucia Sacchi, Guergana K. Savova
AIME4
2017 Improving Implicit Semantic Role Labeling by Predicting Semantic Frame Arguments
abstract
Implicit semantic role labeling (iSRL) is the task of predicting the semantic roles of a predicate that do not appear as explicit arguments, but rather regard common sense knowledge or are mentioned earlier in the discourse. We introduce an approach to iSRL based on a predictive recurrent neural semantic frame model (PRNSFM) that uses a large unannotated corpus to learn the probability of a sequence of semantic arguments given a predicate. We leverage the sequence probabilities predicted by the PRNSFM to estimate selectional preferences for predicates and their arguments. On the NomBank iSRL test set, our approach improves state-of-the-art performance on implicit semantic role labeling with less reliance than prior work on manually constructed language resources.
Quynh Ngoc Thi Do, Steven Bethard, Marie-Francine Moens
IJCNLP(1)2
2017 Towards generalizable entity-centric clinical coreference resolution
Timothy A. Miller, Dmitriy Dligach, Steven Bethard, Chen Lin 0002, Guergana K. Savova
J. Biomed. Informatics3
2016 Domain Adaptation for Authorship Attribution: Improved Structural Correspondence Learning
Upendra Sapkota, Thamar Solorio, Manuel Montes-y-Gómez, Steven Bethard
ACL (1)4
2016 Feature Portability in Cross-domain Clinical Coreference
Timothy A. Miller, Dmitriy Dligach, Chen Lin 0002, Steven Bethard, Guergana K. Savova
AMIA4
2016 Facing the most difficult case of Semantic Role Labeling: A collaboration of word embeddings and co-training
abstract
We present a successful collaboration of word embeddings and co-training to tackle in the most difficult test case of semantic role labeling: predicting out-of-domain and unseen semantic frames. Despite the fact that co-training is a successful traditional semi-supervised method, its application in SRL is very limited especially when a huge amount of labeled data is available. In this work, co-training is used together with word embeddings to improve the performance of a system trained on a large training dataset. We also introduce a semantic role labeling system with a simple learning architecture and effective inference that is easily adaptable to semi-supervised settings with new training data and/or new features. On the out-of-domain testing set of the standard benchmark CoNLL 2009 data our simple approach achieves high performance and improves state-of-the-art results.
Quynh Ngoc Thi Do, Steven Bethard, Marie-Francine Moens
COLING2
2016 A Semantically Compositional Annotation Scheme for Time Normalization
Steven Bethard, Jonathan Parker
LREC1
2016 Age and Gender Prediction on Health Forum Data
Prasha Shrestha, Nicolas Rey-Villamizar, Farig Sadeque, Ted Pedersen, Steven Bethard, Thamar Solorio
LREC5
2016 Multilayered temporal modeling for the clinical domain
abstract
OBJECTIVE: To develop an open-source temporal relation discovery system for the clinical domain. The system is capable of automatically inferring temporal relations between events and time expressions using a multilayered modeling strategy. It can operate at different levels of granularity--from rough temporality expressed as event relations to the document creation time (DCT) to temporal containment to fine-grained classic Allen-style relations. MATERIALS AND METHODS: We evaluated our systems on 2 clinical corpora. One is a subset of the Temporal Histories of Your Medical Events (THYME) corpus, which was used in SemEval 2015 Task 6: Clinical TempEval. The other is the 2012 Informatics for Integrating Biology and the Bedside (i2b2) challenge corpus. We designed multiple supervised machine learning models to compute the DCT relation and within-sentence temporal relations. For the i2b2 data, we also developed models and rule-based methods to recognize cross-sentence temporal relations. We used the official evaluation scripts of both challenges to make our results comparable with results of other participating systems. In addition, we conducted a feature ablation study to find out the contribution of various features to the system's performance. RESULTS: Our system achieved state-of-the-art performance on the Clinical TempEval corpus and was on par with the best systems on the i2b2 2012 corpus. Particularly, on the Clinical TempEval corpus, our system established a new F1 score benchmark, statistically significant as compared to the baseline and the best participating system. CONCLUSION: Presented here is the first open-source clinical temporal relation discovery system. It was built using a multilayered temporal modeling strategy and achieved top performance in 2 major shared tasks.
Chen Lin 0002, Dmitriy Dligach, Timothy A. Miller, Steven Bethard, Guergana K. Savova
J. Am. Medical Informatics Assoc.4
2016 Efficient identification of nationally mandated reportable cancer cases using natural language processing and machine learning
abstract
OBJECTIVE: To help cancer registrars efficiently and accurately identify reportable cancer cases. MATERIAL AND METHODS: The Cancer Registry Control Panel (CRCP) was developed to detect mentions of reportable cancer cases using a pipeline built on the Unstructured Information Management Architecture - Asynchronous Scaleout (UIMA-AS) architecture containing the National Library of Medicine's UIMA MetaMap annotator as well as a variety of rule-based UIMA annotators that primarily act to filter out concepts referring to nonreportable cancers. CRCP inspects pathology reports nightly to identify pathology records containing relevant cancer concepts and combines this with diagnosis codes from the Clinical Electronic Data Warehouse to identify candidate cancer patients using supervised machine learning. Cancer mentions are highlighted in all candidate clinical notes and then sorted in CRCP's web interface for faster validation by cancer registrars. RESULTS: CRCP achieved an accuracy of 0.872 and detected reportable cancer cases with a precision of 0.843 and a recall of 0.848. CRCP increases throughput by 22.6% over a baseline (manual review) pathology report inspection system while achieving a higher precision and recall. Depending on registrar time constraints, CRCP can increase recall to 0.939 at the expense of precision by incorporating a data source information feature. CONCLUSION: CRCP demonstrates accurate results when applying natural language processing features to the problem of detecting patients with cases of reportable cancer from clinical notes. We show that implementing only a portion of cancer reporting rules in the form of regular expressions is sufficient to increase the precision, recall, and speed of the detection of reportable cancer cases when combined with off-the-shelf information extraction software and machine learning.
John D. Osborne, Matthew C. Wyatt, Andrew O. Westfall, James H. Willig, Steven Bethard, Geoffrey D. Gordon
J. Am. Medical Informatics Assoc.5
2015 Adapting Coreference Resolution for Narrative Processing
abstract
Domain adaptation is a challenge for supervised NLP systems because of expensive and time-consuming manual annotated resources.We present a novel method to adapt a supervised coreference resolution system trained on newswire to short narrative stories without retraining the system.The idea is to perform inference via an Integer Linear Programming (ILP) formulation with the features of narratives adopted as soft constraints.When testing on the UMIREC 1 and N2 2 corpora with the-stateof-the-art Berkeley coreference resolution system trained on OntoNotes 3 , our inference substantially outperforms the original inference on the CoNLL 2011 metric.
Quynh Ngoc Thi Do, Steven Bethard, Marie-Francine Moens
EMNLP2
2015 Feature-Rich Two-Stage Logistic Regression for Monolingual Alignment
abstract
Monolingual alignment is the task of pairing semantically similar units from two pieces of text.We report a top-performing supervised aligner that operates on short text snippets.We employ a large feature set to ( 1) encode similarities among semantic units (words and named entities) in context, and (2) address cooperation and competition for alignment among units in the same snippet.These features are deployed in a two-stage logistic regression framework for alignment.On two benchmark data sets, our aligner achieves F 1 scores of 92.1% and 88.5%, with statistically significant error reductions of 4.8% and 7.3% over the previous best aligner.It produces top results in extrinsic evaluation as well.
Md. Arafat Sultan, Steven Bethard, Tamara Sumner
EMNLP2
2015 Not All Character N-grams Are Created Equal: A Study in Authorship Attribution
abstract
Upendra Sapkota, Steven Bethard, Manuel Montes, Thamar Solorio. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Upendra Sapkota, Steven Bethard, Manuel Montes-y-Gómez, Thamar Solorio
HLT-NAACL2
2015 A survey on the application of recurrent neural networks to statistical language modeling
abstract
In this paper, we present a survey on the application of recurrent neural networks to the task of statistical language modeling. Although it has been shown that these models obtain good performance on this task, often superior to other state-of-the-art techniques, they suffer from some important drawbacks, including a very long training time and limitations on the number of context words that can be taken into account in practice. Recent extensions to recurrent neural network models have been developed in an attempt to address these drawbacks. This paper gives an overview of the most important extensions. Each technique is described and its performance on statistical language modeling, as described in the existing literature, is discussed. Our structured overview makes it possible to detect the most promising techniques in the field of recurrent neural networks, applied to language modeling, but it also highlights the techniques for which further research is required.
Wim De Mulder, Steven Bethard, Marie-Francine Moens
Comput. Speech Lang.2
2015 Domain Adaptation in Semantic Role Labeling Using a Neural Language Model and Linguistic Resources
abstract
We propose a method for adapting Semantic Role Labeling (SRL) systems from a source domain to a target domain by combining a neural language model and linguistic resources to generate additional training examples. We primarily aim to improve the results of Location, Time, Manner and Direction roles. In our methodology, main words of selected predicates and arguments in the source-domain training data are replaced with words from the target domain. The replacement words are generated by a language model and then filtered by several linguistic filters (including Part-Of-Speech (POS), WordNet and Predicate constraints). In experiments on the out-of-domain CoNLL 2009 data, with the Recurrent Neural Network Language Model (RNNLM) and a well-known semantic parser from Lund University, we show enhanced recall and F1 without penalizing precision on the four targeted roles. These results improve the results of the same SRL system without using the language model and the linguistic resources, and are better than the results of the same SRL system that is trained with examples that are enriched with word embeddings. We also demonstrate the importance of using a language model and the vocabulary of the target domain when generating new training examples.
Quynh Ngoc Thi Do, Steven Bethard, Marie-Francine Moens
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Easy does it: more usable CAPTCHAs
abstract
Websites present users with puzzles called CAPTCHAs to curb abuse caused by computer algorithms masquerading as people. While CAPTCHAs are generally effective at stopping abuse, they might impair website usability if they are not properly designed. In this paper we describe how we designed two new CAPTCHA schemes for Google that focus on maximizing usability. We began by running an evaluation on Amazon Mechanical Turk with over 27,000 respondents to test the usability of different feature combinations. Then we studied user preferences using Google's consumer survey infrastructure. Finally, drawing on the insights gleaned during those studies, we tested our new captcha schemes first on Mechanical Turk and then on a fraction of production traffic. The resulting scheme is now an integral part of our production system and is served to millions of users. Our scheme achieved a 95.3% human accuracy, a 6.7.
Elie Bursztein, Angelique Moscicki, Celine Fabry, Steven Bethard, John C. Mitchell, Daniel Jurafsky
CHI4
2014 Cross-Topic Authorship Attribution: Will Out-Of-Topic Data Help?
Upendra Sapkota, Thamar Solorio, Manuel Montes-y-Gómez, Steven Bethard, Paolo Rosso
COLING4
2014 ClearTK 2.0: Design Patterns for Machine Learning in UIMA
Steven Bethard, Philip V. Ogren, Lee Becker
LREC1
2014 Discovering body site and severity modifiers in clinical texts
abstract
OBJECTIVE: To research computational methods for discovering body site and severity modifiers in clinical texts. METHODS: We cast the task of discovering body site and severity modifiers as a relation extraction problem in the context of a supervised machine learning framework. We utilize rich linguistic features to represent the pairs of relation arguments and delegate the decision about the nature of the relationship between them to a support vector machine model. We evaluate our models using two corpora that annotate body site and severity modifiers. We also compare the model performance to a number of rule-based baselines. We conduct cross-domain portability experiments. In addition, we carry out feature ablation experiments to determine the contribution of various feature groups. Finally, we perform error analysis and report the sources of errors. RESULTS: The performance of our method for discovering body site modifiers achieves F1 of 0.740-0.908 and our method for discovering severity modifiers achieves F1 of 0.905-0.929. DISCUSSION: Results indicate that both methods perform well on both in-domain and out-domain data, approaching the performance of human annotators. The most salient features are token and named entity features, although syntactic dependency features also contribute to the overall performance. The dominant sources of errors are infrequent patterns in the data and inability of the system to discern deeper semantic structures. CONCLUSIONS: We investigated computational methods for discovering body site and severity modifiers in clinical texts. Our best system is released open source as part of the clinical Text Analysis and Knowledge Extraction System (cTAKES).
Dmitriy Dligach, Steven Bethard, Lee Becker, Timothy A. Miller, Guergana K. Savova
J. Am. Medical Informatics Assoc.2
2014 Dense Event Ordering with a Multi-Pass Architecture
abstract
The past 10 years of event ordering research has focused on learning partial orderings over document events and time expressions. The most popular corpus, the TimeBank, contains a small subset of the possible ordering graph. Many evaluations follow suit by only testing certain pairs of events (e.g., only main verbs of neighboring sentences). This has led most research to focus on specific learners for partial labelings. This paper attempts to nudge the discussion from identifying some relations to all relations. We present new experiments on strongly connected event graphs that contain ∼10 times more relations per document than the TimeBank. We also describe a shift away from the single learner to a sieve-based architecture that naturally blends multiple learners into a precision-ranked cascade of sieves. Each sieve adds labels to the event graph one at a time, and earlier sieves inform later ones through transitive closure. This paper thus describes innovations in both approach and task. We experiment on the densest event graphs to date and show a 14% gain over state-of-the-art.
Nathanael Chambers, Taylor Cassidy, Bill McDowell, Steven Bethard
Trans. Assoc. Comput. Linguistics4
2014 Temporal Annotation in the Clinical Domain
abstract
This article discusses the requirements of a formal specification for the annotation of temporal information in clinical narratives. We discuss the implementation and extension of ISO-TimeML for annotating a corpus of clinical notes, known as the THYME corpus. To reflect the information task and the heavily inference-based reasoning demands in the domain, a new annotation guideline has been developed, "the THYME Guidelines to ISO-TimeML (THYME-TimeML)". To clarify what relations merit annotation, we distinguish between linguistically-derived and inferentially-derived temporal orderings in the text. We also apply a top performing TempEval 2013 system against this new resource to measure the difficulty of adapting systems to the clinical domain. The corpus is available to the community and has been proposed for use in a SemEval 2015 task.
William F. Styler IV, Steven Bethard, Sean Finan, Martha Palmer, Sameer Pradhan, Piet C. de Groen, Bradley James Erickson, Timothy A. Miller, Chen Lin 0002, Guergana K. Savova, James Pustejovsky
Trans. Assoc. Comput. Linguistics2
2014 Back to Basics for Monolingual Alignment: Exploiting Word Similarity and Contextual Evidence
abstract
We present a simple, easy-to-replicate monolingual aligner that demonstrates state-of-the-art performance while relying on almost no supervision and a very small number of external resources. Based on the hypothesis that words with similar meanings represent potential pairs for alignment if located in similar contexts, we propose a system that operates by finding such pairs. In two intrinsic evaluations on alignment test data, our system achieves F1 scores of 88–92%, demonstrating 1–3% absolute improvement over the previous best system. Moreover, in two extrinsic evaluations our aligner outperforms existing aligners, and even a naive application of the aligner approaches state-of-the-art performance in each extrinsic task.
Md. Arafat Sultan, Steven Bethard, Tamara Sumner
Trans. Assoc. Comput. Linguistics2
2013 Discovering Time Expressions in Clinical Text
Timothy A. Miller, Dmitriy Dligach, Steven Bethard, Sameer Pradhan, Chen Lin 0002, Guergana K. Savova
AMIA3
2013 A Synchronous Context Free Grammar for Time Normalization
abstract
⇒2013-04-12) based on a synchronous context free grammar. Synchronous rules map the source language to formally defined operators for manipulating times (FindEnclosed, StartAtEndOf, etc.). Time expressions are then parsed using an extended CYK+ algorithm, and converted to a normalized form by applying the operators recursively. For evaluation, a small set of synchronous rules for English time expressions were developed. Our model outperforms HeidelTime, the best time normalization system in TempEval 2013, on four different time normalization corpora.
Steven Bethard
EMNLP1
2013 Characterizing and Predicting the Multifaceted Nature of Quality in Educational Web Resources
abstract
Efficient learning from Web resources can depend on accurately assessing the quality of each resource. We present a methodology for developing computational models of quality that can assist users in assessing Web resources. The methodology consists of four steps: 1) a meta-analysis of previous studies to decompose quality into high-level dimensions and low-level indicators, 2) an expert study to identify the key low-level indicators of quality in the target domain, 3) human annotation to provide a collection of example resources where the presence or absence of quality indicators has been tagged, and 4) training of a machine learning model to predict quality indicators based on content and link features of Web resources. We find that quality is a multifaceted construct, with different aspects that may be important to different users at different times. We show that machine learning models can predict this multifaceted nature of quality, both in the context of aiding curators as they evaluate resources submitted to digital libraries, and in the context of aiding teachers as they develop online educational resources. Finally, we demonstrate how computational models of quality can be provided as a service, and embedded into applications such as Web search.
Philipp G. Wetzler, Steven Bethard, Heather Leary, Kirsten R. Butcher, Soheil Danesh Bahreini, James H. Martin, Tamara Sumner
ACM Trans. Interact. Intell. Syst.2
2012 Extracting Narrative Timelines as Temporal Dependency Structures
Oleksandr Kolomiyets, Steven Bethard, Marie-Francine Moens
ACL (1)2
2012 Skip N-grams and Ranking Functions for Predicting Script Events
Bram Jans, Steven Bethard, Ivan Vulic, Marie-Francine Moens
EACL2
2012 Annotating Story Timelines as Temporal Dependency Structures
Steven Bethard, Oleksandr Kolomiyets, Marie-Francine Moens
LREC1
2012 Citation-based bootstrapping for large-scale author disambiguation
abstract
We present a new, two‐stage, self‐supervised algorithm for author disambiguation in large bibliographic databases. In the first “bootstrap” stage, a collection of high‐precision features is used to bootstrap a training set with positive and negative examples of coreferring authors. A supervised feature‐based classifier is then trained on the bootstrap clusters and used to cluster the authors in a larger unlabeled dataset. Our self‐supervised approach shares the advantages of unsupervised approaches (no need for expensive hand labels) as well as supervised approaches (a rich set of features that can be discriminatively trained). The algorithm disambiguates 54,000,000 author instances in Thomson Reuters' Web of Knowledge with B3 F1 of.807. We analyze parameters and features, particularly those from citation networks, which have not been deeply investigated in author disambiguation. The most important citation feature is self‐citation, which can be approximated without expensive extraction of the full network. For the supervised stage, the minor improvement due to other citation features (increasing F1 from.748 to.767) suggests they may not be worth the trouble of extracting from databases that don't already have them. A lean feature set without expensive abstract and title features performs 130 times faster with about equal F1.
Michael Levin 0004, Stefan Krawczyk, Steven Bethard, Daniel Jurafsky
J. Assoc. Inf. Sci. Technol.3
2010 Who should I cite: learning literature search models from citation behavior
abstract
Scientists depend on literature search to find prior work that is relevant to their research ideas. We introduce a retrieval model for literature search that incorporates a wide variety of factors important to researchers, and learns the weights of each of these factors by observing citation patterns. We introduce features like topical similarity and author behavioral patterns, and combine these with features from related work like citation count and recency of publication. We present an iterative process for learning weights for these features that alternates between retrieving articles with the current retrieval model, and updating model weights by training a supervised classifier on these articles. We propose a new task for evaluating the resulting retrieval models, where the retrieval system takes only an abstract as its input and must produce as output the list of references at the end of the abstract's article. We evaluate our model on a collection of journal, conference and workshop articles from the ACL Anthology Reference Corpus. Our model achieves a mean average precision of 28.7, a 12.8 point improvement over a term similarity baseline, and a significant improvement both over models using only features from related work and over models without our iterative learning.
Steven Bethard, Daniel Jurafsky
CIKM1
2010 How Good Are Humans at Solving CAPTCHAs? A Large Scale Evaluation
abstract
Captchas are designed to be easy for humans but hard for machines. However, most recent research has focused only on making them hard for machines. In this paper, we present what is to the best of our knowledge the first large scale evaluation of captchas from the human perspective, with the goal of assessing how much friction captchas present to the average user. For the purpose of this study we have asked workers from Amazon's Mechanical Turk and an underground captchabreaking service to solve more than 318 000 captchas issued from the 21 most popular captcha schemes (13 images schemes and 8 audio scheme). Analysis of the resulting data reveals that captchas are often difficult for humans, with audio captchas being particularly problematic. We also find some demographic trends indicating, for example, that non-native speakers of English are slower in general and less accurate on English-centric captcha schemes. Evidence from a week's worth of eBay captchas (14,000,000 samples) suggests that the solving accuracies found in our study are close to real-world values, and that improving audio captchas should become a priority, as nearly 1% of all captchas are delivered as audio rather than images. Finally our study also reveals that it is more effective for an attacker to use Mechanical Turk to solve captchas than an underground service.
Elie Bursztein, Steven Bethard, Celine Fabry, John C. Mitchell, Daniel Jurafsky
IEEE Symposium on Security and Privacy2
2009 Towards Temporal Relation Discovery from the Clinical Narrative
Guergana K. Savova, Steven Bethard, William F. Styler IV, James H. Martin, Martha Palmer, James J. Masanz, Wayne H. Ward
AMIA2
2008 Building a Corpus of Temporal-Causal Structure
Steven Bethard, William J. Corvey, Sara Klingenstein, James H. Martin
LREC1
2008 Semantic role labeling for protein transport predicates
abstract
BACKGROUND: Automatic semantic role labeling (SRL) is a natural language processing (NLP) technique that maps sentences to semantic representations. This technique has been widely studied in the recent years, but mostly with data in newswire domains. Here, we report on a SRL model for identifying the semantic roles of biomedical predicates describing protein transport in GeneRIFs - manually curated sentences focusing on gene functions. To avoid the computational cost of syntactic parsing, and because the boundaries of our protein transport roles often did not match up with syntactic phrase boundaries, we approached this problem with a word-chunking paradigm and trained support vector machine classifiers to classify words as being at the beginning, inside or outside of a protein transport role. RESULTS: We collected a set of 837 GeneRIFs describing movements of proteins between cellular components, whose predicates were annotated for the semantic roles AGENT, PATIENT, ORIGIN and DESTINATION. We trained these models with the features of previous word-chunking models, features adapted from phrase-chunking models, and features derived from an analysis of our data. Our models were able to label protein transport semantic roles with 87.6% precision and 79.0% recall when using manually annotated protein boundaries, and 87.0% precision and 74.5% recall when using automatically identified ones. CONCLUSION: We successfully adapted the word-chunking classification paradigm to semantic role labeling, applying it to a new domain with predicates completely absent from any previous studies. By combining the traditional word and phrasal role labeling features with biomedical features like protein boundaries and MEDPOST part of speech tags, we were able to address the challenges posed by the new domain data and subsequently build robust models that achieved F-measures as high as 83.1. This system for extracting protein transport information from GeneRIFs performs well even with proteins identified automatically, and is therefore more robust than the rule-based methods previously used to extract protein transport roles.
Steven Bethard, Zhiyong Lu, James H. Martin, Lawrence Hunter
BMC Bioinform.1
2006 Identification of Event Mentions and their Semantic Class
Steven Bethard, James H. Martin
EMNLP1