EDBT 2026 Demo / reviewers in the wild / expert
Steven Bethard
dblp:52/5246 · also Steven J. Bethard
· DBLP profile ↗
56ranked-venue papers
8as first author
10since 2021 · last 2025
0000-0001-9560-6491ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 7 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
16 papers |
Information extraction and text analysis · 44% Transfer learning and domain adaptation · 19% Language models and text generation · 14% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 96% Machine learning and data management · 4% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 30 heaviest of 35, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
temporal information extraction |
2.0 | 4 | 2025 | A Semantic Parsing Framework for End-to-End Time Normalization · NeurIPS 2025 Transformer-Based Temporal Information Extraction and Application: A Review · EMNLP 2025 A Synchronous Context Free Grammar for Time Normalization · EMNLP 2013 |
Natural language and speech › Language models and text generation › language modeling › language model architecture
transformer language model |
0.9 | 1 | 2025 | Transformer-Based Temporal Information Extraction and Application: A Review · EMNLP 2025 |
Compilers and program optimization
code generation |
0.9 | 1 | 2025 | A Semantic Parsing Framework for End-to-End Time Normalization · NeurIPS 2025 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.8 | 2 | 2022 | A Comparison of Strategies for Source-Free Domain Adaptation · ACL (1) 2022 Domain Adaptation for Authorship Attribution: Improved Structural Correspondence Learning · ACL (1) 2016 |
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering |
0.8 | 2 | 2020 | Unsupervised Alignment-based Iterative Evidence Retrieval for Multi-hop Question Answering · ACL 2020 Quick and (not so) Dirty: Unsupervised Selection of Justification Sentences for Multi-hop Question Answering · EMNLP/IJCNLP (1) 2019 |
Machine learning › Learning theory › statistical learning theory
bias-variance tradeoff |
0.7 | 1 | 2023 | Two-Stage Fine-Tuning for Improved Bias and Variance for Large Pretrained Language Models · ACL (1) 2023 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.7 | 1 | 2023 | Two-Stage Fine-Tuning for Improved Bias and Variance for Large Pretrained Language Models · ACL (1) 2023 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
source-free domain adaptation |
0.6 | 1 | 2022 | A Comparison of Strategies for Source-Free Domain Adaptation · ACL (1) 2022 |
Natural language and speech › Information extraction and text analysis › fact-checking
evidence retrieval |
0.4 | 1 | 2020 | Unsupervised Alignment-based Iterative Evidence Retrieval for Multi-hop Question Answering · ACL 2020 |
Natural language and speech › Information extraction and text analysis › text normalization
medical term normalization |
0.4 | 1 | 2020 | A Generate-and-Rank Framework with Semantic Type Regularization for Biomedical Concept Normalization · ACL 2020 |
Natural language and speech › Language models and text generation › natural language understanding › question answering
neural question answering |
0.4 | 1 | 2020 | Having Your Cake and Eating it Too: Training Neural Retrieval for Language Inference without Losing Lexical Match · SIGIR 2020 |
Natural language and speech › Question answering and dialogue systems
open-domain question answering |
0.4 | 1 | 2020 | Having Your Cake and Eating it Too: Training Neural Retrieval for Language Inference without Losing Lexical Match · SIGIR 2020 |
Machine learning › Trustworthy machine learning › interpretability › neural network interpretation
transformer interpretability |
0.4 | 1 | 2020 | How does BERT's attention change when you fine-tune? An analysis methodology and a case study in negation scope · ACL 2020 |
Information retrieval › retrieval models › probabilistic retrieval model
BM25 |
0.4 | 1 | 2020 | Having Your Cake and Eating it Too: Training Neural Retrieval for Language Inference without Losing Lexical Match · SIGIR 2020 |
Information retrieval › similarity measure
lexical matching |
0.4 | 1 | 2020 | Having Your Cake and Eating it Too: Training Neural Retrieval for Language Inference without Losing Lexical Match · SIGIR 2020 |
Information retrieval › retrieval models
neural retrieval |
0.4 | 1 | 2020 | Having Your Cake and Eating it Too: Training Neural Retrieval for Language Inference without Losing Lexical Match · SIGIR 2020 |
Information retrieval
retrieval models |
0.4 | 1 | 2020 | Having Your Cake and Eating it Too: Training Neural Retrieval for Language Inference without Losing Lexical Match · SIGIR 2020 |
Medical and health informatics › mental health informatics
depression detection |
0.3 | 1 | 2018 | Measuring the Latency of Depression Detection in Social Media · WSDM 2018 |
Natural language and speech › Language models and text generation
large language model |
0.3 | 1 | 2025 | A Semantic Parsing Framework for End-to-End Time Normalization · NeurIPS 2025 |
Natural language and speech › Information extraction and text analysis › text mining › authorship analysis
authorship attribution |
0.2 | 1 | 2016 | Domain Adaptation for Authorship Attribution: Improved Structural Correspondence Learning · ACL (1) 2016 |
Natural language and speech › Information extraction and text analysis
coreference resolution |
0.2 | 1 | 2015 | Adapting Coreference Resolution for Narrative Processing · EMNLP 2015 |
Natural language and speech › Information extraction and text analysis
semantic role labeling |
0.2 | 1 | 2015 | Domain Adaptation in Semantic Role Labeling Using a Neural Language Model and Linguistic Resources · IEEE ACM Trans. Audio Speech Lang. Process. 2015 |
Machine learning › Kernel, tree and ensemble methods
ensemble learning |
0.2 | 1 | 2023 | Two-Stage Fine-Tuning for Improved Bias and Variance for Large Pretrained Language Models · ACL (1) 2023 |
Usability and user experience research
usability evaluation |
0.2 | 1 | 2014 | Easy does it: more usable CAPTCHAs · CHI 2014 |
Authentication and access control › human interactive proofs
CAPTCHA |
0.2 | 1 | 2014 | Easy does it: more usable CAPTCHAs · CHI 2014 |
Natural language and speech › Information extraction and text analysis › syntactic parsing › dependency parsing
temporal dependency parsing |
0.1 | 1 | 2012 | Extracting Narrative Timelines as Temporal Dependency Structures · ACL (1) 2012 |
Machine learning › Transfer learning and domain adaptation › domain alignment
unsupervised alignment |
0.1 | 1 | 2020 | Unsupervised Alignment-based Iterative Evidence Retrieval for Multi-hop Question Answering · ACL 2020 |
Usable security › security tool usability
CAPTCHA usability |
0.1 | 1 | 2010 | How Good Are Humans at Solving CAPTCHAs? A Large Scale Evaluation · IEEE Symposium on Security and Privacy 2010 |
Machine learning and data management
transfer learning |
0.1 | 1 | 2016 | Domain Adaptation for Authorship Attribution: Improved Structural Correspondence Learning · ACL (1) 2016 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
narrative processing |
0.1 | 1 | 2015 | Adapting Coreference Resolution for Narrative Processing · EMNLP 2015 |
Methods — techniques the papers use, named apart from their topics
data augmentation · 2.3large language model · 1.7RoBERTa · 1.3literature review · 0.9variance decomposition · 0.7ensemble methods · 0.7self-training · 0.6active learning · 0.6transformer · 0.4list-wise ranking · 0.4embedding-based alignment · 0.4BM25 · 0.4BERT · 0.4survey · 0.4production deployment · 0.4crowdsourced evaluation · 0.4latency-weighted f1 · 0.3ERDE metric analysis · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Transformer-Based Temporal Information Extraction and Application: A ReviewabstractTemporal information extraction (IE) aims to extract structured temporal information from unstructured text, thereby uncovering the implicit timelines within.This technique is applied across domains such as healthcare, newswire, and intelligence analysis, aiding models in these areas to perform temporal reasoning and enabling human users to grasp the temporal structure of text.Transformer-based pre-trained language models have produced revolutionary advancements in natural language processing, demonstrating exceptional performance across a multitude of tasks.Despite the achievements garnered by Transformer-based approaches in temporal IE, there is a lack of comprehensive reviews on these endeavors.In this paper, we aim to bridge this gap by systematically summarizing and analyzing the body of work on temporal IE using Transformers while highlighting potential future research directions. Xin Su 0008, Phillip Howard, Steven Bethard |
EMNLP | 3 |
| 2025 | A Semantic Parsing Framework for End-to-End Time NormalizationabstractTime normalization is the task of converting natural language temporal expressions into machine-readable representations. It underpins many downstream applications in information retrieval, question answering, and clinical decision-making. Traditional systems based on the ISO-TimeML schema limit expressivity and struggle with complex constructs such as compositional, event-relative, and multi-span time expressions. In this work, we introduce a novel formulation of time normalization as a code generation task grounded in the SCATE framework, which defines temporal semantics through symbolic and compositional operators. We implement a fully executable SCATE Python library and demonstrate that large language models (LLMs) can generate executable SCATE code. Leveraging this capability, we develop an automatic data augmentation pipeline using LLMs to synthesize large-scale annotated data with code-level validation. Our experiments show that small, locally deployable models trained on this augmented data can achieve strong performance, outperforming even their LLM parents and enabling practical, accurate, and interpretable time normalization. Xin Su 0008, Sungduk Yu, Phillip Howard, Steven Bethard |
NeurIPS | 4 |
| 2025 | Identifying task groupings for multi-task learning using pointwise V-usable information
Yingya Li, Timothy A. Miller, Steven Bethard, Guergana K. Savova |
J. Biomed. Informatics | 3 |
| 2024 | Semi-Structured Chain-of-Thought: Integrating Multiple Sources of Knowledge for Improved Language Model ReasoningabstractXin Su, Tiep Le, Steven Bethard, Phillip Howard. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Xin Su 0008, Tiep Le, Steven Bethard, Phillip Howard |
NAACL-HLT | 3 |
| 2023 | Two-Stage Fine-Tuning for Improved Bias and Variance for Large Pretrained Language ModelsabstractThe bias-variance tradeoff is the idea that learning methods need to balance model complexity with data size to minimize both under-fitting and over-fitting.Recent empirical work and theoretical analyses with over-parameterized neural networks challenge the classic bias-variance trade-off notion suggesting that no such trade-off holds: as the width of the network grows, bias monotonically decreases while variance initially increases followed by a decrease.In this work, we first provide a variance decomposition-based justification criteria to examine whether large pretrained neural models in a fine-tuning setting are generalizable enough to have low bias and variance.We then perform theoretical and empirical analysis using ensemble methods explicitly designed to decrease variance due to optimization.This results in essentially a two-stage fine-tuning algorithm that first ratchets down bias and variance iteratively, and then uses a selected fixed-bias model to further reduce variance due to optimization by ensembling.We also analyze the nature of variance change with the ensemble size in low-and high-resource classes.Empirical results show that this two-stage method obtains strong results on SuperGLUE tasks and clinical information extraction tasks.Code and settings are available: https://github.com/christa60/ bias-var-fine-tuning-plms.git Lijing Wang 0001, Yingya Li, Timothy A. Miller, Steven Bethard, Guergana K. Savova |
ACL (1) | 4 |
| 2023 | Addressing structural hurdles for metadata extraction from environmental impact statementsabstractAbstract Natural language processing techniques can be used to analyze the linguistic content of a document to extract missing pieces of metadata. However, accurate metadata extraction may not depend solely on the linguistics, but also on structural problems such as extremely large documents, unordered multi‐file documents, and inconsistency in manually labeled metadata. In this work, we start from two standard machine learning solutions to extract pieces of metadata from Environmental Impact Statements, environmental policy documents that are regularly produced under the US National Environmental Policy Act of 1969. We present a series of experiments where we evaluate how these standard approaches are affected by different issues derived from real‐world data. We find that metadata extraction can be strongly influenced by nonlinguistic factors such as document length and volume ordering and that the standard machine learning solutions often do not scale well to long documents. We demonstrate how such solutions can be better adapted to these scenarios, and conclude with suggestions for other NLP practitioners cataloging large document collections. Egoitz Laparra, Alex Binford-Walsh, Kirk Emerson, Marc L. Miller, Laura López-Hoffman, Faiz Currim, Steven Bethard |
J. Assoc. Inf. Sci. Technol. | 7 |
| 2022 | A Comparison of Strategies for Source-Free Domain AdaptationabstractData sharing restrictions are common in NLP, especially in the clinical domain, but there is limited research on adapting models to new domains without access to the original training data, a setting known as source-free domain adaptation.We take algorithms that traditionally assume access to the source-domain training data-active learning, self-training, and data augmentation-and adapt them for source-free domain adaptation.Then we systematically compare these different strategies across multiple tasks and domains.We find that active learning yields consistent gains across all SemEval 2021 Task 10 tasks and domains, but though the shared task saw successful self-trained and data augmented models, our systematic comparison finds these strategies to be unreliable for source-free domain adaptation. Xin Su 0008, Yiyun Zhao, Steven Bethard |
ACL (1) | 3 |
| 2021 | Do pretrained transformers infer telicity like humans?abstractPretrained transformer-based language models achieve state-of-the-art performance in many NLP tasks, but it is an open question whether the knowledge acquired by the models during pretraining resembles the linguistic knowledge of humans.We present both humans and pretrained transformers with descriptions of events, and measure their preference for telic interpretations (the event has a natural endpoint) or atelic interpretations (the event does not have a natural endpoint).To measure these preferences and determine what factors influence them, we design an English test and a novel-word test that include a variety of linguistic cues (noun phrase quantity, resultative structure, contextual information, temporal units) that bias toward certain interpretations.We find that humans' choice of telicity interpretation is reliably influenced by theoretically-motivated cues, transformer models (BERT and RoBERTa) are influenced by some (though not all) of the cues, and transformer models often rely more heavily on temporal units than humans do. Yiyun Zhao, Jian Gang Ngui, Lucy Hall Hartley, Steven Bethard |
CoNLL | 4 |
| 2021 | Explainable Multi-hop Verbal Reasoning Through Internal MonologueabstractMany state-of-the-art (SOTA) language models have achieved high accuracy on several multi-hop reasoning problems.However, these approaches tend to not be interpretable because they do not make the intermediate reasoning steps explicit.Moreover, models trained on simpler tasks tend to fail when directly tested on more complex problems.We propose the Explainable multi-hop Verbal Reasoner (EVR) to solve these limitations by (a) decomposing multi-hop reasoning problems into several simple ones, and (b) using natural language to guide the intermediate reasoning hops.We implement EVR by extending the classic reasoning paradigm General Problem Solver (GPS) with a SOTA generative language model to generate subgoals and perform inference in natural language at each reasoning step.Evaluation of EVR on Clark et al. (2020)'s synthetic question answering (QA) dataset shows that EVR achieves SOTA performance while being able to generate all reasoning steps in natural language.Furthermore, EVR generalizes better than other strong methods when trained on simpler tasks or less training data (up to 35.7% and 7.7% absolute improvement respectively). 1 Zhengzhong Liang, Steven Bethard, Mihai Surdeanu |
NAACL-HLT | 2 |
| 2021 | If You Want to Go Far Go Together: Unsupervised Joint Candidate Evidence Retrieval for Multi-hop Question AnsweringabstractMulti-hop reasoning requires aggregation and inference from multiple facts.To retrieve such facts, we propose a simple approach that retrieves and reranks set of evidence facts jointly.Our approach first generates unsupervised clusters of sentences as candidate evidence by accounting links between sentences and coverage with the given query.Then, a RoBERTa-based reranker is trained to bring the most representative evidence cluster to the top.We specifically emphasize on the importance of retrieving evidence jointly by showing several comparative analyses to other methods that retrieve and rerank evidence sentences individually.First, we introduce several attention-and embedding-based analyses, which indicate that jointly retrieving and reranking approaches can learn compositional knowledge required for multi-hop reasoning.Second, our experiments show that jointly retrieving candidate evidence leads to substantially higher evidence retrieval performance when fed to the same supervised reranker.In particular, our joint retrieval and then reranking approach achieves new state-of-the-art evidence retrieval performance on two multi-hop question answering (QA) datasets: 30.5 Recall@2 on QASC, and 67.6% F1 on MultiRC.When the evidence text from our joint retrieval approach is fed to a RoBERTa-based answer selection classifier, we achieve new state-ofthe-art QA performance on MultiRC and second best result on QASC. Vikas Yadav, Steven Bethard, Mihai Surdeanu |
NAACL-HLT | 2 |
| 2020 | A Generate-and-Rank Framework with Semantic Type Regularization for Biomedical Concept NormalizationabstractConcept normalization, the task of linking textual mentions of concepts to concepts in an ontology, is challenging because ontologies are large.In most cases, annotated datasets cover only a small sample of the concepts, yet concept normalizers are expected to predict all concepts in the ontology.In this paper, we propose an architecture consisting of a candidate generator and a list-wise ranker based on BERT.The ranker considers pairings of concept mentions and candidate concepts, allowing it to make predictions for any concept, not just those seen during training.We further enhance this list-wise approach with a semantic type regularizer that allows the model to incorporate semantic type information from the ontology during training.Our proposed concept normalization framework achieves stateof-the-art performance on multiple datasets. Dongfang Xu, Zeyu Zhang 0002, Steven Bethard |
ACL | 3 |
| 2020 | Unsupervised Alignment-based Iterative Evidence Retrieval for Multi-hop Question AnsweringabstractEvidence retrieval is a critical stage of question answering (QA), necessary not only to improve performance, but also to explain the decisions of the corresponding QA method.We introduce a simple, fast, and unsupervised iterative evidence retrieval method, which relies on three ideas: (a) an unsupervised alignment approach to soft-align questions and answers with justification sentences using only GloVe embeddings, (b) an iterative process that reformulates queries focusing on terms that are not covered by existing justifications, which (c) a stopping criterion that terminates retrieval when the terms in the given question and candidate answers are covered by the retrieved justifications.Despite its simplicity, our approach outperforms all the previous methods (including supervised methods) on the evidence selection task on two datasets: MultiRC and QASC.When these evidence sentences are fed into a RoBERTa answer classification component, we achieve state-of-the-art QA performance on these two datasets. Vikas Yadav, Steven Bethard, Mihai Surdeanu |
ACL | 2 |
| 2020 | How does BERT's attention change when you fine-tune? An analysis methodology and a case study in negation scopeabstractLarge pretrained language models like BERT, after fine-tuning to a downstream task, have achieved high performance on a variety of NLP problems.Yet explaining their decisions is difficult despite recent work probing their internal representations.We propose a procedure and analysis methods that take a hypothesis of how a transformer-based model might encode a linguistic phenomenon, and test the validity of that hypothesis based on a comparison between knowledge-related downstream tasks with downstream control tasks, and measurement of cross-dataset consistency.We apply this methodology to test BERT and RoBERTa on a hypothesis that some attention heads will consistently attend from a word in negation scope to the negation cue.We find that after fine-tuning BERT and RoBERTa on a negation scope task, the average attention head improves its sensitivity to negation and its attention consistency across negation datasets compared to the pre-trained models.However, only the base models (not the large models) improve compared to a control task, indicating there is evidence for a shallow encoding of negation only in the base models. Yiyun Zhao, Steven Bethard |
ACL | 2 |
| 2020 | A Dataset and Evaluation Framework for Complex Geographical Description ParsingabstractMuch previous work on geoparsing has focused on identifying and resolving individual toponyms in text like Adrano, S.Maria di Licodia or Catania.However, geographical locations occur not only as individual toponyms, but also as compositions of reference geolocations joined and modified by connectives, e.g., ". . .between the towns of Adrano and S.Maria di Licodia, 32 kilometres northwest of Catania".Ideally, a geoparser should be able to take such text, and the geographical shapes of the toponyms referenced within it, and parse these into a geographical shape, formed by a set of coordinates, that represents the location described.But creating a dataset for this complex geoparsing task is difficult and, if done manually, would require a huge amount of effort to annotate the geographical shapes of not only the geolocation described but also the reference toponyms.We present an approach that automates most of the process by combining Wikipedia and OpenStreetMap.As a result, we have gathered a collection of 360,187 uncurated complex geolocation descriptions, from which we have manually curated 1,000 examples intended to be used as a test set.To accompany the data, we define a new geoparsing evaluation framework along with a scoring methodology and a set of baselines. Egoitz Laparra, Steven Bethard |
COLING | 2 |
| 2020 | Having Your Cake and Eating it Too: Training Neural Retrieval for Language Inference without Losing Lexical MatchabstractWe present a study on the importance of information retrieval (IR) techniques for both the interpretability and the performance of neural question answering (QA) methods. We show that the current state-of-the-art transformer methods (like RoBERTa) encode poorly simple information retrieval (IR) concepts such as lexical overlap between query and the document. To mitigate this limitation, we introduce a supervised RoBERTa QA method that is trained to mimic the behavior of BM25 and the soft-matching idea behind embedding-based alignment methods. We show that fusing the simple lexical-matching IR concepts in transformer techniques results in improvement a) of their (lexical-matching) interpretability, b) retrieval performance, and c) the QA performance on two multi-hop QA datasets. We further highlight the lexical-chasm gap bridging capabilities of transformer methods by analyzing the attention distributions of the supervised RoBERTa classifier over the context versus lexically-matched token pairs. Vikas Yadav, Steven Bethard, Mihai Surdeanu |
SIGIR | 2 |
| 2020 | Does BERT need domain adaptation for clinical negation detection?abstractINTRODUCTION: Classifying whether concepts in an unstructured clinical text are negated is an important unsolved task. New domain adaptation and transfer learning methods can potentially address this issue. OBJECTIVE: We examine neural unsupervised domain adaptation methods, introducing a novel combination of domain adaptation with transformer-based transfer learning methods to improve negation detection. We also want to better understand the interaction between the widely used bidirectional encoder representations from transformers (BERT) system and domain adaptation methods. MATERIALS AND METHODS: We use 4 clinical text datasets that are annotated with negation status. We evaluate a neural unsupervised domain adaptation algorithm and BERT, a transformer-based model that is pretrained on massive general text datasets. We develop an extension to BERT that uses domain adversarial training, a neural domain adaptation method that adds an objective to the negation task, that the classifier should not be able to distinguish between instances from 2 different domains. RESULTS: The domain adaptation methods we describe show positive results, but, on average, the best performance is obtained by plain BERT (without the extension). We provide evidence that the gains from BERT are likely not additive with the gains from domain adaptation. DISCUSSION: Our results suggest that, at least for the task of clinical negation detection, BERT subsumes domain adaptation, implying that BERT is already learning very general representations of negation phenomena such that fine-tuning even on a specific corpus does not lead to much overfitting. CONCLUSION: Despite being trained on nonclinical text, the large training sets of models like BERT lead to large gains in performance for the clinical negation detection task. Chen Lin 0002, Steven Bethard, Dmitriy Dligach, Farig Sadeque, Guergana K. Savova, Timothy A. Miller |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | Unified Medical Language System resources improve sieve-based generation and Bidirectional Encoder Representations from Transformers (BERT)-based ranking for concept normalizationabstractOBJECTIVE: Concept normalization, the task of linking phrases in text to concepts in an ontology, is useful for many downstream tasks including relation extraction, information retrieval, etc. We present a generate-and-rank concept normalization system based on our participation in the 2019 National NLP Clinical Challenges Shared Task Track 3 Concept Normalization. MATERIALS AND METHODS: The shared task provided 13 609 concept mentions drawn from 100 discharge summaries. We first design a sieve-based system that uses Lucene indices over the training data, Unified Medical Language System (UMLS) preferred terms, and UMLS synonyms to generate a list of possible concepts for each mention. We then design a listwise classifier based on the BERT (Bidirectional Encoder Representations from Transformers) neural network to rank the candidate concepts, integrating UMLS semantic types through a regularizer. RESULTS: Our generate-and-rank system was third of 33 in the competition, outperforming the candidate generator alone (81.66% vs 79.44%) and the previous state of the art (76.35%). During postevaluation, the model's accuracy was increased to 83.56% via improvements to how training data are generated from UMLS and incorporation of our UMLS semantic type regularizer. DISCUSSION: Analysis of the model shows that prioritizing UMLS preferred terms yields better performance, that the UMLS semantic type regularizer results in qualitatively better concept predictions, and that the model performs well even on concepts not seen during training. CONCLUSIONS: Our generate-and-rank framework for UMLS concept normalization integrates key UMLS features like preferred terms and semantic types with a neural network-based ranking model to accurately link phrases in text to UMLS concepts. Dongfang Xu, Manoj Gopale, Kris Brown, Edmon Begoli, Steven Bethard |
J. Am. Medical Informatics Assoc. | 6 |
| 2019 | Quick and (not so) Dirty: Unsupervised Selection of Justification Sentences for Multi-hop Question AnsweringabstractVikas Yadav, Steven Bethard, Mihai Surdeanu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Vikas Yadav, Steven Bethard, Mihai Surdeanu |
EMNLP/IJCNLP (1) | 2 |
| 2018 | A Survey on Recent Advances in Named Entity Recognition from Deep Learning modelsabstractNamed Entity Recognition (NER) is a key component in NLP systems for question answering, information retrieval, relation extraction, etc. NER systems have been studied and developed widely for decades, but accurate systems using deep neural networks (NN) have only been introduced in the last few years. We present a comprehensive survey of deep neural network architectures for NER, and contrast them with previous approaches to NER based on feature engineering and other supervised or semi-supervised learning algorithms. Our results highlight the improvements achieved by neural networks, and show how incorporating some of the lessons learned from past work on feature-based NER systems can yield further improvements. Vikas Yadav, Steven Bethard |
COLING | 2 |
| 2018 | Measuring the Latency of Depression Detection in Social MediaabstractDetecting depression is a key public health challenge, as almost 12% of all disabilities can be attributed to depression. Computational models for depression detection must prove not only that can they detect depression, but that they can do it early enough for an intervention to be plausible. However, current evaluations of depression detection are poor at measuring model latency. We identify several issues with the currently popular ERDE metric, and propose a latency-weighted F1 metric that addresses these concerns. We then apply this evaluation to several models from the recent eRisk 2017 shared task on depression detection, and show how our proposed measure can better capture system differences. Farig Sadeque, Dongfang Xu, Steven Bethard |
WSDM | 3 |
| 2018 | From Characters to Time Intervals: New Paradigms for Evaluation and Neural Parsing of Time NormalizationsabstractThis paper presents the first model for time normalization trained on the SCATE corpus. In the SCATE schema, time expressions are annotated as a semantic composition of time entities. This novel schema favors machine learning approaches, as it can be viewed as a semantic parsing task. In this work, we propose a character level multi-output neural network that outperforms previous state-of-the-art built on the TimeML schema. To compare predictions of systems that follow both SCATE and TimeML, we present a new scoring metric for time intervals. We also apply this new metric to carry out a comparative analysis of the annotations of both schemes in the same corpus. Egoitz Laparra, Dongfang Xu, Steven Bethard |
Trans. Assoc. Comput. Linguistics | 3 |
| 2017 | Recurrent Neural Network Architectures for Event Extraction from Italian Medical Reports
Natalia Viani, Timothy A. Miller, Dmitriy Dligach, Steven Bethard, Carlo Napolitano, Silvia G. Priori, Riccardo Bellazzi, Lucia Sacchi, Guergana K. Savova |
AIME | 4 |
| 2017 | Improving Implicit Semantic Role Labeling by Predicting Semantic Frame ArgumentsabstractImplicit semantic role labeling (iSRL) is the task of predicting the semantic roles of a predicate that do not appear as explicit arguments, but rather regard common sense knowledge or are mentioned earlier in the discourse. We introduce an approach to iSRL based on a predictive recurrent neural semantic frame model (PRNSFM) that uses a large unannotated corpus to learn the probability of a sequence of semantic arguments given a predicate. We leverage the sequence probabilities predicted by the PRNSFM to estimate selectional preferences for predicates and their arguments. On the NomBank iSRL test set, our approach improves state-of-the-art performance on implicit semantic role labeling with less reliance than prior work on manually constructed language resources. Quynh Ngoc Thi Do, Steven Bethard, Marie-Francine Moens |
IJCNLP(1) | 2 |
| 2017 | Towards generalizable entity-centric clinical coreference resolution
Timothy A. Miller, Dmitriy Dligach, Steven Bethard, Chen Lin 0002, Guergana K. Savova |
J. Biomed. Informatics | 3 |
| 2016 | Domain Adaptation for Authorship Attribution: Improved Structural Correspondence Learning
Upendra Sapkota, Thamar Solorio, Manuel Montes-y-Gómez, Steven Bethard |
ACL (1) | 4 |
| 2016 | Feature Portability in Cross-domain Clinical Coreference
Timothy A. Miller, Dmitriy Dligach, Chen Lin 0002, Steven Bethard, Guergana K. Savova |
AMIA | 4 |
| 2016 | Facing the most difficult case of Semantic Role Labeling: A collaboration of word embeddings and co-trainingabstractWe present a successful collaboration of word embeddings and co-training to tackle in the most difficult test case of semantic role labeling: predicting out-of-domain and unseen semantic frames. Despite the fact that co-training is a successful traditional semi-supervised method, its application in SRL is very limited especially when a huge amount of labeled data is available. In this work, co-training is used together with word embeddings to improve the performance of a system trained on a large training dataset. We also introduce a semantic role labeling system with a simple learning architecture and effective inference that is easily adaptable to semi-supervised settings with new training data and/or new features. On the out-of-domain testing set of the standard benchmark CoNLL 2009 data our simple approach achieves high performance and improves state-of-the-art results. Quynh Ngoc Thi Do, Steven Bethard, Marie-Francine Moens |
COLING | 2 |
| 2016 | A Semantically Compositional Annotation Scheme for Time Normalization
Steven Bethard, Jonathan Parker |
LREC | 1 |
| 2016 | Age and Gender Prediction on Health Forum Data
Prasha Shrestha, Nicolas Rey-Villamizar, Farig Sadeque, Ted Pedersen, Steven Bethard, Thamar Solorio |
LREC | 5 |
| 2016 | Multilayered temporal modeling for the clinical domainabstractOBJECTIVE: To develop an open-source temporal relation discovery system for the clinical domain. The system is capable of automatically inferring temporal relations between events and time expressions using a multilayered modeling strategy. It can operate at different levels of granularity--from rough temporality expressed as event relations to the document creation time (DCT) to temporal containment to fine-grained classic Allen-style relations. MATERIALS AND METHODS: We evaluated our systems on 2 clinical corpora. One is a subset of the Temporal Histories of Your Medical Events (THYME) corpus, which was used in SemEval 2015 Task 6: Clinical TempEval. The other is the 2012 Informatics for Integrating Biology and the Bedside (i2b2) challenge corpus. We designed multiple supervised machine learning models to compute the DCT relation and within-sentence temporal relations. For the i2b2 data, we also developed models and rule-based methods to recognize cross-sentence temporal relations. We used the official evaluation scripts of both challenges to make our results comparable with results of other participating systems. In addition, we conducted a feature ablation study to find out the contribution of various features to the system's performance. RESULTS: Our system achieved state-of-the-art performance on the Clinical TempEval corpus and was on par with the best systems on the i2b2 2012 corpus. Particularly, on the Clinical TempEval corpus, our system established a new F1 score benchmark, statistically significant as compared to the baseline and the best participating system. CONCLUSION: Presented here is the first open-source clinical temporal relation discovery system. It was built using a multilayered temporal modeling strategy and achieved top performance in 2 major shared tasks. Chen Lin 0002, Dmitriy Dligach, Timothy A. Miller, Steven Bethard, Guergana K. Savova |
J. Am. Medical Informatics Assoc. | 4 |
| 2016 | Efficient identification of nationally mandated reportable cancer cases using natural language processing and machine learningabstractOBJECTIVE: To help cancer registrars efficiently and accurately identify reportable cancer cases. MATERIAL AND METHODS: The Cancer Registry Control Panel (CRCP) was developed to detect mentions of reportable cancer cases using a pipeline built on the Unstructured Information Management Architecture - Asynchronous Scaleout (UIMA-AS) architecture containing the National Library of Medicine's UIMA MetaMap annotator as well as a variety of rule-based UIMA annotators that primarily act to filter out concepts referring to nonreportable cancers. CRCP inspects pathology reports nightly to identify pathology records containing relevant cancer concepts and combines this with diagnosis codes from the Clinical Electronic Data Warehouse to identify candidate cancer patients using supervised machine learning. Cancer mentions are highlighted in all candidate clinical notes and then sorted in CRCP's web interface for faster validation by cancer registrars. RESULTS: CRCP achieved an accuracy of 0.872 and detected reportable cancer cases with a precision of 0.843 and a recall of 0.848. CRCP increases throughput by 22.6% over a baseline (manual review) pathology report inspection system while achieving a higher precision and recall. Depending on registrar time constraints, CRCP can increase recall to 0.939 at the expense of precision by incorporating a data source information feature. CONCLUSION: CRCP demonstrates accurate results when applying natural language processing features to the problem of detecting patients with cases of reportable cancer from clinical notes. We show that implementing only a portion of cancer reporting rules in the form of regular expressions is sufficient to increase the precision, recall, and speed of the detection of reportable cancer cases when combined with off-the-shelf information extraction software and machine learning. John D. Osborne, Matthew C. Wyatt, Andrew O. Westfall, James H. Willig, Steven Bethard, Geoffrey D. Gordon |
J. Am. Medical Informatics Assoc. | 5 |
| 2015 | Adapting Coreference Resolution for Narrative ProcessingabstractDomain adaptation is a challenge for supervised NLP systems because of expensive and time-consuming manual annotated resources.We present a novel method to adapt a supervised coreference resolution system trained on newswire to short narrative stories without retraining the system.The idea is to perform inference via an Integer Linear Programming (ILP) formulation with the features of narratives adopted as soft constraints.When testing on the UMIREC 1 and N2 2 corpora with the-stateof-the-art Berkeley coreference resolution system trained on OntoNotes 3 , our inference substantially outperforms the original inference on the CoNLL 2011 metric. Quynh Ngoc Thi Do, Steven Bethard, Marie-Francine Moens |
EMNLP | 2 |
| 2015 | Feature-Rich Two-Stage Logistic Regression for Monolingual AlignmentabstractMonolingual alignment is the task of pairing semantically similar units from two pieces of text.We report a top-performing supervised aligner that operates on short text snippets.We employ a large feature set to ( 1) encode similarities among semantic units (words and named entities) in context, and (2) address cooperation and competition for alignment among units in the same snippet.These features are deployed in a two-stage logistic regression framework for alignment.On two benchmark data sets, our aligner achieves F 1 scores of 92.1% and 88.5%, with statistically significant error reductions of 4.8% and 7.3% over the previous best aligner.It produces top results in extrinsic evaluation as well. Md. Arafat Sultan, Steven Bethard, Tamara Sumner |
EMNLP | 2 |
| 2015 | Not All Character N-grams Are Created Equal: A Study in Authorship AttributionabstractUpendra Sapkota, Steven Bethard, Manuel Montes, Thamar Solorio. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Upendra Sapkota, Steven Bethard, Manuel Montes-y-Gómez, Thamar Solorio |
HLT-NAACL | 2 |
| 2015 | A survey on the application of recurrent neural networks to statistical language modelingabstractIn this paper, we present a survey on the application of recurrent neural networks to the task of statistical language modeling. Although it has been shown that these models obtain good performance on this task, often superior to other state-of-the-art techniques, they suffer from some important drawbacks, including a very long training time and limitations on the number of context words that can be taken into account in practice. Recent extensions to recurrent neural network models have been developed in an attempt to address these drawbacks. This paper gives an overview of the most important extensions. Each technique is described and its performance on statistical language modeling, as described in the existing literature, is discussed. Our structured overview makes it possible to detect the most promising techniques in the field of recurrent neural networks, applied to language modeling, but it also highlights the techniques for which further research is required. Wim De Mulder, Steven Bethard, Marie-Francine Moens |
Comput. Speech Lang. | 2 |
| 2015 | Domain Adaptation in Semantic Role Labeling Using a Neural Language Model and Linguistic ResourcesabstractWe propose a method for adapting Semantic Role Labeling (SRL) systems from a source domain to a target domain by combining a neural language model and linguistic resources to generate additional training examples. We primarily aim to improve the results of Location, Time, Manner and Direction roles. In our methodology, main words of selected predicates and arguments in the source-domain training data are replaced with words from the target domain. The replacement words are generated by a language model and then filtered by several linguistic filters (including Part-Of-Speech (POS), WordNet and Predicate constraints). In experiments on the out-of-domain CoNLL 2009 data, with the Recurrent Neural Network Language Model (RNNLM) and a well-known semantic parser from Lund University, we show enhanced recall and F1 without penalizing precision on the four targeted roles. These results improve the results of the same SRL system without using the language model and the linguistic resources, and are better than the results of the same SRL system that is trained with examples that are enriched with word embeddings. We also demonstrate the importance of using a language model and the vocabulary of the target domain when generating new training examples. Quynh Ngoc Thi Do, Steven Bethard, Marie-Francine Moens |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Easy does it: more usable CAPTCHAsabstractWebsites present users with puzzles called CAPTCHAs to curb abuse caused by computer algorithms masquerading as people. While CAPTCHAs are generally effective at stopping abuse, they might impair website usability if they are not properly designed. In this paper we describe how we designed two new CAPTCHA schemes for Google that focus on maximizing usability. We began by running an evaluation on Amazon Mechanical Turk with over 27,000 respondents to test the usability of different feature combinations. Then we studied user preferences using Google's consumer survey infrastructure. Finally, drawing on the insights gleaned during those studies, we tested our new captcha schemes first on Mechanical Turk and then on a fraction of production traffic. The resulting scheme is now an integral part of our production system and is served to millions of users. Our scheme achieved a 95.3% human accuracy, a 6.7. Elie Bursztein, Angelique Moscicki, Celine Fabry, Steven Bethard, John C. Mitchell, Daniel Jurafsky |
CHI | 4 |
| 2014 | Cross-Topic Authorship Attribution: Will Out-Of-Topic Data Help?
Upendra Sapkota, Thamar Solorio, Manuel Montes-y-Gómez, Steven Bethard, Paolo Rosso |
COLING | 4 |
| 2014 | ClearTK 2.0: Design Patterns for Machine Learning in UIMA
Steven Bethard, Philip V. Ogren, Lee Becker |
LREC | 1 |
| 2014 | Discovering body site and severity modifiers in clinical textsabstractOBJECTIVE: To research computational methods for discovering body site and severity modifiers in clinical texts. METHODS: We cast the task of discovering body site and severity modifiers as a relation extraction problem in the context of a supervised machine learning framework. We utilize rich linguistic features to represent the pairs of relation arguments and delegate the decision about the nature of the relationship between them to a support vector machine model. We evaluate our models using two corpora that annotate body site and severity modifiers. We also compare the model performance to a number of rule-based baselines. We conduct cross-domain portability experiments. In addition, we carry out feature ablation experiments to determine the contribution of various feature groups. Finally, we perform error analysis and report the sources of errors. RESULTS: The performance of our method for discovering body site modifiers achieves F1 of 0.740-0.908 and our method for discovering severity modifiers achieves F1 of 0.905-0.929. DISCUSSION: Results indicate that both methods perform well on both in-domain and out-domain data, approaching the performance of human annotators. The most salient features are token and named entity features, although syntactic dependency features also contribute to the overall performance. The dominant sources of errors are infrequent patterns in the data and inability of the system to discern deeper semantic structures. CONCLUSIONS: We investigated computational methods for discovering body site and severity modifiers in clinical texts. Our best system is released open source as part of the clinical Text Analysis and Knowledge Extraction System (cTAKES). Dmitriy Dligach, Steven Bethard, Lee Becker, Timothy A. Miller, Guergana K. Savova |
J. Am. Medical Informatics Assoc. | 2 |
| 2014 | Dense Event Ordering with a Multi-Pass ArchitectureabstractThe past 10 years of event ordering research has focused on learning partial orderings over document events and time expressions. The most popular corpus, the TimeBank, contains a small subset of the possible ordering graph. Many evaluations follow suit by only testing certain pairs of events (e.g., only main verbs of neighboring sentences). This has led most research to focus on specific learners for partial labelings. This paper attempts to nudge the discussion from identifying some relations to all relations. We present new experiments on strongly connected event graphs that contain ∼10 times more relations per document than the TimeBank. We also describe a shift away from the single learner to a sieve-based architecture that naturally blends multiple learners into a precision-ranked cascade of sieves. Each sieve adds labels to the event graph one at a time, and earlier sieves inform later ones through transitive closure. This paper thus describes innovations in both approach and task. We experiment on the densest event graphs to date and show a 14% gain over state-of-the-art. Nathanael Chambers, Taylor Cassidy, Bill McDowell, Steven Bethard |
Trans. Assoc. Comput. Linguistics | 4 |
| 2014 | Temporal Annotation in the Clinical DomainabstractThis article discusses the requirements of a formal specification for the annotation of temporal information in clinical narratives. We discuss the implementation and extension of ISO-TimeML for annotating a corpus of clinical notes, known as the THYME corpus. To reflect the information task and the heavily inference-based reasoning demands in the domain, a new annotation guideline has been developed, "the THYME Guidelines to ISO-TimeML (THYME-TimeML)". To clarify what relations merit annotation, we distinguish between linguistically-derived and inferentially-derived temporal orderings in the text. We also apply a top performing TempEval 2013 system against this new resource to measure the difficulty of adapting systems to the clinical domain. The corpus is available to the community and has been proposed for use in a SemEval 2015 task. William F. Styler IV, Steven Bethard, Sean Finan, Martha Palmer, Sameer Pradhan, Piet C. de Groen, Bradley James Erickson, Timothy A. Miller, Chen Lin 0002, Guergana K. Savova, James Pustejovsky |
Trans. Assoc. Comput. Linguistics | 2 |
| 2014 | Back to Basics for Monolingual Alignment: Exploiting Word Similarity and Contextual EvidenceabstractWe present a simple, easy-to-replicate monolingual aligner that demonstrates state-of-the-art performance while relying on almost no supervision and a very small number of external resources. Based on the hypothesis that words with similar meanings represent potential pairs for alignment if located in similar contexts, we propose a system that operates by finding such pairs. In two intrinsic evaluations on alignment test data, our system achieves F1 scores of 88–92%, demonstrating 1–3% absolute improvement over the previous best system. Moreover, in two extrinsic evaluations our aligner outperforms existing aligners, and even a naive application of the aligner approaches state-of-the-art performance in each extrinsic task. Md. Arafat Sultan, Steven Bethard, Tamara Sumner |
Trans. Assoc. Comput. Linguistics | 2 |
| 2013 | Discovering Time Expressions in Clinical Text
Timothy A. Miller, Dmitriy Dligach, Steven Bethard, Sameer Pradhan, Chen Lin 0002, Guergana K. Savova |
AMIA | 3 |
| 2013 | A Synchronous Context Free Grammar for Time Normalizationabstract⇒2013-04-12) based on a synchronous context free grammar. Synchronous rules map the source language to formally defined operators for manipulating times (FindEnclosed, StartAtEndOf, etc.). Time expressions are then parsed using an extended CYK+ algorithm, and converted to a normalized form by applying the operators recursively. For evaluation, a small set of synchronous rules for English time expressions were developed. Our model outperforms HeidelTime, the best time normalization system in TempEval 2013, on four different time normalization corpora. Steven Bethard |
EMNLP | 1 |
| 2013 | Characterizing and Predicting the Multifaceted Nature of Quality in Educational Web ResourcesabstractEfficient learning from Web resources can depend on accurately assessing the quality of each resource. We present a methodology for developing computational models of quality that can assist users in assessing Web resources. The methodology consists of four steps: 1) a meta-analysis of previous studies to decompose quality into high-level dimensions and low-level indicators, 2) an expert study to identify the key low-level indicators of quality in the target domain, 3) human annotation to provide a collection of example resources where the presence or absence of quality indicators has been tagged, and 4) training of a machine learning model to predict quality indicators based on content and link features of Web resources. We find that quality is a multifaceted construct, with different aspects that may be important to different users at different times. We show that machine learning models can predict this multifaceted nature of quality, both in the context of aiding curators as they evaluate resources submitted to digital libraries, and in the context of aiding teachers as they develop online educational resources. Finally, we demonstrate how computational models of quality can be provided as a service, and embedded into applications such as Web search. Philipp G. Wetzler, Steven Bethard, Heather Leary, Kirsten R. Butcher, Soheil Danesh Bahreini, James H. Martin, Tamara Sumner |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2012 | Extracting Narrative Timelines as Temporal Dependency Structures
Oleksandr Kolomiyets, Steven Bethard, Marie-Francine Moens |
ACL (1) | 2 |
| 2012 | Skip N-grams and Ranking Functions for Predicting Script Events
Bram Jans, Steven Bethard, Ivan Vulic, Marie-Francine Moens |
EACL | 2 |
| 2012 | Annotating Story Timelines as Temporal Dependency Structures
Steven Bethard, Oleksandr Kolomiyets, Marie-Francine Moens |
LREC | 1 |
| 2012 | Citation-based bootstrapping for large-scale author disambiguationabstractWe present a new, two‐stage, self‐supervised algorithm for author disambiguation in large bibliographic databases. In the first “bootstrap” stage, a collection of high‐precision features is used to bootstrap a training set with positive and negative examples of coreferring authors. A supervised feature‐based classifier is then trained on the bootstrap clusters and used to cluster the authors in a larger unlabeled dataset. Our self‐supervised approach shares the advantages of unsupervised approaches (no need for expensive hand labels) as well as supervised approaches (a rich set of features that can be discriminatively trained). The algorithm disambiguates 54,000,000 author instances in Thomson Reuters' Web of Knowledge with B3 F1 of.807. We analyze parameters and features, particularly those from citation networks, which have not been deeply investigated in author disambiguation. The most important citation feature is self‐citation, which can be approximated without expensive extraction of the full network. For the supervised stage, the minor improvement due to other citation features (increasing F1 from.748 to.767) suggests they may not be worth the trouble of extracting from databases that don't already have them. A lean feature set without expensive abstract and title features performs 130 times faster with about equal F1. Michael Levin 0004, Stefan Krawczyk, Steven Bethard, Daniel Jurafsky |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2010 | Who should I cite: learning literature search models from citation behaviorabstractScientists depend on literature search to find prior work that is relevant to their research ideas. We introduce a retrieval model for literature search that incorporates a wide variety of factors important to researchers, and learns the weights of each of these factors by observing citation patterns. We introduce features like topical similarity and author behavioral patterns, and combine these with features from related work like citation count and recency of publication. We present an iterative process for learning weights for these features that alternates between retrieving articles with the current retrieval model, and updating model weights by training a supervised classifier on these articles. We propose a new task for evaluating the resulting retrieval models, where the retrieval system takes only an abstract as its input and must produce as output the list of references at the end of the abstract's article. We evaluate our model on a collection of journal, conference and workshop articles from the ACL Anthology Reference Corpus. Our model achieves a mean average precision of 28.7, a 12.8 point improvement over a term similarity baseline, and a significant improvement both over models using only features from related work and over models without our iterative learning. Steven Bethard, Daniel Jurafsky |
CIKM | 1 |
| 2010 | How Good Are Humans at Solving CAPTCHAs? A Large Scale EvaluationabstractCaptchas are designed to be easy for humans but hard for machines. However, most recent research has focused only on making them hard for machines. In this paper, we present what is to the best of our knowledge the first large scale evaluation of captchas from the human perspective, with the goal of assessing how much friction captchas present to the average user. For the purpose of this study we have asked workers from Amazon's Mechanical Turk and an underground captchabreaking service to solve more than 318 000 captchas issued from the 21 most popular captcha schemes (13 images schemes and 8 audio scheme). Analysis of the resulting data reveals that captchas are often difficult for humans, with audio captchas being particularly problematic. We also find some demographic trends indicating, for example, that non-native speakers of English are slower in general and less accurate on English-centric captcha schemes. Evidence from a week's worth of eBay captchas (14,000,000 samples) suggests that the solving accuracies found in our study are close to real-world values, and that improving audio captchas should become a priority, as nearly 1% of all captchas are delivered as audio rather than images. Finally our study also reveals that it is more effective for an attacker to use Mechanical Turk to solve captchas than an underground service. Elie Bursztein, Steven Bethard, Celine Fabry, John C. Mitchell, Daniel Jurafsky |
IEEE Symposium on Security and Privacy | 2 |
| 2009 | Towards Temporal Relation Discovery from the Clinical Narrative
Guergana K. Savova, Steven Bethard, William F. Styler IV, James H. Martin, Martha Palmer, James J. Masanz, Wayne H. Ward |
AMIA | 2 |
| 2008 | Building a Corpus of Temporal-Causal Structure
Steven Bethard, William J. Corvey, Sara Klingenstein, James H. Martin |
LREC | 1 |
| 2008 | Semantic role labeling for protein transport predicatesabstractBACKGROUND: Automatic semantic role labeling (SRL) is a natural language processing (NLP) technique that maps sentences to semantic representations. This technique has been widely studied in the recent years, but mostly with data in newswire domains. Here, we report on a SRL model for identifying the semantic roles of biomedical predicates describing protein transport in GeneRIFs - manually curated sentences focusing on gene functions. To avoid the computational cost of syntactic parsing, and because the boundaries of our protein transport roles often did not match up with syntactic phrase boundaries, we approached this problem with a word-chunking paradigm and trained support vector machine classifiers to classify words as being at the beginning, inside or outside of a protein transport role. RESULTS: We collected a set of 837 GeneRIFs describing movements of proteins between cellular components, whose predicates were annotated for the semantic roles AGENT, PATIENT, ORIGIN and DESTINATION. We trained these models with the features of previous word-chunking models, features adapted from phrase-chunking models, and features derived from an analysis of our data. Our models were able to label protein transport semantic roles with 87.6% precision and 79.0% recall when using manually annotated protein boundaries, and 87.0% precision and 74.5% recall when using automatically identified ones. CONCLUSION: We successfully adapted the word-chunking classification paradigm to semantic role labeling, applying it to a new domain with predicates completely absent from any previous studies. By combining the traditional word and phrasal role labeling features with biomedical features like protein boundaries and MEDPOST part of speech tags, we were able to address the challenges posed by the new domain data and subsequently build robust models that achieved F-measures as high as 83.1. This system for extracting protein transport information from GeneRIFs performs well even with proteins identified automatically, and is therefore more robust than the rule-based methods previously used to extract protein transport roles. Steven Bethard, Zhiyong Lu, James H. Martin, Lawrence Hunter |
BMC Bioinform. | 1 |
| 2006 | Identification of Event Mentions and their Semantic Class
Steven Bethard, James H. Martin |
EMNLP | 1 |