VLDB 2026 Research / reviewers in the wild / expert
Byron C. Wallace
dblp:00/8247
· DBLP profile ↗
82ranked-venue papers
15as first author
35since 2021 · last 2025
0000-0003-2409-7735ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 63 · 10 first-author · 31 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 10 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-authorHuman-computer interaction and ubiquitous computing · 5 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NNsight and NDIF: Democratizing Access to Open-Weight Foundation Model InternalsabstractWe introduce NNsight and NDIF, technologies that work in tandem to enable scientific study of the representations and computations learned by very large neural networks. NNsight is an open-source system that extends PyTorch to introduce deferred remote execution. The National Deep Inference Fabric (NDIF) is a scalable inference service that executes NNsight requests, allowing users to share GPU resources and pretrained models. These technologies are enabled by the Intervention Graph, an architecture developed to decouple experimental design from model runtime. Together, this framework provides transparent and efficient access to the internals of deep neural networks such as very large language models (LLMs) without imposing the cost or complexity of hosting customized models individually. We conduct a quantitative survey of the machine learning literature that reveals a growing gap in the study of the internals of large-scale AI. We demonstrate the design and use of our framework to address this gap by enabling a range of research methods on huge models. Finally, we conduct benchmarks to compare performance with previous approaches.
Code, documentation, and tutorials are available at https://nnsight.net/. Jaden Fiotto-Kaufman, Alexander R. Loftus, Eric Todd, Jannik Brinkmann, Koyena Pal, Dmitrii Troitskii, Michael Ripa, Adam Belfki, Can Rager, Caden Juang, Aaron Mueller, Samuel Marks, Arnab Sen Sharma, Francesca Lucchetti, Nikhil Prakash, Carla E. Brodley, Arjun Guha, Jonathan Bell 0001, Byron C. Wallace, David Bau |
ICLR | 19 |
| 2025 | Do Automatic Factuality Metrics Measure Factuality? A Critical EvaluationabstractModern LLMs can now produce highly readable abstractive summaries, to the point that traditional automated metrics for evaluating summary quality, such as ROUGE, have saturated.
However, LLMs still sometimes introduce inaccuracies into summaries, i.e., information inconsistent with or unsupported by the corresponding source.
Measuring the occurrence of these often subtle factual inconsistencies automatically has proved challenging.
This in turn has motivated development of metrics intended to measure the factual consistency of generated summaries against sources.
But are these approaches measuring what they purport to? Or are they mostly exploiting artifacts?
In this work, we stress test a range of automatic factuality metrics—including specialized model-based approaches and LLM-based prompting methods—to probe what they actually capture. Using a shallow classifier to separate “easy” examples for factual evaluation—where surface features suffice—from “hard” cases requiring deeper reasoning, we find that all metrics show substantial performance drops on the latter.
Furthermore, some metrics are more sensitive to benign, fact-preserving edits than to factual corrections. Building on this observation, we demonstrate that most automatic factuality metrics can be gamed—that is, their scores can be artificially inflated by appending innocuous, content-free sentences to summaries. Among the metrics tested, the LLM prompt-based ChatGPT-DA approach is the most robust and reliable; however, it exhibits a notable caveat: it likely relies more on parametric knowledge than on the provided source when making judgments. Taken together, our findings call into question the reliability of current factuality metrics and prompt a broader reflection on what these metrics are truly measuring. We conclude with concrete recommendations for improving both benchmark design and metric robustness, particularly in light of their vulnerability to superficial manipulations. Sanjana Ramprasad, Byron C. Wallace |
NeurIPS | 2 |
| 2025 | Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language ModelsabstractFor an LLM to correctly respond to an instruction it must understand both the semantics and the domain (i.e., subject area) of a given task-instruction pair. However, syntax can also convey implicit information. Recent work shows that \textit{syntactic templates}---frequent sequences of Part-of-Speech (PoS) tags---are prevalent in training data and often appear in model outputs. In this work we characterize syntactic templates, domain, and semantics in task-instruction pairs. We identify cases of spurious correlations between syntax and domain, where models learn to associate a domain with syntax during training; this can sometimes override prompt semantics. Using a synthetic training dataset, we find that the syntactic-domain correlation can lower performance (mean 0.51 +/- 0.06) on entity knowledge tasks in OLMo-2 models (1B-13B). We introduce an evaluation framework to detect this phenomenon in trained models, and show that it occurs on a subset of the FlanV2 dataset in open (OLMo-2-7B; Llama-4-Maverick), and closed (GPT-4o) models. Finally, we present a case study on the implications for LLM security, showing that unintended syntactic-domain correlations can be used to bypass refusals in OLMo-2-7B Instruct and GPT-4o. Our findings highlight two needs: (1) to explicitly test for syntactic-domain correlations, and (2) to ensure \textit{syntactic} diversity in training data, specifically within domains, to prevent such spurious correlations. Chantal Shaib, Vinith M. Suriyakumar, Byron C. Wallace, Marzyeh Ghassemi |
NeurIPS | 3 |
| 2024 | FactPICO: Factuality Evaluation for Plain Language Summarization of Medical EvidenceabstractSebastian Joseph, Lily Chen, Jan Trienes, Hannah Göke, Monika Coers, Wei Xu, Byron Wallace, Junyi Jessy Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Sebastian Joseph, Lily Chen, Jan Trienes, Hannah Louisa Göke, Monika Coers, Wei Xu 0004, Byron C. Wallace, Junyi Jessy Li |
ACL (1) | 7 |
| 2024 | InfoLossQA: Characterizing and Recovering Information Loss in Text SimplificationabstractJan Trienes, Sebastian Joseph, Jörg Schlötterer, Christin Seifert, Kyle Lo, Wei Xu, Byron Wallace, Junyi Jessy Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Jan Trienes, Sebastian Joseph, Jörg Schlötterer, Christin Seifert, Kyle Lo, Wei Xu 0004, Byron C. Wallace, Junyi Jessy Li |
ACL (1) | 7 |
| 2024 | Token Erasure as a Footprint of Implicit Vocabulary Items in LLMsabstractLLMs process text as sequences of tokens that roughly correspond to words, where less common words are represented by multiple tokens.However, individual tokens are often semantically unrelated to the meanings of the words/concepts they comprise.For example, Llama-2-7b's tokenizer splits the word "northeastern" into the tokens [_n, ort, he, astern], none of which correspond to semantically meaningful units like "north" or "east."Similarly, the overall meanings of named entities like "Neil Young" and multi-word expressions like "break a leg" cannot be directly inferred from their constituent tokens.Mechanistically, how do LLMs convert such arbitrary groups of tokens into useful higher-level representations?In this work, we find that last token representations of named entities and multi-token words exhibit a pronounced "erasure" effect, where information about previous and current tokens is rapidly forgotten in early layers.Using this observation, we propose a method to "read out" the implicit vocabulary of an autoregressive LLM by examining differences in token representations across layers, and present results of this method for Llama-2-7b and Llama-3-8b.To our knowledge, this is the first attempt to probe the implicit vocabulary of an LLM. 1 Sheridan Feucht, David Atkinson, Byron C. Wallace, David Bau |
EMNLP | 3 |
| 2024 | Detection and Measurement of Syntactic Templates in Generated TextabstractThe diversity of text can be measured beyond word-level features, however existing diversity evaluation focuses primarily on word-level features.Here we propose a method for evaluating diversity over syntactic features to characterize general repetition in models, beyond frequent n-grams.Specifically, we define syntactic templates (e.g., strings comprising parts-of-speech) and show that models tend to produce templated text in downstream tasks at a higher rate than what is found in human-reference texts We find that most (76%) templates in modelgenerated text can be found in pre-training data (compared to only 35% of human-authored text), and are not overwritten during fine-tuning or alignment processes such as RLHF.The connection between templates in generated text and the pre-training data allows us to analyze syntactic templates in models where we do not have the pre-training data.We also find that templates as features are able to differentiate between models, tasks, and domains, and are useful for qualitatively evaluating common model constructions.Finally, we demonstrate the use of templates as a useful tool for analyzing style memorization of training data in LLMs 1 . Chantal Shaib, Yanai Elazar, Junyi Jessy Li, Byron C. Wallace |
EMNLP | 4 |
| 2024 | Investigating Mysteries of CoT-Augmented DistillationabstractEliciting chain of thought (CoT) rationales - sequences of token that convey a “reasoning” process has been shown to consistently improve LLM performance on tasks like question answering. More recent efforts have shown that such rationales can also be used for model distillation: Including CoT sequences (elicited from a large “teacher” model) in addition to target labels when fine-tuning a small student model yields (often substantial) improvements. In this work we ask: Why and how does this additional training signal help in model distillation? We perform ablations to interrogate this, and report some potentially surprising results. Specifically: (1) Placing CoT sequences after labels (rather than before) realizes consistently better downstream performance – this means that no student “reasoning” is necessary at test time to realize gains. (2) When rationales are appended in this way, they need not be coherent reasoning sequences to yield improvements; performance increases are robust to permutations of CoT tokens, for example. In fact, (3) a small number of key tokens are sufficient to achieve improvements equivalent to those observed when full rationales are used in model distillation. Somin Wadhwa, Silvio Amir, Byron C. Wallace |
EMNLP | 3 |
| 2024 | Learning from Natural Language Explanations for Generalizable Entity MatchingabstractEntity matching is the task of linking records from different sources that refer to the same real-world entity.Past work has primarily treated entity linking as a standard supervised learning problem.However, supervised entity matching models often do not generalize well to new data, and collecting exhaustive labeled training data is often cost prohibitive.Further, recent efforts have adopted LLMs for this task in few/zero-shot settings, exploiting their general knowledge.But LLMs are prohibitively expensive for performing inference at scale for real-world entity matching tasks.As an efficient alternative, we re-cast entity matching as a conditional generation task as opposed to binary classification.This enables us to "distill" LLM reasoning into smaller entity matching models via natural language explanations.This approach achieves strong performance, especially on out-of-domain generalization tests (↑10.85%F-1) where standalone generative methods struggle.We perform ablations that highlight the importance of explanations, both for performance and model robustness.Explain matching label class given the entity descriptions: Label: Match E_a: Nike Sportswear AF-1 488298-436 MN Navy.E_b: Air Force 1 [BRAND] Somin Wadhwa, Adit Krishnan, Runhui Wang, Byron C. Wallace, Luyang Kong |
EMNLP | 4 |
| 2024 | Evaluating the Zero-shot Robustness of Instruction-tuned Language ModelsabstractInstruction fine-tuning has recently emerged as a promising approach for improving the zero-shot capabilities of Large Language Models (LLMs) on new tasks. This technique has shown particular strength in improving the performance of modestly sized LLMs, sometimes inducing performance competitive with much larger model variants. In this paper, we ask two questions: (1) How sensitive are instruction-tuned models to the particular phrasings of instructions, and, (2) How can we make them more robust to such natural language variation? To answer the former, we collect a set of 319 instructions manually written by NLP practitioners for over 80 unique tasks included in widely used benchmarks, and we evaluate the variance and average performance of these instructions as compared to instruction phrasings observed during instruction fine-tuning. We find that using novel (unobserved) but appropriate instruction phrasings consistently degrades model performance, sometimes substantially so. Further, such natural instructions yield a wide variance in downstream performance, despite their semantic equivalence. Put another way, instruction-tuned models are not especially robust to instruction re-phrasings.
We propose a simple method to mitigate this issue by introducing ``soft prompt'' embedding parameters and optimizing these to maximize the similarity between representations of semantically equivalent instructions. We show that this method consistently improves the robustness of instruction-tuned models. Jiuding Sun, Chantal Shaib, Byron C. Wallace |
ICLR | 3 |
| 2024 | Function Vectors in Large Language ModelsabstractWe report the presence of a simple neural mechanism that represents an input-output function as a vector within autoregressive transformer language models (LMs). Using causal mediation analysis on a diverse range of in-context-learning (ICL) tasks, we find that a small number attention heads transport a compact representation of the demonstrated task, which we call a function vector (FV). FVs are robust to changes in context, i.e., they trigger execution of the task on inputs such as zero-shot and natural text settings that do not resemble the ICL contexts from which they are collected. We test FVs across a range of tasks, models, and layers and find strong causal effects across settings in middle layers. We investigate the internal structure of FVs and find while that they often contain information that encodes the output space of the function, this information alone is not sufficient to reconstruct an FV. Finally, we test semantic vector composition in FVs, and find that to some extent they can be summed to create vectors that trigger new complex tasks. Our findings show that compact, causal internal vector representations of function abstractions can be explicitly extracted from LLMs. Eric Todd, Millicent Li, Arnab Sen Sharma, Aaron Mueller, Byron C. Wallace, David Bau |
ICLR | 5 |
| 2024 | Towards Reducing Diagnostic Errors with Interpretable Risk Predictionabstracta confident diagnosis can be made. We use an LLM to retrieve an initial pool of evidence, but then refine this set of evidence according to correlations learned by the model. We conduct an in-depth evaluation of the usefulness of our approach by simulating how it might be used by a clinician to decide between a pre-defined list of differential diagnoses. Denis Jered McInerney, William Dickinson, Lucy C. Flynn, Andrea Young, Geoffrey S. Young, Jan-Willem van de Meent, Byron C. Wallace |
NAACL-HLT | 7 |
| 2024 | On-the-fly Definition Augmentation of LLMs for Biomedical NERabstractMonica Munnangi, Sergey Feldman, Byron Wallace, Silvio Amir, Tom Hope, Aakanksha Naik. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Monica Munnangi, Sergey Feldman, Byron C. Wallace, Silvio Amir, Tom Hope, Aakanksha Naik |
NAACL-HLT | 3 |
| 2024 | Question answering systems for health professionals at the point of care - a systematic reviewabstractOBJECTIVES: Question answering (QA) systems have the potential to improve the quality of clinical care by providing health professionals with the latest and most relevant evidence. However, QA systems have not been widely adopted. This systematic review aims to characterize current medical QA systems, assess their suitability for healthcare, and identify areas of improvement. MATERIALS AND METHODS: We searched PubMed, IEEE Xplore, ACM Digital Library, ACL Anthology, and forward and backward citations on February 7, 2023. We included peer-reviewed journal and conference papers describing the design and evaluation of biomedical QA systems. Two reviewers screened titles, abstracts, and full-text articles. We conducted a narrative synthesis and risk of bias assessment for each study. We assessed the utility of biomedical QA systems. RESULTS: We included 79 studies and identified themes, including question realism, answer reliability, answer utility, clinical specialism, systems, usability, and evaluation methods. Clinicians' questions used to train and evaluate QA systems were restricted to certain sources, types and complexity levels. No system communicated confidence levels in the answers or sources. Many studies suffered from high risks of bias and applicability concerns. Only 8 studies completely satisfied any criterion for clinical utility, and only 7 reported user evaluations. Most systems were built with limited input from clinicians. DISCUSSION: While machine learning methods have led to increased accuracy, most studies imperfectly reflected real-world healthcare information needs. Key research priorities include developing more realistic healthcare QA datasets and considering the reliability of answer sources, rather than merely focusing on accuracy. Gregory Kell, Angus Roberts, Serge Umansky, Linglong Qian, Frank Soboczenski, Byron C. Wallace, Nikhil Patel, Iain James Marshall |
J. Am. Medical Informatics Assoc. | 7 |
| 2024 | Leveraging generative AI for clinical evidence synthesis needs to ensure trustworthiness
Qiao Jin 0001, Denis Jered McInerney, Yong Chen 0016, Fei Wang 0001, Curtis L. Cole, Qian Yang 0004, Yanshan Wang, Bradley A. Malin, Mor Peleg, Byron C. Wallace, Zhiyong Lu, Chunhua Weng, Yifan Peng 0002 |
J. Biomed. Informatics | 11 |
| 2024 | Do Multi-Document Summarization Models Synthesize?abstractAbstract Multi-document summarization entails producing concise synopses of collections of inputs. For some applications, the synopsis should accurately synthesize inputs with respect to a key aspect, e.g., a synopsis of film reviews written about a particular movie should reflect the average critic consensus. As a more consequential example, narrative summaries that accompany biomedical systematic reviews of clinical trial results should accurately summarize the potentially conflicting results from individual trials. In this paper we ask: To what extent do modern multi-document summarization models implicitly perform this sort of synthesis? We run experiments over opinion and evidence synthesis datasets using a suite of summarization models, from fine-tuned transformers to GPT-4. We find that existing models partially perform synthesis, but imperfectly: Even the best performing models are over-sensitive to changes in input ordering and under-sensitive to changes in input compositions (e.g., ratio of positive to negative reviews). We propose a simple, general, effective method for improving model synthesis capabilities by generating an explicitly diverse set of candidate outputs, and then selecting from these the string best aligned with the expected aggregate measure for the inputs, or abstaining when the model produces no good candidate. Jay DeYoung, Stephanie C. Martinez, Iain James Marshall, Byron C. Wallace |
Trans. Assoc. Comput. Linguistics | 4 |
| 2023 | Revisiting Relation Extraction in the era of Large Language ModelsabstractRelation extraction (RE) is the core NLP task of inferring semantic relationships between entities from text.Standard supervised RE techniques entail training modules to tag tokens comprising entity spans and then predict the relationship between them.Recent work has instead treated the problem as a sequence-tosequence task, linearizing relations between entities as target strings to be generated conditioned on the input.Here we push the limits of this approach, using larger language models (GPT-3 and Flan-T5 large) than considered in prior work and evaluating their performance on standard RE tasks under varying levels of supervision.We address issues inherent to evaluating generative approaches to RE by doing human evaluations, in lieu of relying on exact matching.Under this refined evaluation, we find that: (1) Few-shot prompting with GPT-3 achieves near SOTA performance, i.e., roughly equivalent to existing fully supervised models; (2) Flan-T5 is not as capable in the fewshot setting, but supervising and fine-tuning it with Chain-of-Thought (CoT) style explanations (generated via GPT-3) yields SOTA results.We release this model as a new baseline for RE tasks 1 . Somin Wadhwa, Silvio Amir, Byron C. Wallace |
ACL (1) | 3 |
| 2023 | Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluationsabstract-gram similarity metrics such as ROUGE. Better automated evaluation metrics are needed, but few resources exist to assess metrics when they are proposed. Therefore, we introduce a dataset of human-assessed summary quality facets and pairwise preferences to encourage and support the development of better automated evaluation methods for literature review MDS. We take advantage of community submissions to the Multi-document Summarization for Literature Review (MSLR) shared task to compile a diverse and representative sample of generated summaries. We analyze how automated summarization evaluation metrics correlate with lexical features of generated summaries, to other automated metrics including several we propose in this work, and to aspects of human-assessed summary quality. We find that not only do automated metrics fail to capture aspects of quality as assessed by humans, in many cases the system rankings produced by these metrics are anti-correlated with rankings according to human annotators. Lucy Lu Wang, Yulia Otmakhova 0001, Jay DeYoung, Hung-Thinh Truong, Bailey Kuehl, Erin Bransom, Byron C. Wallace |
ACL (1) | 7 |
| 2023 | Future Lens: Anticipating Subsequent Tokens from a Single Hidden StateabstractWe conjecture that hidden state vectors corresponding to individual input tokens encode information sufficient to accurately predict several tokens ahead.More concretely, in this paper we ask: Given a hidden (internal) representation of a single token at position t in an input, can we reliably anticipate the tokens that will appear at positions ≥ t + 2? To test this, we measure linear approximation and causal intervention methods in GPT-J-6B to evaluate the degree to which individual hidden states in the network contain signal rich enough to predict future hidden states and, ultimately, token outputs.We find that, at some layers, we can approximate a model's output with more than 48% accuracy with respect to its prediction of subsequent tokens through a single hidden state.Finally we present a "Future Lens" visualization that uses these methods to create a new view of transformer states. Koyena Pal, Jiuding Sun, Andrew Yuan, Byron C. Wallace, David Bau |
CoNLL | 4 |
| 2023 | How Many and Which Training Points Would Need to be Removed to Flip this Prediction?abstractWe consider the problem of identifying a minimal subset of training data S t such that if the instances comprising S t had been removed prior to training, the categorization of a given test point x t would have been different.Identifying such a set may be of interest for a few reasons.First, the cardinality of S t provides a measure of robustness (if |S t | is small for x t , we might be less confident in the corresponding prediction), which we show is correlated with but complementary to predicted probabilities.Second, interrogation of S t may provide a novel mechanism for contesting a particular model prediction: If one can make the case that the points in S t are wrongly labeled or irrelevant, this may argue for overturning the associated prediction.Identifying S t via bruteforce is intractable.We propose comparatively fast approximation methods to find S t based on influence functions, and find that-for simple convex text classification models-these approaches can often successfully identify relatively small sets of training examples which, if removed, would flip the prediction. 1 Byron C. Wallace |
EACL | 3 |
| 2023 | Multilingual Simplification of Medical TextsabstractAutomated text simplification aims to produce simple versions of complex texts.This task is especially useful in the medical domain, where the latest medical findings are typically communicated via complex, technical articles.This creates barriers for laypeople seeking access to up-to-date medical findings, consequently impeding progress on health literacy.Most existing work on medical text simplification has focused on monolingual settings, with the result that such evidence would be available only in just one language (most often, English).This work addresses this limitation via multilingual simplification, i.e., directly simplifying complex texts into simplified texts in multiple languages.We introduce MULTICOCHRANE, the first sentence-aligned multilingual text simplification dataset for the medical domain in four languages: English, Spanish, French, and Farsi.We evaluate fine-tuned and zero-shot models across these languages with extensive human assessments and analyses.Although models can generate viable simplified texts, we identify several outstanding challenges that this dataset might be used to address. MultiCochrane Complex SimplePreclinical studies have suggested that RIC may have beneficial effects in ischaemic stroke patients and those at risk of ischaemic stroke.English: Studies have suggested that RIC may have beneficial effects for preventing and treating ischaemic stroke.Spanish: Los estudios han indicado que el CIR puede tener efectos beneficiosos en la prevención y el tratamiento del accidente cerebrovascular isquémico.French: Des études ont suggéré que le CID pourrait avoir des effets bénéfiques sur la prévention et le traitement de l'AVC ischémique.Farsi: ﮐﮫ اﻧد ﮐرده ﭘﯾﺷﻧﮭﺎد طﺎﻟﻌﺎت RIC درﻣﺎن و ﭘﯾﺷﮕﯾری ﺑرای ﻣﻔﯾدی اﺛرات اﺳت ﻣﻣﮑن ﺑﺎﺷد داﺷﺗﮫ اﯾﺳﮑﻣﯾﮏ ﻣﻐزی .ﺳﮑﺗﮫ Human evaluation of system outputsEnglish: Interventions have suggested that RAP may have beneficial effects in ischaemic stroke patients, those at risk of stroke.Spanish: Las intervenciones pueden ser efectivas para los pacientes que se accidente cerebrovascular isquémico y los que se encuentran en riesgo del accidente cerebrovascular isquémico.Gloss: The interventions can be effective for the patients that accident themselves ischemic stroke and those that find themselves at risk of the ischemic stroke. Sebastian Joseph, Kathryn Kazanas, Keziah Reina, Vishnesh J. Ramanathan, Wei Xu 0004, Byron C. Wallace, Junyi Jessy Li |
EMNLP | 6 |
| 2023 | Appraising the Potential Uses and Harms of LLMs for Medical Systematic ReviewsabstractMedical systematic reviews play a vital role in healthcare decision making and policy.However, their production is time-consuming, limiting the availability of high-quality and up-todate evidence summaries.Recent advances in large language models (LLMs) offer the potential to automatically generate literature reviews on demand, addressing this issue.However, LLMs sometimes generate inaccurate (and potentially misleading) texts by "hallucination" or omission.In healthcare, this can make LLMs unusable at best and dangerous at worst.We conducted 16 interviews with international systematic review experts to characterize the perceived utility and risks of LLMs in the specific context of medical evidence reviews.Experts indicated that LLMs can assist in the writing process by drafting summaries, generating templates, distilling information, and crosschecking information.But they also raised concerns regarding confidently composed but inaccurate LLM outputs and other potential downstream harms, including decreased accountability and proliferation of low-quality reviews.Informed by this qualitative analysis, we identify criteria for rigorous evaluation of biomedical LLMs aligned with domain expert views. Hye Sun Yun, Iain James Marshall, Thomas A. Trikalinos, Byron C. Wallace |
EMNLP | 4 |
| 2023 | Accomodating User Expressivity while Maintaining Safety for a Virtual Alcohol Misuse CounselorabstractClient-centered counseling, in which individuals are prompted to talk about their behavior, is the standard treatment for Alcohol misuse. However, open-ended conversations with virtual agent counselor raise potential safety concerns if the agent misunderstands and provides erroneous advice. Thus, while generative machine learning models have been successful for language understanding and generation tasks, these approaches may not be effective or safe for counseling. We present a hybrid dialog system that uses a machine-learning model to generate responses to individual client speech combined with a rule-based approach to transition through structured counseling sessions. The dialog system is used to drive a virtual agent alcohol misuse counselor. We evaluated this hybrid system by comparing it to a functionally equivalent system in which the dialog is driven by fully-constraining user utterances via multiple-choice menus among individuals with problematic drinking. Participants who interacted with the agent using the hybrid dialog system reported a higher degree of readiness to change their drinking habits compared to the system using constrained input. Additionally, the outputs of the models used by the system were judged to be safe and appropriate for the task by expert counselors. Stefan Olafsson, Paola Pedrelli, Byron C. Wallace, Timothy W. Bickmore |
IVA | 3 |
| 2022 | Evaluating Factuality in Text Simplificationabstractmodels aim to make input texts more readable. Such methods have the potential to make complex information accessible to a wider audience, e.g., providing access to recent medical literature which might otherwise be impenetrable for a lay reader. However, such models risk introducing errors into automatically simplified texts, for instance by inserting statements unsupported by the corresponding original text, or by omitting key information. Providing more readable but inaccurate versions of texts may in many cases be worse than providing no such access at all. The problem of factual accuracy (and the lack thereof) has received heightened attention in the context of summarization models, but the factuality of automatically simplified texts has not been investigated. We introduce a taxonomy of errors that we use to analyze both references drawn from standard simplification datasets and state-of-the-art model outputs. We find that errors often appear in both that are not captured by existing evaluation metrics, motivating a need for research into ensuring the factual accuracy of automated simplification models. Ashwin Devaraj, William Sheffield, Byron C. Wallace, Junyi Jessy Li |
ACL (1) | 3 |
| 2022 | That's the Wrong Lung! Evaluating and Improving the Interpretability of Unsupervised Multimodal Encoders for Medical DataabstractPretraining multimodal models on Electronic Health Records (EHRs) provides a means of learning representations that can transfer to downstream tasks with minimal supervision. Recent multimodal models induce soft local alignments between image regions and sentences. This is of particular interest in the medical domain, where alignments might highlight regions in an image relevant to specific phenomena described in free-text. While past work has suggested that attention "heatmaps" can be interpreted in this manner, there has been little evaluation of such alignments. We compare alignments from a state-of-the-art multimodal (image and text) model for EHR with human annotations that link image regions to sentences. Our main finding is that the text has an often weak or unintuitive influence on attention; alignments do not consistently reflect basic anatomical information. Moreover, synthetic modifications - such as substituting "left" for "right" - do not substantially influence highlights. Simple techniques such as allowing the model to opt out of attending to the image and few-shot finetuning show promise in terms of their ability to improve alignments with very little or no supervision. We make our code and checkpoints open-source. Denis Jered McInerney, Geoffrey S. Young, Jan-Willem van de Meent, Byron C. Wallace |
EMNLP | 4 |
| 2022 | PHEE: A Dataset for Pharmacovigilance Event Extraction from TextabstractThe primary goal of drug safety researchers and regulators is to promptly identify adverse drug reactions.Doing so may in turn prevent or reduce the harm to patients and ultimately improve public health.Evaluating and monitoring drug safety (i.e., pharmacovigilance) involves analyzing an ever growing collection of spontaneous reports from health professionals, physicians, and pharmacists, and information voluntarily submitted by patients.In this scenario, facilitating analysis of such reports via automation has the potential to rapidly identify safety signals.Unfortunately, public resources for developing natural language models for this task are scant.We present PHEE, a novel dataset for pharmacovigilance comprising over 5000 annotated events from medical case reports and biomedical literature, making it the largest such public dataset to date.We describe the hierarchical event schema designed to provide coarse and fine-grained information about patients' demographics, treatments and (side) effects.Along with the discussion of the dataset, we present a thorough experimental evaluation of current state-of-the-art approaches for biomedical event extraction, point out their limitations, and highlight open challenges to foster future research in this area 1 . Zhaoyue Sun, Jiazheng Li 0002, Gabriele Pergola, Byron C. Wallace, Bino John, Nigel Greene, Joseph Kim, Yulan He 0001 |
EMNLP | 4 |
| 2021 | Identifying Communication Behavior Indicators in Secure Messages
Dezon Finch, Lina Bouayad, Timothy P. Hogan, Sarah L. Cutrona, Byron C. Wallace, Stephen Luther, Bridget Smith, Stephanie L. Shimada |
AMIA | 5 |
| 2021 | Applying State of the Art Language Models to Enable Better Clinical Natural Language Processing
Bryan D. Steitz, Emily Alsentzer, Hoo Chang Shin, Byron C. Wallace, Adam Wright |
AMIA | 4 |
| 2021 | Unsupervised Data Augmentation with Naive Augmentation and without Unlabeled DataabstractUnsupervised Data Augmentation (UDA) is a semi-supervised technique that applies a consistency loss to penalize differences between a model's predictions on (a) observed (unlabeled) examples; and (b) corresponding 'noised' examples produced via data augmentation.While UDA has gained popularity for text classification, open questions linger over which of its components are important, and how to extend the method to sequence labeling tasks; this paper addresses these questions.Our main contribution is an empirical study of UDA to establish which components of the algorithm confer benefits in NLP.Notably, although prior work has emphasized use of clever augmentation techniques including back-translation, we find that enforcing consistency between predictions assigned to observed and randomly substituted words often yields comparable (or greater) benefits compared to these more complex perturbation models.Furthermore, we find that applying UDA's consistency loss affords meaningful gains without any unlabeled data at all, i.e., in a standard supervised setting.In short, UDA need not be unsupervised to realize much of its noted benefits, and does not require complex data augmentation to be effective. David Lowell, Brian E. Howard, Zachary C. Lipton, Byron C. Wallace |
EMNLP (1) | 4 |
| 2021 | Disentangling Representations of Text by Masking TransformersabstractRepresentations from large pretrained models such as BERT encode a range of features into monolithic vectors, affording strong predictive accuracy across a range of downstream tasks.In this paper we explore whether it is possible to learn disentangled representations by identifying existing subnetworks within pretrained models that encode distinct, complementary aspects.Concretely, we learn binary masks over transformer weights or hidden units to uncover subsets of features that correlate with a specific factor of variation; this eliminates the need to train a disentangled model from scratch for a particular task.We evaluate this method with respect to its ability to disentangle representations of sentiment from genre in movie reviews, toxicity from dialect in Tweets, and syntax from semantics.By combining masking with magnitude pruning we find that we can identify sparse subnetworks within BERT that strongly encode particular aspects (e.g., semantics) while only weakly encoding others (e.g., syntax).Moreover, despite only learning masks, disentanglement-via-masking performs as well as -and often better thanpreviously proposed methods based on variational autoencoders and adversarial training. Xiongyi Zhang, Jan-Willem van de Meent, Byron C. Wallace |
EMNLP (1) | 3 |
| 2021 | On the Impact of Random Seeds on the Fairness of Clinical ClassifiersabstractSilvio Amir, Jan-Willem van de Meent, Byron Wallace. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Silvio Amir, Jan-Willem van de Meent, Byron C. Wallace |
NAACL-HLT | 3 |
| 2021 | Paragraph-level Simplification of Medical TextsabstractWe consider the problem of learning to simplify medical texts. This is important because most reliable, up-to-date information in biomedicine is dense with jargon and thus practically inaccessible to the lay audience. Furthermore, manual simplification does not scale to the rapidly growing body of biomedical literature, motivating the need for automated approaches. Unfortunately, there are no large-scale resources available for this task. In this work we introduce a new corpus of parallel texts in English comprising technical and lay summaries of all published evidence pertaining to different clinical topics. We then propose a new metric based on likelihood scores from a masked language model pretrained on scientific texts. We show that this automated measure better differentiates between technical and lay summaries than existing heuristics. We introduce and evaluate baseline encoder-decoder Transformer models for simplification and propose a novel augmentation to these in which we explicitly penalize the decoder for producing 'jargon' terms; we find that this yields improvements over baselines in terms of readability. Ashwin Devaraj, Iain James Marshall, Byron C. Wallace, Junyi Jessy Li |
NAACL-HLT | 3 |
| 2021 | Does BERT Pretrained on Clinical Notes Reveal Sensitive Data?abstractEric Lehman, Sarthak Jain, Karl Pichotta, Yoav Goldberg, Byron Wallace. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Eric P. Lehman, Karl Pichotta, Yoav Goldberg, Byron C. Wallace |
NAACL-HLT | 5 |
| 2021 | An Empirical Comparison of Instance Attribution Methods for NLPabstractPouya Pezeshkpour, Sarthak Jain, Byron Wallace, Sameer Singh. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Pouya Pezeshkpour, Byron C. Wallace, Sameer Singh 0001 |
NAACL-HLT | 3 |
| 2021 | Interpretability Analysis for Named Entity Recognition to Understand System Predictions and How They Can ImproveabstractAbstract Named entity recognition systems achieve remarkable performance on domains such as English news. It is natural to ask: What are these models actually learning to achieve this? Are they merely memorizing the names themselves? Or are they capable of interpreting the text and inferring the correct entity type from the linguistic context? We examine these questions by contrasting the performance of several variants of architectures for named entity recognition, with some provided only representations of the context as features. We experiment with GloVe-based BiLSTM-CRF as well as BERT. We find that context does influence predictions, but the main factor driving high performance is learning the named tokens themselves. Furthermore, we find that BERT is not always better at recognizing predictive contexts compared to a BiLSTM-CRF model. We enlist human annotators to evaluate the feasibility of inferring entity types from context alone and find that humans are also mostly unable to infer entity types for the majority of examples on which the context-only system made errors. However, there is room for improvement: A system should be able to recognize any named entity in a predictive context correctly and our experiments indicate that current systems may be improved by such capability. Our human study also revealed that systems and humans do not always learn the same contextual clues, and context-only systems are sometimes correct even when humans fail to recognize the entity type from the context. Finally, we find that one issue contributing to model errors is the use of “entangled” representations that encode both contextual and local token information into a single vector, which can obscure clues. Our results suggest that designing models that explicitly operate over representations of local inputs and context, respectively, may in some cases improve performance. In light of these and related findings, we highlight directions for future work. Oshin Agarwal, Yinfei Yang, Byron C. Wallace, Ani Nenkova |
Comput. Linguistics | 3 |
| 2020 | ERASER: A Benchmark to Evaluate Rationalized NLP ModelsabstractJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, Byron C. Wallace. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Jay DeYoung, Nazneen Fatema Rajani, Eric P. Lehman, Caiming Xiong, Richard Socher, Byron C. Wallace |
ACL | 7 |
| 2020 | Explaining Black Box Predictions and Unveiling Data Artifacts through Influence FunctionsabstractModern deep learning models for NLP are notoriously opaque.This has motivated the development of methods for interpreting such models, e.g., via gradient-based saliency maps or the visualization of attention weights.Such approaches aim to provide explanations for a particular model prediction by highlighting important words in the corresponding input text.While this might be useful for tasks where decisions are explicitly influenced by individual tokens in the input, we suspect that such highlighting is not always suitable for tasks where model decisions should be driven by more complex reasoning.In this work, we investigate the use of influence functions for NLP, providing an alternative approach to interpreting neural text classifiers.Influence functions explain the decisions of a model by identifying influential training examples.Despite the promise of this approach, influence functions have not yet been extensively evaluated in the context of NLP, a gap addressed by this work.We conduct a comparison between influence functions and common word-saliency methods on representative tasks.As suspected, we find that influence functions are particularly useful for natural language inference, a task in which 'saliency maps' may not provide clear interpretation.Furthermore, we develop a new quantitative measure based on influence functions that can reveal artifacts in training data. Xiaochuang Han, Byron C. Wallace, Yulia Tsvetkov |
ACL | 2 |
| 2020 | Learning to Faithfully Rationalize by ConstructionabstractIn many settings it is important for one to be able to understand why a model made a particular prediction. In NLP this often entails extracting snippets of an input text ‘responsible for’ corresponding model output; when such a snippet comprises tokens that indeed informed the model’s prediction, it is a faithful explanation. In some settings, faithfulness may be critical to ensure transparency. Lei et al. (2016) proposed a model to produce faithful rationales for neural text classification by defining independent snippet extraction and prediction modules. However, the discrete selection over input tokens performed by this method complicates training, leading to high variance and requiring careful hyperparameter tuning. We propose a simpler variant of this approach that provides faithful explanations by construction. In our scheme, named FRESH, arbitrary feature importance scores (e.g., gradients from a trained model) are used to induce binary labels over token inputs, which an extractor can be trained to predict. An independent classifier module is then trained exclusively on snippets provided by the extractor; these snippets thus constitute faithful explanations, even if the classifier is arbitrarily complex. In both automatic and manual evaluations we find that variants of this simple framework yield predictive performance superior to ‘end-to-end’ approaches, while being more general and easier to train. Code is available at https://github.com/successar/FRESH. Sarah Wiegreffe, Yuval Pinter, Byron C. Wallace |
ACL | 4 |
| 2020 | Trialstreamer: A living, automatically updated database of clinical trial reportsabstractOBJECTIVE: Randomized controlled trials (RCTs) are the gold standard method for evaluating whether a treatment works in health care but can be difficult to find and make use of. We describe the development and evaluation of a system to automatically find and categorize all new RCT reports. MATERIALS AND METHODS: Trialstreamer continuously monitors PubMed and the World Health Organization International Clinical Trials Registry Platform, looking for new RCTs in humans using a validated classifier. We combine machine learning and rule-based methods to extract information from the RCT abstracts, including free-text descriptions of trial PICO (populations, interventions/comparators, and outcomes) elements and map these snippets to normalized MeSH (Medical Subject Headings) vocabulary terms. We additionally identify sample sizes, predict the risk of bias, and extract text conveying key findings. We store all extracted data in a database, which we make freely available for download, and via a search portal, which allows users to enter structured clinical queries. Results are ranked automatically to prioritize larger and higher-quality studies. RESULTS: As of early June 2020, we have indexed 673 191 publications of RCTs, of which 22 363 were published in the first 5 months of 2020 (142 per day). We additionally include 304 111 trial registrations from the International Clinical Trials Registry Platform. The median trial sample size was 66. CONCLUSIONS: We present an automated system for finding and categorizing RCTs. This yields a novel resource: a database of structured information automatically extracted for all published RCTs in humans. We make daily updates of this database available on our website (https://trialstreamer.robotreviewer.net). Iain James Marshall, Benjamin E. Nye, Joël Kuiper, Anna Noel-Storr, Rachel Marshall, Rory Maclean, Frank Soboczenski, Ani Nenkova, James Thomas 0001, Byron C. Wallace |
J. Am. Medical Informatics Assoc. | 10 |
| 2019 | Structured Neural Topic Models for ReviewsabstractWe present Variational Aspect-based Latent Topic Allocation (VALTA), a family of autoencoding topic models that learn aspect-based representations of reviews. VALTA defines a user-item encoder that maps bag-of-words vectors for combined reviews associated with each paired user and item onto structured embeddings, which in turn define per-aspect topic weights. We model individual reviews in a structured manner by inferring an aspect assignment for each sentence in a given review, where the per-aspect topic weights obtained by the user-item encoder serve to define a mixture over topics, conditioned on the aspect. The result is an autoencoding neural topic model for reviews, which can be trained in a fully unsupervised manner to learn topics that are structured into aspects. Experimental evaluation on large number of datasets demonstrates that aspects are interpretable, yield higher coherence scores than non-structured autoencoding topic model variants, and can be utilized to perform aspect-based comparison and genre discovery. Babak Esmaeili 0001, Hongyi Huang, Byron C. Wallace, Jan-Willem van de Meent |
AISTATS | 3 |
| 2019 | Practical Obstacles to Deploying Active LearningabstractDavid Lowell, Zachary C. Lipton, Byron C. Wallace. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. David Lowell, Zachary C. Lipton, Byron C. Wallace |
EMNLP/IJCNLP (1) | 3 |
| 2019 | What Does the Evidence Say? Models to Help Make Sense of the Biomedical LiteratureabstractIdeally decisions regarding medical treatments would be informed by the totality of the available evidence. The best evidence we currently have is in published natural language articles describing the conduct and results of clinical trials. Because these are unstructured, it is difficult for domain experts (e.g., physicians) to sort through and appraise the evidence pertaining to a given clinical question. Natural language technologies have the potential to improve access to the evidence via semi-automated processing of the biomedical literature. In this brief paper I highlight work on developing tasks, corpora, and models to support semi-automated evidence retrieval and extraction. The aim is to design models that can consume articles describing clinical trials and automatically extract from these key clinical variables and findings, and estimate their reliability. Completely automating 'machine reading' of evidence remains a distant aim given current technologies; the more immediate hope is to use such technologies to help domain experts access and make sense of unstructured biomedical evidence more efficiently, with the ultimate aim of improving patient care. Aside from their practical importance, these tasks pose core NLP challenges that directly motivate methodological innovation. Byron C. Wallace |
IJCAI | 1 |
| 2019 | Explainable modeling of annotations in crowdsourcingabstractAggregation models for improving the quality of annotations collected via crowdsourcing have been widely studied, but far less has been done to explain why annotators make the mistakes that they do. To this end, we propose a joint aggregation and worker clustering model that detects patterns underlying crowd worker labels to characterize varieties of labeling errors. We evaluate our approach on a Named Entity Recognition dataset labeled by Mechanical Turk workers in both a retrospective experiment and a small human study. The former shows that our joint model improves the quality of clusters vs. aggregation followed by clustering. Results of the latter suggest that clusters aid human sense-making in interpreting worker labels and predicting worker mistakes. By enabling better explanation of annotator mistakes, our model creates a new opportunity to help Requesters improve task instructions and to help crowd annotators learn from their mistakes. Source code, data, and supplementary material is shared online. An T. Nguyen 0001, Matthew Lease, Byron C. Wallace |
IUI | 3 |
| 2018 | An Interpretable Joint Graphical Model for Fact-Checking From CrowdsabstractAssessing the veracity of claims made on the Internet is an important, challenging, and timely problem. While automated fact-checking models have potential to help people better assess what they read, we argue such models must be explainable, accurate, and fast to be useful in practice; while prediction accuracy is clearly important, model transparency is critical in order for users to trust the system and integrate their own knowledge with model predictions. To achieve this, we propose a novel probabilistic graphical model (PGM) which combines machine learning with crowd annotations. Nodes in our model correspond to claim veracity, article stance regarding claims, reputation of news sources, and annotator reliabilities. We introduce a fast variational method for parameter estimation. Evaluation across two real-world datasets and three scenarios shows that: (1) joint modeling of sources, claims and crowd annotators in a PGM improves the predictive performance and interpretability for predicting claim veracity; and (2) our variational inference method achieves scalably fast parameter estimation, with only modest degradation in performance compared to Gibbs sampling. Regarding model transparency, we designed and deployed a prototype fact-checker Web tool, including a visual interface for explaining model predictions. Results of a small user study indicate that model explanations improve user satisfaction and trust in model predictions. We share our web demo, model source code, and the 13K crowd labels we collected. An T. Nguyen 0001, Aditya Kharosekar, Matthew Lease, Byron C. Wallace |
AAAI | 4 |
| 2018 | A Corpus with Multi-Level Annotations of Patients, Interventions and Outcomes to Support Language Processing for Medical LiteratureabstractWe present a corpus of 5,000 richly annotated abstracts of medical articles describing clinical randomized controlled trials. Annotations include demarcations of text spans that describe the Patient population enrolled, the Interventions studied and to what they were Compared, and the Outcomes measured (the 'PICO' elements). These spans are further annotated at a more granular level, e.g., individual interventions within them are marked and mapped onto a structured medical vocabulary. We acquired annotations from a diverse set of workers with varying levels of expertise and cost. We describe our data collection process and the corpus itself in detail. We then outline a set of challenging NLP tasks that would aid searching of the medical literature and the practice of evidence-based medicine. Benjamin E. Nye, Junyi Jessy Li, Roma Patel, Yinfei Yang, Iain James Marshall, Ani Nenkova, Byron C. Wallace |
ACL (1) | 7 |
| 2018 | Learning Disentangled Representations of Texts with Application to Biomedical AbstractsabstractWe propose a method for learning disentangled representations of texts that code for distinct and complementary aspects, with the aim of affording efficient model transfer and interpretability.To induce disentangled embeddings, we propose an adversarial objective based on the (dis)similarity between triplets of documents with respect to specific aspects.Our motivating application is embedding biomedical abstracts describing clinical trials in a manner that disentangles the populations, interventions, and outcomes in a given trial.We show that our method learns representations that encode these clinically salient aspects, and that these can be effectively used to perform aspect-specific retrieval.We demonstrate that the approach generalizes beyond our motivating application in experiments on two multi-aspect review corpora. Edward Banner, Jan-Willem van de Meent, Iain James Marshall, Byron C. Wallace |
EMNLP | 5 |
| 2018 | Structured Multi-Label Biomedical Text Tagging via Attentive Neural Tree DecodingabstractWe propose a model for tagging unstructured texts with an arbitrary number of terms drawn from a tree-structured vocabulary (i.e., an ontology).We treat this as a special case of sequence-to-sequence learning in which the decoder begins at the root node of an ontological tree and recursively elects to expand child nodes as a function of the input text, the current node, and the latent decoder state.In our experiments the proposed method outperforms state-of-the-art approaches on the important task of automatically assigning MeSH terms to biomedical abstracts. Gaurav Singh 0001, James Thomas 0001, Iain James Marshall, John Shawe-Taylor, Byron C. Wallace |
EMNLP | 5 |
| 2018 | Believe it or not: Designing a Human-AI Partnership for Mixed-Initiative Fact-CheckingabstractFact-checking, the task of assessing the veracity of claims, is an important, timely, and challenging problem. While many automated fact-checking systems have been recently proposed, the human side of the partnership has been largely neglected: how might people understand, interact with, and establish trust with an AI fact-checking system? Does such a system actually help people better assess the factuality of claims? In this paper, we present the design and evaluation of a mixed-initiative approach to fact-checking, blending human knowledge and experience with the efficiency and scalability of automated information retrieval and ML. In a user study in which participants used our system to aid their own assessment of claims, our results suggest that individuals tend to trust the system: participant accuracy assessing claims improved when exposed to correct model predictions. However, this trust perhaps goes too far: when the model was wrong, exposure to its predictions often degraded human accuracy. Participants given the option to interact with these incorrect predictions were often able improve their own performance. This suggests that transparent models are key to facilitating effective human interaction with fallible AI models. An T. Nguyen 0001, Aditya Kharosekar, Saumyaa Krishnan, Siddhesh Krishnan, Elizabeth Tate, Byron C. Wallace, Matthew Lease |
UIST | 6 |
| 2018 | Neural information retrieval: at the end of the early yearsabstractA recent “third wave” of neural network (NN) approaches now delivers state-of-the-art performance in many machine learning tasks, spanning speech recognition, computer vision, and natural language processing. Because these modern NNs often comprise multiple interconnected layers, work in this area is often referred to as deep learning . Recent years have witnessed an explosive growth of research into NN-based approaches to information retrieval (IR). A significant body of work has now been created. In this paper, we survey the current landscape of Neural IR research, paying special attention to the use of learned distributed representations of textual units. We highlight the successes of neural IR thus far, catalog obstacles to its wider adoption, and suggest potentially promising directions for future research. Kezban Dilek Onal, Ismail Sengör Altingövde, Pinar Karagöz, Alexander Braylan, Brandon Dang, Heng-Lu Chang, Henna Kim, Quinten McNamara, Aaron Angert, Edward Banner, Vivek Khetan, Tyler McDonnell, An T. Nguyen 0001, Byron C. Wallace, Maarten de Rijke, Matthew Lease |
Inf. Retr. J. | 17 |
| 2017 | Active Discriminative Text Representation LearningabstractWe propose a new active learning (AL) method for text classification with convolutional neural networks (CNNs). In AL, one selects the instances to be manually labeled with the aim of maximizing model performance with minimal effort. Neural models capitalize on word embeddings as representations (features), tuning these to the task at hand. We argue that AL strategies for multi-layered neural models should focus on selecting instances that most affect the embedding space (i.e., induce discriminative word representations). This is in contrast to traditional AL approaches (e.g., entropy-based uncertainty sampling), which specify higher level objectives. We propose a simple approach for sentence classification that selects instances containing words whose embeddings are likely to be updated with the greatest magnitude, thereby rapidly learning discriminative, task-specific embeddings. We extend this approach to document classification by jointly considering: (1) the expected changes to the constituent word representations; and (2) the model’s current overall uncertainty regarding the instance. The relative emphasis placed on these criteria is governed by a stochastic process that favors selecting instances likely to improve representations at the outset of learning, and then shifts toward general uncertainty sampling as AL progresses. Empirical results show that our method outperforms baseline AL approaches on both sentence and document classification tasks. We also show that, as expected, the method quickly learns discriminative word embeddings. To the best of our knowledge, this is the first work on AL addressing neural models for text classification. Matthew Lease, Byron C. Wallace |
AAAI | 3 |
| 2017 | Aggregating and Predicting Sequence Labels from Crowd AnnotationsabstractDespite sequences being core to NLP, scant work has considered how to handle noisy sequence labels from multiple annotators for the same text. Given such annotations, we consider two complementary tasks: (1) aggregating sequential crowd labels to infer a best single set of consensus annotations; and (2) using crowd annotations as training data for a model that can predict sequences in unannotated text. For aggregation, we propose a novel Hidden Markov Model variant. To predict sequences in unannotated text, we propose a neural approach using Long Short Term Memory. We evaluate a suite of methods across two different applications and text genres: Named-Entity Recognition in news articles and Information Extraction from biomedical abstracts. Results show improvement over strong baselines. Our source code and data are available online. An T. Nguyen 0001, Byron C. Wallace, Junyi Jessy Li, Ani Nenkova, Matthew Lease |
ACL (1) | 2 |
| 2017 | A Neural Candidate-Selector Architecture for Automatic Structured Clinical Text AnnotationabstractWe consider the task of automatically annotating free texts describing clinical trials with concepts from a controlled, structured medical vocabulary. Specifically, we aim to build a model to infer distinct sets of (ontological) concepts describing complementary clinically salient aspects of the underlying trials: the populations enrolled, the interventions administered and the outcomes measured, i.e., the PICO elements. This important practical problem poses a few key challenges. One issue is that the output space is vast, because the vocabulary comprises many unique concepts. Compounding this problem, annotated data in this domain is expensive to collect and hence sparse. Furthermore, the outputs (sets of concepts for each PICO element) are correlated: specific populations (e.g., diabetics) will render certain intervention concepts likely (insulin therapy) while effectively precluding others (radiation therapy). Such correlations should be exploited. We propose a novel neural model that addresses these challenges. We introduce a Candidate-Selector architecture in which the model considers setes of candidate concepts for PICO elements, and assesses their plausibility conditioned on the input text to be annotated. This relies on a 'candidate set' generator, which may be learned or relies on heuristics. A conditional discriminative neural model then jointly selects candidate concepts, given the input text. We compare the predictive performance of our approach to strong baselines, and show that it outperforms them. Finally, we perform a qualitative evaluation of the generated annotations by asking domain experts to assess their quality. Gaurav Singh 0001, Iain James Marshall, James Thomas 0001, John Shawe-Taylor, Byron C. Wallace |
CIKM | 5 |
| 2017 | A Sensitivity Analysis of (and Practitioners' Guide to) Convolutional Neural Networks for Sentence ClassificationabstractConvolutional Neural Networks (CNNs) have recently achieved remarkably strong performance on the practically important task of sentence classification (Kim, 2014; Kalchbrenner et al., 2014; Johnson and Zhang, 2014; Zhang et al., 2016). However, these models require practitioners to specify an exact model architecture and set accompanying hyperparameters, including the filter region size, regularization parameters, and so on. It is currently unknown how sensitive model performance is to changes in these configurations for the task of sentence classification. We thus conduct a sensitivity analysis of one-layer CNNs to explore the effect of architecture components on model performance; our aim is to distinguish between important and comparatively inconsequential design decisions for sentence classification. We focus on one-layer CNNs (to the exclusion of more complex models) due to their comparative simplicity and strong empirical performance, which makes it a modern standard baseline method akin to Support Vector Machine (SVMs) and logistic regression. We derive practical advice from our extensive empirical results for those interested in getting the most out of CNNs for sentence classification in real world settings. Byron C. Wallace |
IJCNLP(1) | 2 |
| 2017 | Identifying reports of randomized controlled trials (RCTs) via a hybrid machine learning and crowdsourcing approachabstractOBJECTIVES: Identifying all published reports of randomized controlled trials (RCTs) is an important aim, but it requires extensive manual effort to separate RCTs from non-RCTs, even using current machine learning (ML) approaches. We aimed to make this process more efficient via a hybrid approach using both crowdsourcing and ML. METHODS: We trained a classifier to discriminate between citations that describe RCTs and those that do not. We then adopted a simple strategy of automatically excluding citations deemed very unlikely to be RCTs by the classifier and deferring to crowdworkers otherwise. RESULTS: Combining ML and crowdsourcing provides a highly sensitive RCT identification strategy (our estimates suggest 95%-99% recall) with substantially less effort (we observed a reduction of around 60%-80%) than relying on manual screening alone. CONCLUSIONS: Hybrid crowd-ML strategies warrant further exploration for biomedical curation/annotation tasks. Byron C. Wallace, Anna Noel-Storr, Iain James Marshall, Aaron M. Cohen, Neil R. Smalheiser, James Thomas 0001 |
J. Am. Medical Informatics Assoc. | 1 |
| 2016 | Modelling Context with User Embeddings for Sarcasm Detection in Social MediaabstractWe introduce a deep neural network for automated sarcasm detection.Recent work has emphasized the need for models to capitalize on contextual features, beyond lexical and syntactic cues present in utterances.For example, different speakers will tend to employ sarcasm regarding different subjects and, thus, sarcasm detection models ought to encode such speaker information.Current methods have achieved this by way of laborious feature engineering.By contrast, we propose to automatically learn and then exploit user embeddings, to be used in concert with lexical signals to recognize sarcasm.Our approach does not require elaborate feature engineering (and concomitant data scraping); fitting user embeddings requires only the text from their previous posts.The experimental results show that the our model outperforms a state-of-the-art approach leveraging an extensive set of carefully crafted features. Silvio Amir, Byron C. Wallace, Paula Carvalho 0001, Mário J. Silva |
CoNLL | 2 |
| 2016 | Rationale-Augmented Convolutional Neural Networks for Text ClassificationabstractWe present a new Convolutional Neural Network (CNN) model for text classification that jointly exploits labels on documents and their constituent sentences.Specifically, we consider scenarios in which annotators explicitly mark sentences (or snippets) that support their overall document categorization, i.e., they provide rationales.Our model exploits such supervision via a hierarchical approach in which each document is represented by a linear combination of the vector representations of its component sentences.We propose a sentence-level convolutional model that estimates the probability that a given sentence is a rationale, and we then scale the contribution of each sentence to the aggregate document representation in proportion to these estimates.Experiments on five classification datasets that have document labels and associated rationales demonstrate that our approach consistently outperforms strong baselines.Moreover, our model naturally provides explanations for its predictions. Iain James Marshall, Byron C. Wallace |
EMNLP | 3 |
| 2016 | Probabilistic Modeling for Crowdsourcing Partially-Subjective RatingsabstractWhile many methods have been proposed to ensure data quality for objective tasks (in which a single correct response is presumed to exist for each item), estimating data quality with subjective tasks remains largely unexplored. Consider the popular task of collecting instance ratings from human judges: while agreement tends be high for instances having extremely good or bad properties, instances with more middling properties naturally elicit a wider variance in opinion. In addition, because such subjectivity permits a valid diversity of responses, it can be difficult to detect if a judge does not undertake the task in good faith. To address this, we propose a probabilistic, heteroskedastic model in which the means and variances of worker responses are modeled as functions of instance attributes. We derive efficient Expectation Maximization (EM) learning and variational inference algorithms for parameter estimation. We apply our model to a large dataset of 24,132 Mechanical Turk ratings of user experience in viewing videos on smartphones with varying hardware capabilities. Results show that our method is effective at both predicting user ratings and in detecting unreliable respondents. An T. Nguyen 0001, Matthew Halpern, Byron C. Wallace, Matthew Lease |
HCOMP | 3 |
| 2016 | MGNC-CNN: A Simple Approach to Exploiting Multiple Word Embeddings for Sentence ClassificationabstractWe introduce a novel, simple convolution neural network (CNN) architecture - multi-group norm constraint CNN (MGNC-CNN) that capitalizes on multiple sets of word embeddings for sentence classification. MGNC-CNN extracts features from input embedding sets independently and then joins these at the penultimate layer in the network to form a final feature vector. We then adopt a group regularization strategy that differentially penalizes weights associated with the subcomponents generated from the respective embedding sets. This model is much simpler than comparable alternative architectures and requires substantially less training time. Furthermore, it is flexible in that it does not require input word embeddings to be of the same dimensionality. We show that MGNC-CNN consistently outperforms baseline models. Stephen Roller, Byron C. Wallace |
HLT-NAACL | 3 |
| 2016 | A Correlated Worker Model for Grouped, Imbalanced and Multitask Data
An T. Nguyen 0001, Byron C. Wallace, Matthew Lease |
UAI | 2 |
| 2016 | RobotReviewer: evaluation of a system for automatically assessing bias in clinical trialsabstractOBJECTIVE: To develop and evaluate RobotReviewer, a machine learning (ML) system that automatically assesses bias in clinical trials. From a (PDF-formatted) trial report, the system should determine risks of bias for the domains defined by the Cochrane Risk of Bias (RoB) tool, and extract supporting text for these judgments. METHODS: We algorithmically annotated 12,808 trial PDFs using data from the Cochrane Database of Systematic Reviews (CDSR). Trials were labeled as being at low or high/unclear risk of bias for each domain, and sentences were labeled as being informative or not. This dataset was used to train a multi-task ML model. We estimated the accuracy of ML judgments versus humans by comparing trials with two or more independent RoB assessments in the CDSR. Twenty blinded experienced reviewers rated the relevance of supporting text, comparing ML output with equivalent (human-extracted) text from the CDSR. RESULTS: By retrieving the top 3 candidate sentences per document (top3 recall), the best ML text was rated more relevant than text from the CDSR, but not significantly (60.4% ML text rated 'highly relevant' v 56.5% of text from reviews; difference +3.9%, [-3.2% to +10.9%]). Model RoB judgments were less accurate than those from published reviews, though the difference was <10% (overall accuracy 71.0% with ML v 78.3% with CDSR). CONCLUSION: Risk of bias assessment may be automated with reasonable accuracy. Automatically identified text supporting bias assessment is of equal quality to the manually identified text in the CDSR. This technology could substantially reduce reviewer workload and expedite evidence syntheses. Iain James Marshall, Joël Kuiper, Byron C. Wallace |
J. Am. Medical Informatics Assoc. | 3 |
| 2016 | Improving the utility of MeSH® terms using the TopicalMeSH representation
Elmer V. Bernstam, Trevor Cohen, Byron C. Wallace, Todd R. Johnson |
J. Biomed. Informatics | 4 |
| 2016 | Extracting PICO Sentences from Clinical Trial Reports using Supervised Distant SupervisionabstractSystematic reviews underpin Evidence Based Medicine (EBM) by addressing precise clinical questions via comprehensive synthesis of all relevant published evidence. Authors of systematic reviews typically define a Population/Problem, Intervention, Comparator, and Outcome (a PICO criteria) of interest, and then retrieve, appraise and synthesize results from all reports of clinical trials that meet these criteria. Identifying PICO elements in the full-texts of trial reports is thus a critical yet time-consuming step in the systematic review process. We seek to expedite evidence synthesis by developing machine learning models to automatically extract sentences from articles relevant to PICO elements. Collecting a large corpus of training data for this task would be prohibitively expensive. Therefore, we derive distant supervision (DS) with which to train models using previously conducted reviews. DS entails heuristically deriving 'soft' labels from an available structured resource. However, we have access only to unstructured, free-text summaries of PICO elements for corresponding articles; we must derive from these the desired sentence-level annotations. To this end, we propose a novel method -- supervised distant supervision (SDS) -- that uses a small amount of direct supervision to better exploit a large corpus of distantly labeled instances by learning to pseudo-annotate articles using the available DS. We show that this approach tends to outperform existing methods with respect to automated PICO extraction. Byron C. Wallace, Joël Kuiper, Aakash Sharma, Mingxi (Brian) Zhu, Iain James Marshall |
J. Mach. Learn. Res. | 1 |
| 2016 | Editorial: special issue on machine learning for health and medicine
Jenna Wiens, Byron C. Wallace |
Mach. Learn. | 2 |
| 2015 | Graph-Sparse LDA: A Topic Model with Structured SparsityabstractTopic modeling is a powerful tool for uncovering latent structure in many domains, including medicine, finance, and vision. The goals for the model vary depending on the application: sometimes the discovered topics are used for prediction or another downstream task. In other cases, the content of the topic may be of intrinsic scientific interest. Unfortunately, even when one uses modern sparse techniques, discovered topics are often difficult to interpret due to the high dimensionality of the underlying space. To improve topic interpretability, we introduce Graph-Sparse LDA, a hierarchical topic model that uses knowledge of relationships between words (e.g., as encoded by an ontology). In our model, topics are summarized by a few latent concept-words from the underlying graph that explain the observed words. Graph-Sparse LDA recovers sparse, interpretable summaries on two real-world biomedical datasets while matching state-of-the-art prediction performance. Finale Doshi-Velez, Byron C. Wallace, Ryan P. Adams |
AAAI | 2 |
| 2015 | Sparse, Contextually Informed Models for Irony Detection: Exploiting User Communities, Entities and SentimentabstractByron C. Wallace, Do Kook Choe, Eugene Charniak. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Byron C. Wallace, Do Kook Choe, Eugene Charniak |
ACL (1) | 1 |
| 2015 | Improving Retrieval of PubMed Articles Using the TopicalMeSH Representation
Elmer V. Bernstam, Trevor Cohen, Byron C. Wallace, Todd R. Johnson |
AMIA | 4 |
| 2015 | Combining Crowd and Expert Labels Using Decision Theoretic Active LearningabstractWe consider a finite-pool data categorization scenario which requires exhaustively classifying a given set of examples with a limited budget. We adopt a hybrid human-machine approach which blends automatic machine learning with human labeling across a tiered workforce composed of domain experts and crowd workers. To effectively achieve high-accuracy labels over the instances in the pool at minimal cost, we develop a novel approach based on decision-theoretic active learning. On the important task of biomedical citation screening for systematic reviews, results on real data show that our method achieves consistent improvements over baseline strategies. To foster further research by others, we have made our data available online. An T. Nguyen 0001, Byron C. Wallace, Matthew Lease |
HCOMP | 2 |
| 2015 | Automating Risk of Bias Assessment for Clinical TrialsabstractSystematic reviews, which summarize the entirety of the evidence pertaining to a specific clinical question, have become critical for evidence-based decision making in healthcare. But such reviews have become increasingly onerous to produce due to the exponentially expanding biomedical literature base. This study proposes a step toward mitigating this problem by automating risk of bias assessment in systematic reviews, in which reviewers determine whether study results may be affected by biases (e.g., poor randomization or blinding). Conducting risk of bias assessment is an important but onerous task. We thus describe a machine learning approach to automate this assessment, using the standard Cochrane Risk of Bias Tool which assesses seven common types of bias. Training such a system would typically require a large labeled corpus, which would be prohibitively expensive to collect here. Instead, we use distant supervision, using data from the Cochrane Database of Systematic Reviews (a large repository of systematic reviews), to pseudoannotate a corpus of 2200 clinical trial reports in PDF format. We then develop a joint model which, using the full text of a clinical trial report as input, predicts the risks of bias while simultaneously extracting the text fragments supporting these assessments. This study represents a step toward automating or semiautomating extraction of data necessary for the synthesis of clinical trials. Iain James Marshall, Joël Kuiper, Byron C. Wallace |
IEEE J. Biomed. Health Informatics | 3 |
| 2014 | Discovering Better AAAI Keywords via Clustering with Community-Sourced ConstraintsabstractSelecting good conference keywords is important because they often determine the composition of review committees and hence which papers are reviewed by whom. But presently conference keywords are generated in an ad-hoc manner by a small set of conference organizers. This approach is plainly not ideal. There is no guarantee, for example, that the generated keyword set aligns with what the community is actually working on and submitting to the conference in a given year. This is especially true in fast moving fields such as AI. The problem is exacerbated by the tendency of organizers to draw heavily on preceding years' keyword lists when generating a new set. Rather than a select few ordaining a keyword set that that represents AI at large, it would be preferable to generate these keywords more directly from the data, with input from research community members. To this end, we solicited feedback from seven AAAI PC members regarding a previously existing keyword set and used these 'community-sourced constraints' to inform a clustering over the abstracts of all submissions to AAAI 2013. We show that the keywords discovered via this data-driven, human-in-the-loop method are at least as preferred (by AAAI PC members) as 2013's manually generated set, and that they include categories previously overlooked by organizers. Many of the discovered terms were used for this year's conference. Kelly Moran, Byron C. Wallace, Carla E. Brodley |
AAAI | 2 |
| 2014 | Identifying Differences in Physician Communication Styles with a Log-Linear Transition Component ModelabstractWe consider the task of grouping doctors with respect to communication patterns exhibited in outpatient visits. We propose a novel approach toward this end in which we model speech act transitions in conversations via a log-linear model incorporating physician specific components. We train this model over transcripts of outpatient visits annotated with speech act codes and then cluster physicians in (a transformation of) this parameter space. We find significant correlations between the induced groupings and patient survey response data comprising ratings of physician communication. Furthermore, the novel sequential component model we leverage to induce this clustering allows us to explore differences across these groups. This work demonstrates how statistical AI might be used to better understand (and ultimately improve) physician communication. Byron C. Wallace, Issa J. Dahabreh, Thomas A. Trikalinos, Michael Barton Laws, Ira B. Wilson, Eugene Charniak |
AAAI | 1 |
| 2014 | Can Cognitive Scientists Help Computers Recognize Irony?
Byron C. Wallace, Laura Kertz |
CogSci | 1 |
| 2014 | Spá: A Web-Based Viewer for Text Mining in Evidence Based Medicine
Joël Kuiper, Iain James Marshall, Byron C. Wallace, Morris A. Swertz |
ECML/PKDD (3) | 3 |
| 2014 | A large-scale quantitative analysis of latent factors and sentiment in online doctor reviewsabstractOnline physician reviews are a massive and potentially rich source of information capturing patient sentiment regarding healthcare. We analyze a corpus comprising nearly 60,000 such reviews with a state-of-the-art probabilistic model of text. We describe a probabilistic generative model that captures latent sentiment across aspects of care (eg, interpersonal manner). We target specific aspects by leveraging a small set of manually annotated reviews. We perform regression analysis to assess whether model output improves correlation with state-level measures of healthcare. We report both qualitative and quantitative results. Model output correlates with state-level measures of quality healthcare, including patient likelihood of visiting their primary care physician within 14 days of discharge (p=0.03), and using the proposed model better predicts this outcome (p=0.10). We find similar results for healthcare expenditure. Generative models of text can recover important information from online physician reviews, facilitating large-scale analyses of such reviews. Byron C. Wallace, Michael J. Paul, Urmimala Sarkar, Thomas A. Trikalinos, Mark Dredze |
J. Am. Medical Informatics Assoc. | 1 |
| 2014 | Improving class probability estimates for imbalanced data
Byron C. Wallace, Issa J. Dahabreh |
Knowl. Inf. Syst. | 1 |
| 2013 | A Generative Joint, Additive, Sequential Model of Topics and Speech Acts in Patient-Doctor CommunicationabstractWe develop a novel generative model of conversation that jointly captures both the topical content and the speech act type associated with each utterance.Our model expresses both token emission and state transition probabilities as log-linear functions of separate components corresponding to topics and speech acts (and their interactions).We apply this model to a dataset comprising annotated patient-physician visits and show that the proposed joint approach outperforms a baseline univariate model. Byron C. Wallace, Thomas A. Trikalinos, Michael Barton Laws, Ira B. Wilson, Eugene Charniak |
EMNLP | 1 |
| 2012 | Class Probability Estimates are Unreliable for Imbalanced Data (and How to Fix Them)abstractObtaining good probability estimates is imperative for many applications. The increased uncertainty and typically asymmetric costs surrounding rare events increases this need. Experts (and classification systems) often rely on probabilities to inform decisions. However, we demonstrate that class probability estimates attained via supervised learning in imbalanced scenarios systematically underestimate the probabilities for minority class instances, despite ostensibly good overall calibration. To our knowledge, this problem has not previously been explored. Motivated by our exposition of this issue, we propose a simple, effective and theoretically motivated method to mitigate the bias of probability estimates for imbalanced data that bags estimators calibrated over balanced bootstrap samples. This approach drastically improves performance on the minority instances without greatly affecting overall calibration. We show that additional uncertainty can be exploited via a Bayesian approach by considering posterior distributions over bagged probability estimates. Byron C. Wallace, Issa J. Dahabreh |
ICDM | 1 |
| 2012 | Multiple Narrative Disentanglement: Unraveling Infinite Jest
Byron C. Wallace |
HLT-NAACL | 1 |
| 2011 | Class Imbalance, ReduxabstractClass imbalance (i.e., scenarios in which classes are unequally represented in the training data) occurs in many real-world learning tasks. Yet despite its practical importance, there is no established theory of class imbalance, and existing methods for handling it are therefore not well motivated. In this work, we approach the problem of imbalance from a probabilistic perspective, and from this vantage identify dataset characteristics (such as dimensionality, sparsity, etc.) that exacerbate the problem. Motivated by this theory, we advocate the approach of bagging an ensemble of classifiers induced over balanced bootstrap training samples, arguing that this strategy will often succeed where others fail. Thus in addition to providing a theoretical understanding of class imbalance, corroborated by our experiments on both simulated and real datasets, we provide practical guidance for the data mining practitioner working with imbalanced data. Byron C. Wallace, Kevin Small, Carla E. Brodley, Thomas A. Trikalinos |
ICDM | 1 |
| 2011 | The Constrained Weight Space SVM: Learning with Ranked Features
Kevin Small, Byron C. Wallace, Carla E. Brodley, Thomas A. Trikalinos |
ICML | 2 |
| 2011 | Who Should Label What? Instance Allocation in Multiple Expert Active LearningabstractThe active learning (AL) framework is an increasingly popular strategy for reducing the amount of human labeling effort required to induce a predictive model. Most work in AL has assumed that a single, infallible oracle provides labels requested by the learner at a fixed cost. However, real-world applications suitable for AL often include multiple domain experts who provide labels of varying cost and quality. We explore this multiple expert active learning (MEAL) scenario and develop a novel algorithm for instance allocation that exploits the meta-cognitive abilities of novice (cheap) experts in order to make the best use of the experienced (expensive) annotators. We demonstrate that this strategy outperforms strong baseline approaches to MEAL on both a sentiment analysis dataset and two datasets from our motivating application of biomedical citation screening. Furthermore, we provide evidence that novice labelers are often aware of which instances they are likely to mislabel. Byron C. Wallace, Kevin Small, Carla E. Brodley, Thomas A. Trikalinos |
SDM | 1 |
| 2010 | Active learning for biomedical citation screeningabstractActive learning (AL) is an increasingly popular strategy for mitigating the amount of labeled data required to train classifiers, thereby reducing annotator effort. We describe a real-world, deployed application of AL to the problem of biomedical citation screening for systematic reviews at the Tufts Medical Center's Evidence-based Practice Center. We propose a novel active learning strategy that exploits a priori domain knowledge provided by the expert (specifically, labeled features)and extend this model via a Linear Programming algorithm for situations where the expert can provide ranked labeled features. Our methods outperform existing AL strategies on three real-world systematic review datasets. We argue that evaluation must be specific to the scenario under consideration. To this end, we propose a new evaluation framework for finite-pool scenarios, wherein the primary aim is to label a fixed set of examples rather than to simply induce a good predictive model. We use a method from medical decision theory for eliciting the relative costs of false positives and false negatives from the domain expert, constructing a utility measure of classification performance that integrates the expert preferences. Our findings suggest that the expert can, and should, provide more information than instance labels alone. In addition to achieving strong empirical results on the citation screening problem, this work outlines many important steps for moving away from simulated active learning and toward deploying AL for real-world applications. Byron C. Wallace, Kevin Small, Carla E. Brodley, Thomas A. Trikalinos |
KDD | 1 |
| 2010 | Semi-automated screening of biomedical citations for systematic reviewsabstractBACKGROUND: Systematic reviews address a specific clinical question by unbiasedly assessing and analyzing the pertinent literature. Citation screening is a time-consuming and critical step in systematic reviews. Typically, reviewers must evaluate thousands of citations to identify articles eligible for a given review. We explore the application of machine learning techniques to semi-automate citation screening, thereby reducing the reviewers' workload. RESULTS: We present a novel online classification strategy for citation screening to automatically discriminate "relevant" from "irrelevant" citations. We use an ensemble of Support Vector Machines (SVMs) built over different feature-spaces (e.g., abstract and title text), and trained interactively by the reviewer(s). Semi-automating the citation screening process is difficult because any such strategy must identify all citations eligible for the systematic review. This requirement is made harder still due to class imbalance; there are far fewer "relevant" than "irrelevant" citations for any given systematic review. To address these challenges we employ a custom active-learning strategy developed specifically for imbalanced datasets. Further, we introduce a novel undersampling technique. We provide experimental results over three real-world systematic review datasets, and demonstrate that our algorithm is able to reduce the number of citations that must be screened manually by nearly half in two of these, and by around 40% in the third, without excluding any of the citations eligible for the systematic review. CONCLUSIONS: We have developed a semi-automated citation screening algorithm for systematic reviews that has the potential to substantially reduce the number of citations reviewers have to manually screen, without compromising the quality and comprehensiveness of the review. Byron C. Wallace, Thomas A. Trikalinos, Joseph Lau, Carla E. Brodley, Christopher H. Schmid |
BMC Bioinform. | 1 |