VLDB 2026 Research / reviewers in the wild / expert
Anna Rumshisky
dblp:63/873
· DBLP profile ↗
44ranked-venue papers
3as first author
12since 2021 · last 2025
0009-0008-5090-3765ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Diverse, not Short: A Length-Controlled Data Selection Strategy for Improving Response Diversity of Language ModelsabstractDiverse language model responses are crucial for creative generation, open-ended tasks, and self-improvement training.We show that common diversity metrics, and even reward models used for preference optimization, systematically bias models toward shorter outputs, limiting expressiveness.To address this, we introduce Diverse, not Short (Diverse-NS), a length-controlled data selection strategy that improves response diversity while maintaining length parity.By generating and filtering preference data that balances diversity, quality, and length, Diverse-NS enables effective training using only 3,000 preference pairs.Applied to LLaMA-3.1-8B and the Olmo-2 family, Diverse-NS substantially enhances lexical and semantic diversity.We show consistent improvement in diversity with minor reduction or gains in response quality on four creative generation tasks: Divergent Associations, Persona Generation, Alternate Uses, and Creative Writing.Surprisingly, experiments with the Olmo-2 model family (7B, and 13B) show that smaller models like Olmo-2-7B can serve as effective "diversity teachers" for larger models.By explicitly addressing length bias, our method efficiently pushes models toward more diverse and expressive outputs 1 . Vijeta Deshpande, Debasmita Ghose, John D. Patterson, Roger E. Beaty, Anna Rumshisky |
EMNLP | 5 |
| 2025 | MergeME: Model Merging Techniques for Homogeneous and Heterogeneous MoEsabstractYuhang Zhou, Giannis Karamanolakis, Victor Soto, Anna Rumshisky, Mayank Kulkarni, Furong Huang, Wei Ai, Jianhua Lu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Giannis Karamanolakis, Victor Soto, Anna Rumshisky, Mayank Kulkarni, Furong Huang, Wei Ai 0002, Jianhua Lu |
NAACL (Long Papers) | 4 |
| 2024 | NarrativeTime: Dense Temporal Annotation on a TimelineabstractFor the past decade, temporal annotation has been sparse: only a small portion of event pairs in a text was annotated. We present NarrativeTime, the first timeline-based annotation framework that achieves full coverage of all possible TLINKs. To compare with the previous SOTA in dense temporal annotation, we perform full re-annotation of the classic TimeBankDense corpus (American English), which shows comparable agreement with a signigicant increase in density. We contribute TimeBankNT corpus (with each text fully annotated by two expert annotators), extensive annotation guidelines, open-source tools for annotation and conversion to TimeML format, and baseline results. Anna Rogers, Marzena Karpinska, Vladislav Lialin, Gregory Smelkov, Anna Rumshisky |
LREC/COLING | 6 |
| 2024 | Deconstructing In-Context Learning: Understanding Prompts via CorruptionabstractThe ability of large language models (LLMs) to “learn in context” based on the provided prompt has led to an explosive growth in their use, culminating in the proliferation of AI assistants such as ChatGPT, Claude, and Bard. These AI assistants are known to be robust to minor prompt modifications, mostly due to alignment techniques that use human feedback. In contrast, the underlying pre-trained LLMs they use as a backbone are known to be brittle in this respect. Building high-quality backbone models remains a core challenge, and a common approach to assessing their quality is to conduct few-shot evaluation. Such evaluation is notorious for being highly sensitive to minor prompt modifications, as well as the choice of specific in-context examples. Prior work has examined how modifying different elements of the prompt can affect model performance. However, these earlier studies tended to concentrate on a limited number of specific prompt attributes and often produced contradictory results. Additionally, previous research either focused on models with fewer than 15 billion parameters or exclusively examined black-box models like GPT-3 or PaLM, making replication challenging. In the present study, we decompose the entire prompt into four components: task description, demonstration inputs, labels, and inline instructions provided for each demonstration. We investigate the effects of structural and semantic corruptions of these elements on model performance. We study models ranging from 1.5B to 70B in size, using ten datasets covering classification and generation tasks. We find that repeating text within the prompt boosts model performance, and bigger models (≥30B) are more sensitive to the semantics of the prompt. Finally, we observe that adding task and inline instructions to the demonstrations enhances model performance even when the instructions are semantically corrupted. The code is available at this URL. Namrata Shivagunde, Vladislav Lialin, Sherin Muckatira, Anna Rumshisky |
LREC/COLING | 4 |
| 2024 | ReLoRA: High-Rank Training Through Low-Rank UpdatesabstractDespite the dominance and effectiveness of scaling, resulting in large networks with hundreds of billions of parameters, the necessity to train overparameterized models remains poorly understood, while training costs grow exponentially. In this paper, we explore parameter-efficient training techniques as an approach to training large neural networks. We introduce a novel method called ReLoRA, which utilizes low-rank updates to train high-rank networks. We apply ReLoRA to training transformer language models with up to 1.3B parameters and demonstrate comparable performance to regular neural network training. ReLoRA saves up to 5.5Gb of RAM per GPU and improves training speed by 9-40% depending on the model size and hardware setup. Our findings show the potential of parameter- efficient techniques for large-scale pre-training. Our code is available on GitHub. Vladislav Lialin, Sherin Muckatira, Namrata Shivagunde, Anna Rumshisky |
ICLR | 4 |
| 2023 | Larger Probes Tell a Different Story: Extending Psycholinguistic Datasets Via In-Context LearningabstractLanguage model probing is often used to test specific capabilities of models.However, conclusions from such studies may be limited when the probing benchmarks are small and lack statistical power.In this work, we introduce new, larger datasets for negation (NEG-1500-SIMP) and role reversal (ROLE-1500) inspired by psycholinguistic studies.We dramatically extend existing NEG-136 and ROLE-88 benchmarks using GPT3, increasing their size from 18 and 44 sentence pairs to 750 each.We also create another version of extended negation dataset (NEG-1500-SIMP-TEMP), created using template-based generation.It consists of 770 sentence pairs.We evaluate 22 models on the extended datasets, seeing model performance dip 20-57% compared to the original smaller benchmarks.We observe high levels of negation sensitivity in models like BERT and ALBERT demonstrating that previous findings might have been skewed due to smaller test sets.Finally, we observe that while GPT3 has generated all the examples in ROLE-1500 is only able to solve 24.6% of them during probing.The datasets and code are available on Github 1 . Namrata Shivagunde, Vladislav Lialin, Anna Rumshisky |
EMNLP | 3 |
| 2023 | Self-Healing Through Error Detection, Attribution, and RetrainingabstractNegative feedback received from users of voice agents can provide valuable training signal to their underlying ML systems. However, such systems tend to have complex inference pipelines consisting of multiple model-based and deterministic components. Therefore, when negative feedback is received, it can be difficult to attribute the system error to a specific sub-component. In this work, we address this challenge by building a system for error attribution and correction. We prototype attributing errors to the ML models used for do-main classification (DC) in the NLU component of an assistant’s pipeline, using a combination of a model and rule based system. We propose a simple method to add these detected errors directly to offline DC model training, and study our system’s effectiveness on a challenging test set of low-frequency utterances. Our experiments on nine domains suggest that augmenting DC training data with our method significantly improves performance on a majority of them. Ansel MacLaughlin, Anna Rumshisky, Rinat Khaziev, Anil Ramakrishna, Yuval Merhav, Rahul Gupta 0001 |
ICASSP | 2 |
| 2023 | Sampling bias in NLU models: Impact and Mitigation
Zefei Li, Anil Ramakrishna, Anna Rumshisky, Andy Rosenbaum, Saleh Soltan, Rahul Gupta 0001 |
INTERSPEECH | 3 |
| 2022 | Down and Across: Introducing Crossword-Solving as a New NLP BenchmarkabstractSolving crossword puzzles requires diverse reasoning capabilities, access to a vast amount of knowledge about language and the world, and the ability to satisfy the constraints imposed by the structure of the puzzle.In this work, we introduce solving crossword puzzles as a new natural language understanding task.We release a corpus of crossword puzzles collected from the New York Times daily crossword spanning 25 years and comprised of a total of around nine thousand puzzles.These puzzles include a diverse set of clues: historic, factual, word meaning, synonyms/antonyms, fill-in-the-blank, abbreviations, prefixes/suffixes, wordplay, and crosslingual, as well as clues that depend on the answers to other clues.We separately release the clue-answer pairs from these puzzles as an open-domain question answering dataset containing over half a million unique clueanswer pairs.For the question answering task, our baselines include several sequence-tosequence and retrieval-based generative models.We also introduce a non-parametric constraint satisfaction baseline for solving the entire crossword puzzle.Finally, we propose an evaluation framework which consists of several complementary performance metrics. Saurabh Kulshreshtha, Olga Kovaleva, Namrata Shivagunde, Anna Rumshisky |
ACL (1) | 4 |
| 2022 | Life after BERT: What do Other Muppets Understand about Language?abstractExisting pre-trained transformer analysis works usually focus only on one or two model families at a time, overlooking the variability of the architecture and pre-training objectives.In our work, we utilize the oLMpics benchmark and psycholinguistic probing datasets for a diverse set of 29 models including T5, BART, and ALBERT.Additionally, we adapt the oLMpics zero-shot setup for autoregressive models and evaluate GPT networks of different sizes.Our findings show that none of these models can resolve compositional questions in a zero-shot fashion, suggesting that this skill is not learnable using existing pre-training objectives.Furthermore, we find that global model decisions such as architecture, directionality, size of the dataset, and pre-training objective are not predictive of a model's linguistic capabilities.The code for this study is available on GitHub 1 . Vladislav Lialin, Kevin Zhao, Namrata Shivagunde, Anna Rumshisky |
ACL (1) | 4 |
| 2022 | An Efficient DP-SGD Mechanism for Large Scale NLU ModelsabstractRecent advances in deep learning have drastically improved performance on many Natural Language Understanding (NLU) tasks. However, the data used to train NLU models may contain private information such as addresses or phone numbers, particularly when drawn from human subjects. It is desirable that underlying models do not expose private information contained in the training data. Differentially Private Stochastic Gradient Descent (DP-SGD) has been proposed as a mechanism to build privacy-preserving models. However, DP-SGD can be prohibitively slow to train. In this work, we propose a more efficient DP-SGD for training using a GPU infrastructure and apply it to fine-tuning models based on LSTM and transformer architectures. We report faster training times, alongside accuracy, theoretical privacy guarantees and success of Membership inference attacks for our models and observe that fine-tuning with proposed variant of DP-SGD can yield competitive models without significant degradation in training time and improvement in privacy protection. We also make observations such as looser theoretical ϵ, δ can translate into significant practical privacy gains. Christophe Dupuy, Radhika Arava, Rahul Gupta 0001, Anna Rumshisky |
ICASSP | 4 |
| 2022 | Federated Learning with Noisy User FeedbackabstractRahul Sharma, Anil Ramakrishna, Ansel MacLaughlin, Anna Rumshisky, Jimit Majmudar, Clement Chung, Salman Avestimehr, Rahul Gupta. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Anil Ramakrishna, Ansel MacLaughlin, Anna Rumshisky, Jimit Majmudar, Clement Chung, Amir Salman Avestimehr, Rahul Gupta 0001 |
NAACL-HLT | 4 |
| 2020 | Getting Closer to AI Complete Question Answering: A Set of Prerequisite Real TasksabstractThe recent explosion in question answering research produced a wealth of both factoid reading comprehension (RC) and commonsense reasoning datasets. Combining them presents a different kind of task: deciding not simply whether information is present in the text, but also whether a confident guess could be made for the missing information. We present QuAIL, the first RC dataset to combine text-based, world knowledge and unanswerable questions, and to provide question type annotation that would enable diagnostics of the reasoning strategies by a given QA system. QuAIL contains 15K multi-choice questions for 800 texts in 4 domains. Crucially, it offers both general and text-specific questions, unlikely to be found in pretraining data. We show that QuAIL poses substantial challenges to the current state-of-the-art systems, with a 30% drop in accuracy compared to the most similar existing dataset. Anna Rogers, Olga Kovaleva, Matthew Downey, Anna Rumshisky |
AAAI | 4 |
| 2020 | When BERT Plays the Lottery, All Tickets Are WinningabstractLarge Transformer-based models were shown to be reducible to a smaller number of selfattention heads and layers.We consider this phenomenon from the perspective of the lottery ticket hypothesis, using both structured and magnitude pruning.For fine-tuned BERT, we show that (a) it is possible to find subnetworks achieving performance that is comparable with that of the full model, and (b) similarly-sized subnetworks sampled from the rest of the model perform worse.Strikingly, with structured pruning even the worst possible subnetworks remain highly trainable, indicating that most pre-trained BERT weights are potentially useful.We also study the "good" subnetworks to see if their success can be attributed to superior linguistic knowledge, but find them unstable, and not explained by meaningful self-attention patterns. Sai Prasanna, Anna Rogers, Anna Rumshisky |
EMNLP (1) | 3 |
| 2020 | A Primer in BERTology: What We Know About How BERT WorksabstractTransformer-based models have pushed state of the art in many areas of NLP, but our understanding of what is behind their success is still limited. This paper is the first survey of over 150 studies of the popular BERT model. We review the current state of knowledge about how BERT works, what kind of information it learns and how it is represented, common modifications to its training objectives and architecture, the overparameterization issue, and approaches to compression. We then outline directions for future research. Anna Rogers, Olga Kovaleva, Anna Rumshisky |
Trans. Assoc. Comput. Linguistics | 3 |
| 2019 | Revealing the Dark Secrets of BERTabstractOlga Kovaleva, Alexey Romanov, Anna Rogers, Anna Rumshisky. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Olga Kovaleva, Alexey Romanov, Anna Rogers, Anna Rumshisky |
EMNLP/IJCNLP (1) | 4 |
| 2019 | MCN: A comprehensive corpus for medical concept normalization
Yen-Fu Luo, Weiyi Sun, Anna Rumshisky |
J. Biomed. Informatics | 3 |
| 2018 | Context-Aware Neural Model for Temporal Information ExtractionabstractWe propose a context-aware neural network model for temporal information extraction, with a uniform architecture for event-event, event-timex and timex-timex pairs.A Global Context Layer (GCL), inspired by the Neural Turing Machine (NTM), stores processed temporal relations in the narrative order, and retrieves them for use when the relevant entities are encountered.Relations are then classified in this larger context.The GCL model uses long-term memory and attention mechanisms to resolve long-distance dependencies that regular RNNs cannot recognize.GCL does not use postprocessing to resolve timegraph conflicts, outperforming previous approaches that do so.To our knowledge, GCL is also the first model to use an NTM-like architecture to incorporate the information about global context into discourse-scale processing of natural text. Yuanliang Meng, Anna Rumshisky |
ACL (1) | 2 |
| 2018 | Triad-based Neural Network for Coreference ResolutionabstractWe propose a triad-based neural network system that generates affinity scores between entity mentions for coreference resolution. The system simultaneously accepts three mentions as input, taking mutual dependency and logical constraints of all three mentions into account, and thus makes more accurate predictions than the traditional pairwise approach. Depending on system choices, the affinity scores can be further used in clustering or mention ranking. Our experiments show that a standard hierarchical clustering using the scores produces state-of-art results with MUC and B 3 metrics on the English portion of CoNLL 2012 Shared Task. The model does not rely on many handcrafted features and is easy to train and use. The triads can also be easily extended to polyads of higher orders. To our knowledge, this is the first neural network system to model mutual dependency of more than two members at mention level. Yuanliang Meng, Anna Rumshisky |
COLING | 2 |
| 2018 | What's in Your Embedding, And How It Predicts Task PerformanceabstractAttempts to find a single technique for general-purpose intrinsic evaluation of word embeddings have so far not been successful. We present a new approach based on scaled-up qualitative analysis of word vector neighborhoods that quantifies interpretable characteristics of a given model (e.g. its preference for synonyms or shared morphological forms as nearest neighbors). We analyze 21 such factors and show how they correlate with performance on 14 extrinsic and intrinsic task datasets (and also explain the lack of correlation between some of them). Our approach enables multi-faceted evaluation, parameter search, and generally – a more principled, hypothesis-driven approach to development of distributional semantic representations. Anna Rogers, Shashwath Hosur Ananthakrishna, Anna Rumshisky |
COLING | 3 |
| 2018 | RuSentiment: An Enriched Sentiment Analysis Dataset for Social Media in RussianabstractThis paper presents RuSentiment, a new dataset for sentiment analysis of social media posts in Russian, and a new set of comprehensive annotation guidelines that are extensible to other languages. RuSentiment is currently the largest in its class for Russian, with 31,185 posts annotated with Fleiss’ kappa of 0.58 (3 annotations per post). To diversify the dataset, 6,950 posts were pre-selected with an active learning-style strategy. We report baseline classification results, and we also release the best-performing embeddings trained on 3.2B tokens of Russian VKontakte posts. Anna Rogers, Alexey Romanov, Anna Rumshisky, Svitlana Volkova, Mikhail Gronas, Alex Gribov |
COLING | 3 |
| 2018 | Similarity-Based Reconstruction Loss for Meaning RepresentationabstractThis paper addresses the problem of representation learning.Using an autoencoder framework, we propose and evaluate several loss functions that can be used as an alternative to the commonly used cross-entropy reconstruction loss.The proposed loss functions use similarities between words in the embedding space, and can be used to train any neural model for text generation.We show that the introduced loss functions amplify semantic diversity of reconstructed sentences, while preserving the original meaning of the input.We test the derived autoencoder-generated representations on paraphrase detection and language inference tasks and demonstrate performance improvement compared to the traditional cross-entropy loss. Olga Kovaleva, Anna Rumshisky, Alexey Romanov |
EMNLP | 2 |
| 2018 | Automatic Labeling of Problem-Solving Dialogues for Computational Microgenetic Learning Analytics
Yuanliang Meng, Anna Rumshisky, Florence R. Sullivan |
LREC | 2 |
| 2017 | Temporal Information Extraction for Question Answering Using Syntactic Dependencies in an LSTM-based ArchitectureabstractIn this paper, we propose to use a set of simple, uniform in architecture LSTMbased models to recover different kinds of temporal relations from text.Using the shortest dependency path between entities as input, the same architecture is implemented to extract intra-sentence, crosssentence, and document creation time relations.A "double-checking" technique reverses entity pairs in classification, boosting the recall of positive cases and reducing misclassifications between opposite classes.An efficient pruning algorithm resolves conflicts globally.Evaluated on QA-TempEval (SemEval2015 Task 5), our proposed technique outperforms state-ofthe-art methods by a large margin.We also conduct intrinsic evaluation and post stateof-the-art results on Timebank-Dense. Yuanliang Meng, Anna Rumshisky, Alexey Romanov |
EMNLP | 2 |
| 2017 | Towards Debate Automation: a Recurrent Model for Predicting Debate WinnersabstractIn this paper we introduce a practical first step towards the creation of an automated debate agent: a state-of-the-art recurrent predictive model for predicting debate winners.By having an accurate predictive model, we are able to objectively rate the quality of a statement made at a specific turn in a debate.The model is based on a recurrent neural network architecture with attention, which allows the model to effectively account for the entire debate when making its prediction.Our model achieves state-of-the-art accuracy on a dataset of debate transcripts annotated with audience favorability of the debate teams.Finally, we discuss how future work can leverage our proposed model for the creation of an automated debate agent.We accomplish this by determining the model input that will maximize audience favorability toward a given side of a debate at an arbitrary turn. Peter Potash, Anna Rumshisky |
EMNLP | 2 |
| 2017 | Here's My Point: Joint Pointer Architecture for Argument MiningabstractIn order to determine argument structure in text, one must understand how individual components of the overall argument are linked. This work presents the first neural network-based approach to link extraction in argument mining. Specifically, we propose a novel architecture that applies Pointer Network sequence-to-sequence attention modeling to structural prediction in discourse parsing tasks. We then develop a joint model that extends this architecture to simultaneously address the link extraction task and the classification of argument components. The proposed joint model achieves state-of-the-art results on two separate evaluation corpora, showing far superior performance than the previously proposed corpus-specific and heavily feature-engineered models. Furthermore, our results demonstrate that jointly optimizing for both tasks is crucial for high performance. Peter Potash, Alexey Romanov, Anna Rumshisky |
EMNLP | 3 |
| 2017 | Length, Interchangeability, and External Knowledge: Observations from Predicting Argument ConvincingnessabstractIn this work, we provide insight into three key aspects related to predicting argument convincingness. First, we explicitly display the power that text length possesses for predicting convincingness in an unsupervised setting. Second, we show that a bag-of-words embedding model posts state-of-the-art on a dataset of arguments annotated for convincingness, outperforming an SVM with numerous hand-crafted features as well as recurrent neural network models that attempt to capture semantic composition. Finally, we assess the feasibility of integrating external knowledge when predicting convincingness, as arguments are often more convincing when they contain abundant information and facts. We finish by analyzing the correlations between the various models we propose. Peter Potash, Robin Bhattacharya, Anna Rumshisky |
IJCNLP(1) | 3 |
| 2016 | MUTT: Metric Unit TesTing for Language Generation TasksabstractPrecise evaluation metrics are important for assessing progress in high-level language generation tasks such as machine translation or image captioning.Historically, these metrics have been evaluated using correlation with human judgment.However, human-derived scores are often alarmingly inconsistent and are also limited in their ability to identify precise areas of weakness.In this paper, we perform a case study for metric evaluation by measuring the effect that systematic sentence transformations (e.g.active to passive voice) have on the automatic metric scores.These sentence "corruptions" serve as unit tests for precisely measuring the strengths and weaknesses of a given metric.We find that not only are human annotations heavily inconsistent in this study, but that the Metric Unit TesT analysis is able to capture precise shortcomings of particular metrics (e.g.comparing passive and active sentences) better than a simple correlation with human judgment can. Willie Boag, Renan Campos, Kate Saenko, Anna Rumshisky |
ACL (1) | 4 |
| 2016 | Interpretable Topic Features for Post-ICU Mortality Prediction
Yen-Fu Luo, Anna Rumshisky |
AMIA | 2 |
| 2015 | GhostWriter: Using an LSTM for Automatic Rap Lyric GenerationabstractThis paper demonstrates the effectiveness of a Long Short-Term Memory language model in our initial efforts to generate unconstrained rap lyrics.The goal of this model is to generate lyrics that are similar in style to that of a given rapper, but not identical to existing lyrics: this is the task of ghostwriting.Unlike previous work, which defines explicit templates for lyric generation, our model defines its own rhyme scheme, line length, and verse length.Our experiments show that a Long Short-Term Memory language model produces better "ghostwritten" lyrics than a baseline model. Peter Potash, Alexey Romanov, Anna Rumshisky |
EMNLP | 3 |
| 2015 | Normalization of relative and incomplete temporal expressions in clinical narrativesabstractOBJECTIVE: To improve the normalization of relative and incomplete temporal expressions (RI-TIMEXes) in clinical narratives. METHODS: We analyzed the RI-TIMEXes in temporally annotated corpora and propose two hypotheses regarding the normalization of RI-TIMEXes in the clinical narrative domain: the anchor point hypothesis and the anchor relation hypothesis. We annotated the RI-TIMEXes in three corpora to study the characteristics of RI-TMEXes in different domains. This informed the design of our RI-TIMEX normalization system for the clinical domain, which consists of an anchor point classifier, an anchor relation classifier, and a rule-based RI-TIMEX text span parser. We experimented with different feature sets and performed an error analysis for each system component. RESULTS: The annotation confirmed the hypotheses that we can simplify the RI-TIMEXes normalization task using two multi-label classifiers. Our system achieves anchor point classification, anchor relation classification, and rule-based parsing accuracy of 74.68%, 87.71%, and 57.2% (82.09% under relaxed matching criteria), respectively, on the held-out test set of the 2012 i2b2 temporal relation challenge. DISCUSSION: Experiments with feature sets reveal some interesting findings, such as: the verbal tense feature does not inform the anchor relation classification in clinical narratives as much as the tokens near the RI-TIMEX. Error analysis showed that underrepresented anchor point and anchor relation classes are difficult to detect. CONCLUSIONS: We formulate the RI-TIMEX normalization problem as a pair of multi-label classification problems. Considering only RI-TIMEX extraction and normalization, the system achieves statistically significant improvement over the RI-TIMEX results of the best systems in the 2012 i2b2 challenge. Weiyi Sun, Anna Rumshisky, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 2 |
| 2015 | Ease of adoption of clinical natural language processing software: An evaluation of five systemsabstractOBJECTIVE: In recognition of potential barriers that may inhibit the widespread adoption of biomedical software, the 2014 i2b2 Challenge introduced a special track, Track 3 - Software Usability Assessment, in order to develop a better understanding of the adoption issues that might be associated with the state-of-the-art clinical NLP systems. This paper reports the ease of adoption assessment methods we developed for this track, and the results of evaluating five clinical NLP system submissions. MATERIALS AND METHODS: A team of human evaluators performed a series of scripted adoptability test tasks with each of the participating systems. The evaluation team consisted of four "expert evaluators" with training in computer science, and eight "end user evaluators" with mixed backgrounds in medicine, nursing, pharmacy, and health informatics. We assessed how easy it is to adopt the submitted systems along the following three dimensions: communication effectiveness (i.e., how effective a system is in communicating its designed objectives to intended audience), effort required to install, and effort required to use. We used a formal software usability testing tool, TURF, to record the evaluators' interactions with the systems and 'think-aloud' data revealing their thought processes when installing and using the systems and when resolving unexpected issues. RESULTS: Overall, the ease of adoption ratings that the five systems received are unsatisfactory. Installation of some of the systems proved to be rather difficult, and some systems failed to adequately communicate their designed objectives to intended adopters. Further, the average ratings provided by the end user evaluators on ease of use and ease of interpreting output are -0.35 and -0.53, respectively, indicating that this group of users generally deemed the systems extremely difficult to work with. While the ratings provided by the expert evaluators are higher, 0.6 and 0.45, respectively, these ratings are still low indicating that they also experienced considerable struggles. DISCUSSION: The results of the Track 3 evaluation show that the adoptability of the five participating clinical NLP systems has a great margin for improvement. Remedy strategies suggested by the evaluators included (1) more detailed and operation system specific use instructions; (2) provision of more pertinent onscreen feedback for easier diagnosis of problems; (3) including screen walk-throughs in use instructions so users know what to expect and what might have gone wrong; (4) avoiding jargon and acronyms in materials intended for end users; and (5) packaging prerequisites required within software distributions so that prospective adopters of the software do not have to obtain each of the third-party components on their own. Kai Zheng 0002, V. G. Vinod Vydiswaran, Yang Liu 0019, Yue Wang 0035, Amber Stubbs, Özlem Uzuner, Anupama E. Gururaj, Samuel Bayer, John S. Aberdeen, Anna Rumshisky, Serguei V. S. Pakhomov, Hua Xu 0001 |
J. Biomed. Informatics | 10 |
| 2014 | Unfolding physiological state: mortality modelling in intensive care unitsabstractAccurate knowledge of a patient's disease state and trajectory is critical in a clinical setting. Modern electronic healthcare records contain an increasingly large amount of data, and the ability to automatically identify the factors that influence patient outcomes stand to greatly improve the efficiency and quality of care. We examined the use of latent variable models (viz. Latent Dirichlet Allocation) to decompose free-text hospital notes into meaningful features, and the predictive power of these features for patient mortality. We considered three prediction regimes: (1) baseline prediction, (2) dynamic (time-varying) outcome prediction, and (3) retrospective outcome prediction. In each, our prediction task differs from the familiar time-varying situation whereby data accumulates; since fewer patients have long ICU stays, as we move forward in time fewer patients are available and the prediction task becomes increasingly difficult. We found that latent topic-derived features were effective in determining patient mortality under three timelines: inhospital, 30 day post-discharge, and 1 year post-discharge mortality. Our results demonstrated that the latent topic features important in predicting hospital mortality are very different from those that are important in post-discharge mortality. In general, latent topic features were more predictive than structured features, and a combination of the two performed best. The time-varying models that combined latent topic features and baseline features had AUCs that reached 0.85, 0.80, and 0.77 for in-hospital, 30 day post-discharge and 1 year post-discharge mortality respectively. Our results agreed with other work suggesting that the first 24 hours of patient information are often the most predictive of hospital mortality. Retrospective models that used a combination of latent topic features and structured features achieved AUCs of 0.96, 0.82, and 0.81 for in-hospital, 30 day, and 1-year mortality prediction. Our work focuses on the dynamic (time-varying) setting because models from this regime could facilitate an on-going severity stratification system that helps direct care-staff resources and inform treatment strategies. Marzyeh Ghassemi, Tristan Naumann, Finale Doshi-Velez, Nicole Brimmer, Rohit Joshi, Anna Rumshisky, Peter Szolovits |
KDD | 6 |
| 2014 | Research and applications: Word sense disambiguation in the clinical domain: a comparison of knowledge-rich and knowledge-poor unsupervised methodsabstractOBJECTIVE: To evaluate state-of-the-art unsupervised methods on the word sense disambiguation (WSD) task in the clinical domain. In particular, to compare graph-based approaches relying on a clinical knowledge base with bottom-up topic-modeling-based approaches. We investigate several enhancements to the topic-modeling techniques that use domain-specific knowledge sources. MATERIALS AND METHODS: The graph-based methods use variations of PageRank and distance-based similarity metrics, operating over the Unified Medical Language System (UMLS). Topic-modeling methods use unlabeled data from the Multiparameter Intelligent Monitoring in Intensive Care (MIMIC II) database to derive models for each ambiguous word. We investigate the impact of using different linguistic features for topic models, including UMLS-based and syntactic features. We use a sense-tagged clinical dataset from the Mayo Clinic for evaluation. RESULTS: The topic-modeling methods achieve 66.9% accuracy on a subset of the Mayo Clinic's data, while the graph-based methods only reach the 40-50% range, with a most-frequent-sense baseline of 56.5%. Features derived from the UMLS semantic type and concept hierarchies do not produce a gain over bag-of-words features in the topic models, but identifying phrases from UMLS and using syntax does help. DISCUSSION: Although topic models outperform graph-based methods, semantic features derived from the UMLS prove too noisy to improve performance beyond bag-of-words. CONCLUSIONS: Topic modeling for WSD provides superior results in the clinical domain; however, integration of knowledge remains to be effectively exploited. Rachel Chasin, Anna Rumshisky, Özlem Uzuner, Peter Szolovits |
J. Am. Medical Informatics Assoc. | 2 |
| 2013 | Evaluating temporal relations in clinical text: 2012 i2b2 ChallengeabstractBACKGROUND: The Sixth Informatics for Integrating Biology and the Bedside (i2b2) Natural Language Processing Challenge for Clinical Records focused on the temporal relations in clinical narratives. The organizers provided the research community with a corpus of discharge summaries annotated with temporal information, to be used for the development and evaluation of temporal reasoning systems. 18 teams from around the world participated in the challenge. During the workshop, participating teams presented comprehensive reviews and analysis of their systems, and outlined future research directions suggested by the challenge contributions. METHODS: The challenge evaluated systems on the information extraction tasks that targeted: (1) clinically significant events, including both clinical concepts such as problems, tests, treatments, and clinical departments, and events relevant to the patient's clinical timeline, such as admissions, transfers between departments, etc; (2) temporal expressions, referring to the dates, times, durations, or frequencies phrases in the clinical text. The values of the extracted temporal expressions had to be normalized to an ISO specification standard; and (3) temporal relations, between the clinical events and temporal expressions. Participants determined pairs of events and temporal expressions that exhibited a temporal relation, and identified the temporal relation between them. RESULTS: For event detection, statistical machine learning (ML) methods consistently showed superior performance. While ML and rule based methods seemed to detect temporal expressions equally well, the best systems overwhelmingly adopted a rule based approach for value normalization. For temporal relation classification, the systems using hybrid approaches that combined ML and heuristics based methods produced the best results. Weiyi Sun, Anna Rumshisky, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 2 |
| 2013 | Temporal reasoning over clinical text: the state of the artabstractOBJECTIVES: To provide an overview of the problem of temporal reasoning over clinical text and to summarize the state of the art in clinical natural language processing for this task. TARGET AUDIENCE: This overview targets medical informatics researchers who are unfamiliar with the problems and applications of temporal reasoning over clinical text. SCOPE: We review the major applications of text-based temporal reasoning, describe the challenges for software systems handling temporal information in clinical text, and give an overview of the state of the art. Finally, we present some perspectives on future research directions that emerged during the recent community-wide challenge on text-based temporal reasoning in the clinical domain. Weiyi Sun, Anna Rumshisky, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 2 |
| 2013 | Annotating temporal information in clinical narratives
Weiyi Sun, Anna Rumshisky, Özlem Uzuner |
J. Biomed. Informatics | 2 |
| 2012 | Using UMLS for Word Sense Disambiguation in Clinical Notes
Anna Rumshisky, Rachel Chasin, Özlem Uzuner, Peter Szolovits |
AMIA | 1 |
| 2012 | Word Sense Inventories by Non-Experts
Anna Rumshisky, Nick Botchan, Sophie Kushkuley, James Pustejovsky |
LREC | 1 |
| 2011 | Domain-specific entity extraction from noisy, unstructured data using ontology-guided search
Sergey Bratus, Anna Rumshisky, Alexy Khrabrov 0001, Rajendra Magar |
Int. J. Document Anal. Recognit. | 2 |
| 2006 | Towards a Generative Lexical Resource: The Brandeis Semantic Ontology
James Pustejovsky, Catherine Havasi, Jessica Littman, Anna Rumshisky, Marc Verhagen |
LREC | 4 |
| 2006 | Inducing Sense-Discriminating Context Patterns from Sense-Tagged Corpora
Anna Rumshisky, James Pustejovsky |
LREC | 1 |
| 2005 | Automating Temporal Annotation with TARSQI
Marc Verhagen, Inderjeet Mani, Roser Saurí, Jessica Littman, Robert Knippen, Seok Bae Jang, Anna Rumshisky, John Phillips, James Pustejovsky |
ACL | 7 |
| 2004 | Automated Induction of Sense in Context
James Pustejovsky, Patrick Hanks, Anna Rumshisky |
COLING | 3 |