Dina Demner-Fushman

dblp:59/2029 · DBLP profile ↗
← Back
127ranked-venue papers
17as first author
13since 2021 · last 2026
0000-0002-4361-5799ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 93 · 15 first-author · 8 since 2021Artificial intelligence and machine learning · 25 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 14 · 2 since 2021Human-computer interaction and ubiquitous computing · 5Graphics, computer vision, multimedia, augmented reality and games · 3
YearPublicationVenuePosition
2026 Lessons from the TREC Plain Language Adaptation of Biomedical Abstracts (PLABA) track
Brian D. Ondov, William Xia, Kush Attal, Ishita Unde, Jerry He, Dina Demner-Fushman
J. Biomed. Informatics6
2024 Empowering Language Model with Guided Knowledge Fusion for Biomedical Document Re-ranking
Dina Demner-Fushman
AIME (1)2
2024 Sentence-Aligned Simplification of Biomedical Abstracts
Brian D. Ondov, Dina Demner-Fushman
AIME (1)2
2024 Towards Answering Health-related Questions from Medical Videos: Datasets and Approaches
abstract
The increase in the availability of online videos has transformed the way we access information and knowledge. A growing number of individuals now prefer instructional videos as they offer a series of step-by-step procedures to accomplish particular tasks. Instructional videos from the medical domain may provide the best possible visual answers to first aid, medical emergency, and medical education questions. This paper focuses on answering health-related questions asked by health consumers by providing visual answers from medical videos. The scarcity of large-scale datasets in the medical domain is a key challenge that hinders the development of applications that can help the public with their health-related questions. To address this issue, we first proposed a pipelined approach to create two large-scale datasets: HealthVidQA-CRF and HealthVidQA-Prompt. Leveraging the datasets, we developed monomodal and multimodal approaches that can effectively provide visual answers from medical videos to natural language questions. We conducted a comprehensive analysis of the results and outlined the findings, focusing on the impact of the created datasets on model training and the significance of visual features in enhancing the performance of the monomodal and multi-modal approaches for medical visual answer localization task.
Kush Attal, Dina Demner-Fushman
LREC/COLING3
2024 Pedagogically Aligned Objectives Create Reliable Automatic Cloze Tests
abstract
The cloze training objective of Masked Language Models makes them a natural choice for generating plausible distractors for human cloze questions. However, distractors must also be both distinct and incorrect, neither of which is directly addressed by existing neural methods. Evaluation of recent models has also relied largely on automated metrics, which cannot demonstrate the reliability or validity of human comprehension tests. In this work, we first formulate the pedagogically motivated objectives of plausibility, incorrectness, and distinctiveness in terms of conditional distributions from language models. Second, we present an unsupervised, interpretable method that uses these objectives to jointly optimize sets of distractors. Third, we test the reliability and validity of the resulting cloze tests compared to other methods with human participants. We find our method has stronger correlation with teacher-created comprehension tests than the state-of-the-art neural method and is more internally consistent. Our implementation is freely available and can quickly create a multiple choice cloze test from any given passage.
Brian D. Ondov, Kush Attal, Dina Demner-Fushman
NAACL-HLT3
2024 On the role of the UMLS in supporting diagnosis generation proposed by Large Language Models
Majid Afshar, Yanjun Gao, Emma Croxford, Dina Demner-Fushman
J. Biomed. Informatics5
2023 The National Library of Medicine indexer assignment dataset: A new large-scale dataset for reviewer assignment research
abstract
MEDLINE is the National Library of Medicine's (NLM) journal citation database. It contains over 28 million references to biomedical and life science journal articles, and a key feature of the database is that all articles are indexed with NLM Medical Subject Headings (MeSH). The library employs a team of MeSH indexers, and in recent years they have been asked to index close to 1 million articles per year in order to keep MEDLINE up to date. An important part of the MEDLINE indexing process is the assignment of articles to indexers. High quality and timely indexing is only possible when articles are assigned to indexers with suitable expertise. This paper introduces the NLM indexer assignment dataset: a large dataset of 4.2 million indexer article assignments for articles indexed between 2011 and 2019. The dataset is shown to be a valuable testbed for expert matching and assignment algorithms, and indexer article assignment is also found to be useful domain-adaptive pre-training for the closely related task of reviewer assignment.
Alastair R. Rae, James G. Mork, Dina Demner-Fushman
J. Assoc. Inf. Sci. Technol.3
2023 Medical image retrieval via nearest neighbor search on pre-trained image features
Russell F. Loane, Soumya Gayen, Dina Demner-Fushman
Knowl. Based Syst.4
2022 A survey of automated methods for biomedical text simplification
abstract
OBJECTIVE: Plain language in medicine has long been advocated as a way to improve patient understanding and engagement. As the field of Natural Language Processing has progressed, increasingly sophisticated methods have been explored for the automatic simplification of existing biomedical text for consumers. We survey the literature in this area with the goals of characterizing approaches and applications, summarizing existing resources, and identifying remaining challenges. MATERIALS AND METHODS: We search English language literature using lists of synonyms for both the task (eg, "text simplification") and the domain (eg, "biomedical"), and searching for all pairs of these synonyms using Google Scholar, Semantic Scholar, PubMed, ACL Anthology, and DBLP. We expand search terms based on results and further include any pertinent papers not in the search results but cited by those that are. RESULTS: We find 45 papers that we deem relevant to the automatic simplification of biomedical text, with data spanning 7 natural languages. Of these (nonexclusively), 32 describe tools or methods, 13 present data sets or resources, and 9 describe impacts on human comprehension. Of the tools or methods, 22 are chiefly procedural and 10 are chiefly neural. CONCLUSIONS: Though neural methods hold promise for this task, scarcity of parallel data has led to continued development of procedural methods. Various low-resource mitigations have been proposed to advance neural methods, including paragraph-level and unsupervised models and augmentation of neural models with procedural elements drawing from knowledge bases. However, high-quality parallel data will likely be crucial for developing fully automated biomedical text simplification.
Brian D. Ondov, Kush Attal, Dina Demner-Fushman
J. Am. Medical Informatics Assoc.3
2022 Question-aware transformer models for consumer health question summarization
Shweta Yadav 0001, Asma Ben Abacha, Dina Demner-Fushman
J. Biomed. Informatics4
2021 Reproducibility in biomedical natural language processing: A FAIR approach to what we need to know
Kevin Cohen 0001, Anna Ripple, Asma Ben Abacha, Olivier Bodenreider, Orin Hargraves, Karin Verspoor, Pierre Zweigenbaum, Dina Demner-Fushman
AMIA8
2021 The 2021 ImageCLEF Benchmark: Multimedia Retrieval in Medical, Nature, Internet and Social Media Applications
Bogdan Ionescu, Henning Müller, Renaud Péteri, Asma Ben Abacha, Dina Demner-Fushman, Sadid A. Hasan, Mourad Sarrouti, Obioma Pelka, Christoph M. Friedrich, Alba Garcia Seco de Herrera, Janadhip Jacutprakart, Vassili Kovalev, Serge Kozlovski, Vitali Liauchuk, Yashin Dicente Cid, Jon Chamberlain, Adrian F. Clark, Antonio C. de A. Campello Jr., Hassan Moustahfid, Thomas Oliver, Abigail Schulz, Paul Brie, Raul Berari, Dimitri Fichou, Andrei Tauteanu, Mihai Dogariu, Liviu-Daniel Stefan, Mihai Gabriel Constantin, Jérôme Deshayes-Chossart, Adrian Popescu 0001
ECIR (2)5
2021 Searching for scientific evidence in a pandemic: An overview of TREC-COVID
abstract
We present an overview of the TREC-COVID Challenge, an information retrieval (IR) shared task to evaluate search on scientific literature related to COVID-19. The goals of TREC-COVID include the construction of a pandemic search test collection and the evaluation of IR methods for COVID-19. The challenge was conducted over five rounds from April to July 2020, with participation from 92 unique teams and 556 individual submissions. A total of 50 topics (sets of related queries) were used in the evaluation, starting at 30 topics for Round 1 and adding 5 new topics per round to target emerging topics at that state of the still-emerging pandemic. This paper provides a comprehensive overview of the structure and results of TREC-COVID. Specifically, the paper provides details on the background, task structure, topic structure, corpus, participation, pooling, assessment, judgments, results, top-performing systems, lessons learned, and benchmark datasets.
Kirk Roberts, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, Kyle Lo, Ian Soboroff, Ellen M. Voorhees, Lucy Lu Wang, William R. Hersh
J. Biomed. Informatics4
2020 Automatic MeSH Indexing: Revisiting the Subheading Attachment Problem
Alastair R. Rae, David O. Pritchard, James G. Mork, Dina Demner-Fushman
AMIA4
2020 Flight of the PEGASUS? Comparing Transformers on Few-shot and Zero-shot Multi-document Abstractive Summarization
abstract
Recent work has shown that pre-trained Transformers obtain remarkable performance on many natural language processing tasks, including automatic summarization. However, most work has focused on (relatively) data-rich single-document summarization settings. In this paper, we explore highly-abstractive multi-document summarization, where the summary is explicitly conditioned on a user-given topic statement or question. We compare the summarization quality produced by three state-of-the-art transformer-based models: BART, T5, and PEGASUS. We report the performance on four challenging summarization datasets: three from the general domain and one from consumer health in both zero-shot and few-shot learning settings. While prior work has shown significant differences in performance for these models on standard summarization tasks, our results indicate that with as few as 10 labeled examples, there is no statistically significant difference in summary quality, suggesting the need for more abstractive benchmark collections when determining state-of-the-art.
Travis R. Goodwin, Max E. Savery, Dina Demner-Fushman
COLING3
2020 HOLMS: Alternative Summary Evaluation with Large Language Models
abstract
Efficient document summarization requires evaluation measures that can not only rank a set of systems based on an average score, but also highlight which individual summary is better than another.However, despite the very active research on summarization approaches, few works have proposed new evaluation measures in the recent years.The standard measures relied upon for the development of summarization systems are most often ROUGE and BLEU which, despite being efficient in overall system ranking, remain lexical in nature and have a limited potential when it comes to training neural networks.In this paper, we present a new hybrid evaluation measure for summarization, called HOLMS, that combines both language models pre-trained on large corpora and lexical similarity measures.Through several experiments, we show that HOLMS outperforms ROUGE and BLEU substantially in its correlation with human judgments on several extractive summarization datasets for both linguistic quality and pyramid scores.
Yassine Mrabet, Dina Demner-Fushman
COLING2
2020 ImageCLEF 2020: Multimedia Retrieval in Lifelogging, Medical, Nature, and Internet Applications
Bogdan Ionescu, Henning Müller, Renaud Péteri, Duc-Tien Dang-Nguyen, Liting Zhou, Luca Piras 0001, Michael Riegler 0001, Pål Halvorsen, Minh-Triet Tran, Mathias Lux, Cathal Gurrin, Jon Chamberlain, Adrian F. Clark, Antonio C. de A. Campello Jr., Alba Garcia Seco de Herrera, Asma Ben Abacha, Vivek V. Datla, Sadid A. Hasan, Joey Liu, Dina Demner-Fushman, Obioma Pelka, Christoph M. Friedrich, Yashin Dicente Cid, Serge Kozlovski, Vitali Liauchuk, Vassili Kovalev, Raul Berari, Paul Brie, Dimitri Fichou, Mihai Dogariu, Liviu-Daniel Stefan, Mihai Gabriel Constantin
ECIR (2)20
2020 Consumer health information and question answering: helping consumers find answers to their health-related information needs
abstract
OBJECTIVE: Consumers increasingly turn to the internet in search of health-related information; and they want their questions answered with short and precise passages, rather than needing to analyze lists of relevant documents returned by search engines and reading each document to find an answer. We aim to answer consumer health questions with information from reliable sources. MATERIALS AND METHODS: We combine knowledge-based, traditional machine and deep learning approaches to understand consumers' questions and select the best answers from consumer-oriented sources. We evaluate the end-to-end system and its components on simple questions generated in a pilot development of MedlinePlus Alexa skill, as well as the short and long real-life questions submitted to the National Library of Medicine by consumers. RESULTS: Our system achieves 78.7% mean average precision and 87.9% mean reciprocal rank on simple Alexa questions, and 44.5% mean average precision and 51.6% mean reciprocal rank on real-life questions submitted by National Library of Medicine consumers. DISCUSSION: The ensemble of deep learning, domain knowledge, and traditional approaches recognizes question type and focus well in the simple questions, but it leaves room for improvement on the real-life consumers' questions. Information retrieval approaches alone are sufficient for finding answers to simple Alexa questions. Answering real-life questions, however, benefits from a combination of information retrieval and inference approaches. CONCLUSION: A pilot practical implementation of research needed to help consumers find reliable answers to their health-related questions demonstrates that for most questions the reliable answers exist and can be found automatically with acceptable accuracy.
Dina Demner-Fushman, Yassine Mrabet, Asma Ben Abacha
J. Am. Medical Informatics Assoc.1
2020 A customizable deep learning model for nosocomial risk prediction from critical care notes with indirect supervision
abstract
OBJECTIVE: Reliable longitudinal risk prediction for hospitalized patients is needed to provide quality care. Our goal is to develop a generalizable model capable of leveraging clinical notes to predict healthcare-associated diseases 24-96 hours in advance. METHODS: We developed a reCurrent Additive Network for Temporal RIsk Prediction (CANTRIP) to predict the risk of hospital acquired (occurring ≥ 48 hours after admission) acute kidney injury, pressure injury, or anemia ≥ 24 hours before it is implicated by the patient's chart, labs, or notes. We rely on the MIMIC III critical care database and extract distinct positive and negative cohorts for each disease. We retrospectively determine the date-of-event using structured and unstructured criteria and use it as a form of indirect supervision to train and evaluate CANTRIP to predict disease risk using clinical notes. RESULTS: Our experiments indicate that CANTRIP, operating on text alone, obtains 74%-87% area under the curve and 77%-85% Specificity. Baseline shallow models showed lower performance on all metrics, while bidirectional long short-term memory obtained the highest Sensitivity at the cost of significantly lower Specificity and Precision. DISCUSSION: Proper model architecture allows clinical text to be successfully harnessed to predict nosocomial disease, outperforming shallow models and obtaining similar performance to disease-specific models reported in the literature. CONCLUSION: Clinical text on its own can provide a competitive alternative to traditional structured features (eg, lab values, vital signs). CANTRIP is able to generalize across nosocomial diseases without disease-specific feature extraction and is available at https://github.com/h4ste/cantrip.
Travis R. Goodwin, Dina Demner-Fushman
J. Am. Medical Informatics Assoc.2
2020 TREC-COVID: rationale and structure of an information retrieval shared task for COVID-19
abstract
TREC-COVID is an information retrieval (IR) shared task initiated to support clinicians and clinical research during the COVID-19 pandemic. IR for pandemics breaks many normal assumptions, which can be seen by examining 9 important basic IR research questions related to pandemic situations. TREC-COVID differs from traditional IR shared task evaluations with special considerations for the expected users, IR modality considerations, topic development, participant requirements, assessment process, relevance criteria, evaluation metrics, iteration process, projected timeline, and the implications of data use as a post-task test collection. This article describes how all these were addressed for the particular requirements of developing IR systems under a pandemic situation. Finally, initial participation numbers are also provided, which demonstrate the tremendous interest the IR community has in this effort.
Kirk Roberts, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, Kyle Lo, Ian Soboroff, Ellen M. Voorhees, Lucy Lu Wang, William R. Hersh
J. Am. Medical Informatics Assoc.4
2020 Understanding spatial language in radiology: Representation framework, annotation, and spatial relation extraction from chest X-ray reports using deep learning
Surabhi Datta, Yuqi Si, Laritza Rodriguez, Sonya E. Shooshan, Dina Demner-Fushman, Kirk Roberts
J. Biomed. Informatics5
2019 On the Summarization of Consumer Health Questions
abstract
Question understanding is one of the main challenges in question answering.In real world applications, users often submit natural language questions that are longer than needed and include peripheral information that increases the complexity of the question, leading to substantially more false positives in answer retrieval.In this paper, we study neural abstractive models for medical question summarization.We introduce the MeQSum corpus of 1,000 summarized consumer health questions.We explore data augmentation methods and evaluate state-of-the-art neural abstractive models on this new task.In particular, we show that semantic augmentation from question datasets improves the overall performance, and that pointer-generator networks outperform sequence-to-sequence attentional models on this task, with a ROUGE-1 score of 44.16%.We also present a detailed error analysis and discuss directions for improvement that are specific to question summarization.
Asma Ben Abacha, Dina Demner-Fushman
ACL (1)2
2019 Deep Learning from Incomplete Data: Detecting Imminent Risk of Hospital-acquired Pneumonia in ICU Patients
Travis R. Goodwin, Dina Demner-Fushman
AMIA2
2019 Classification Types: A New Feature in the SPECIALIST Lexicon
Chris J. Lu, Amanda Payne, Dina Demner-Fushman
AMIA3
2019 Drug-drug Interaction Extraction via Transfer Learning
Kin Wah Fung, Dina Demner-Fushman
AMIA3
2019 A High Recall Classifier for Selecting Articles for MEDLINE Indexing
Alastair R. Rae, Max E. Savery, James G. Mork, Dina Demner-Fushman
AMIA4
2019 Evaluation of System for Selective Indexing Classification
Max E. Savery, Melanie Huston, James G. Mork, Olga Printseva, Susan Schmidt, Alastair R. Rae, Dina Demner-Fushman
AMIA7
2019 ImageCLEF 2019: Multimedia Retrieval in Lifelogging, Medical, Nature, and Security Applications
Bogdan Ionescu, Henning Müller, Renaud Péteri, Duc-Tien Dang-Nguyen, Luca Piras 0001, Michael Riegler 0001, Minh-Triet Tran, Mathias Lux, Cathal Gurrin, Yashin Dicente Cid, Vitali Liauchuk, Vassili Kovalev, Asma Ben Abacha, Sadid A. Hasan, Vivek V. Datla, Joey Liu, Dina Demner-Fushman, Obioma Pelka, Christoph M. Friedrich, Jon Chamberlain, Adrian F. Clark, Alba Garcia Seco de Herrera, Narciso García, Ergina Kavallieratou, Carlos R. del-Blanco, Carlos Cuevas, Nikos Vasilopoulos, Konstantinos Karampidis
ECIR (2)17
2019 A question-entailment approach to question answering
abstract
BACKGROUND: One of the challenges in large-scale information retrieval (IR) is developing fine-grained and domain-specific methods to answer natural language questions. Despite the availability of numerous sources and datasets for answer retrieval, Question Answering (QA) remains a challenging problem due to the difficulty of the question understanding and answer extraction tasks. One of the promising tracks investigated in QA is mapping new questions to formerly answered questions that are "similar". RESULTS: We propose a novel QA approach based on Recognizing Question Entailment (RQE) and we describe the QA system and resources that we built and evaluated on real medical questions. First, we compare logistic regression and deep learning methods for RQE using different kinds of datasets including textual inference, question similarity, and entailment in both the open and clinical domains. Second, we combine IR models with the best RQE method to select entailed questions and rank the retrieved answers. To study the end-to-end QA approach, we built the MedQuAD collection of 47,457 question-answer pairs from trusted medical sources which we introduce and share in the scope of this paper. Following the evaluation process used in TREC 2017 LiveQA, we find that our approach exceeds the best results of the medical task with a 29.8% increase over the best official score. CONCLUSIONS: The evaluation results support the relevance of question entailment for QA and highlight the effectiveness of combining IR and RQE for future QA efforts. Our findings also show that relying on a restricted set of reliable answer sources can bring a substantial improvement in medical QA.
Asma Ben Abacha, Dina Demner-Fushman
BMC Bioinform.2
2019 Spell checker for consumer language (CSpell)
abstract
Objective: Automated understanding of consumer health inquiries might be hindered by misspellings. To detect and correct various types of spelling errors in consumer health questions, we developed a distributable spell-checking tool, CSpell, that handles nonword errors, real-word errors, word boundary infractions, punctuation errors, and combinations of the above. Methods: We developed a novel approach of using dual embedding within Word2vec for context-dependent corrections. This technique was used in combination with dictionary-based corrections in a 2-stage ranking system. We also developed various splitters and handlers to correct word boundary infractions. All correction approaches are integrated to handle errors in consumer health questions. Results: Our approach achieves an F1 score of 80.93% and 69.17% for spelling error detection and correction, respectively. Discussion: The dual-embedding model shows a significant improvement (9.13%) in F1 score compared with the general practice of using cosine similarity with word vectors in Word2vec for context ranking. Our 2-stage ranking system shows a 4.94% improvement in F1 score compared with the best 1-stage ranking system. Conclusion: CSpell improves over the state of the art and provides near real-time automatic misspelling detection and correction in consumer health questions. The software and the CSpell test set are available at https://umlslex.nlm.nih.gov/cSpell.
Chris J. Lu, Alan R. Aronson, Sonya E. Shooshan, Dina Demner-Fushman
J. Am. Medical Informatics Assoc.4
2018 Finding medication doses in the literature
Dina Demner-Fushman, James G. Mork, Willie J. Rogers, Sonya E. Shooshan, Laritza Rodriguez, Alan R. Aronson
AMIA1
2018 Adverse Reactions and Drug-Drug Interaction Extraction tracks at the Text Analysis Conference (TAC)
Dina Demner-Fushman, Joseph M. Tonning, Kin Wah Fung, Phong Do, Richard D. Boyce, Kirk Roberts
AMIA1
2018 Increasing UMLS Coverage and Reducing Ambiguity via Automated Creation of Synonymous Terms
François-Michel Lang, James G. Mork, Dina Demner-Fushman, Alan R. Aronson
AMIA3
2018 Improving Spelling Correction with Consumer Health Terminology
Chris J. Lu, Dina Demner-Fushman
AMIA2
2018 Benchmarking Information Retrieval for Precision Oncology: the TREC Precision Medicine Track
Kirk Roberts, Dina Demner-Fushman, Ellen M. Voorhees, William R. Hersh, Steven Bedrick, Alexander J. Lazar, Shubham Pant
AMIA2
2018 Semantic annotation of consumer health questions
abstract
BACKGROUND: Consumers increasingly use online resources for their health information needs. While current search engines can address these needs to some extent, they generally do not take into account that most health information needs are complex and can only fully be expressed in natural language. Consumer health question answering (QA) systems aim to fill this gap. A major challenge in developing consumer health QA systems is extracting relevant semantic content from the natural language questions (question understanding). To develop effective question understanding tools, question corpora semantically annotated for relevant question elements are needed. In this paper, we present a two-part consumer health question corpus annotated with several semantic categories: named entities, question triggers/types, question frames, and question topic. The first part (CHQA-email) consists of relatively long email requests received by the U.S. National Library of Medicine (NLM) customer service, while the second part (CHQA-web) consists of shorter questions posed to MedlinePlus search engine as queries. Each question has been annotated by two annotators. The annotation methodology is largely the same between the two parts of the corpus; however, we also explain and justify the differences between them. Additionally, we provide information about corpus characteristics, inter-annotator agreement, and our attempts to measure annotation confidence in the absence of adjudication of annotations. RESULTS: The resulting corpus consists of 2614 questions (CHQA-email: 1740, CHQA-web: 874). Problems are the most frequent named entities, while treatment and general information questions are the most common question types. Inter-annotator agreement was generally modest: question types and topics yielded highest agreement, while the agreement for more complex frame annotations was lower. Agreement in CHQA-web was consistently higher than that in CHQA-email. Pairwise inter-annotator agreement proved most useful in estimating annotation confidence. CONCLUSIONS: To our knowledge, our corpus is the first focusing on annotation of uncurated consumer health questions. It is currently used to develop machine learning-based methods for question understanding. We make the corpus publicly available to stimulate further research on consumer health QA.
Halil Kilicoglu, Asma Ben Abacha, Yassine Mrabet, Sonya E. Shooshan, Laritza Rodriguez, Kate Masterton, Dina Demner-Fushman
BMC Bioinform.7
2017 TextFlow: A Text Similarity Measure based on Continuous Sequences
abstract
Text similarity measures are used in multiple tasks such as plagiarism detection, information ranking and recognition of paraphrases and textual entailment.While recent advances in deep learning highlighted further the relevance of sequential models in natural language generation, existing similarity measures do not fully exploit the sequential nature of language.Examples of such similarity measures include ngrams and skip-grams overlap which rely on distinct slices of the input texts.In this paper we present a novel text similarity measure inspired from a common representation in DNA sequence alignment algorithms.The new measure, called TextFlow, represents input text pairs as continuous curves and uses both the actual position of the words and sequence matching to compute the similarity value.Our experiments on eight different datasets show very encouraging results in paraphrase detection, textual entailment recognition and ranking relevance.
Yassine Mrabet, Halil Kilicoglu, Dina Demner-Fushman
ACL (1)3
2017 CTB: A Custom Taxonomy Builder for Named Entity Extraction
Dina Demner-Fushman, Willie J. Rogers
AMIA1
2017 Evaluation of Clinical Text Segmentation to Facilitate Cohort Retrieval
Tracy Edinger, Dina Demner-Fushman, Aaron M. Cohen, Steven Bedrick, William R. Hersh
AMIA2
2017 New MetaMap Features for Processing Numerical Tables
François-Michel Lang, Dina Demner-Fushman
AMIA2
2017 Information Retrieval for Biomedical Datasets: The 2016 bioCADDIE Challenge
Kirk Roberts, Anupama E. Gururaj, Saeid Pournejati, Trevor Cohen, William R. Hersh, Dina Demner-Fushman, Lucila Ohno-Machado, Hua Xu 0001
AMIA7
2017 Mining the literature for genes associated with placenta-mediated maternal diseases
Laritza Rodriguez, Stephanie M. Morrison, Kathleen Greenberg, Dina Demner-Fushman
AMIA4
2017 Named entity recognition in functional neuroimaging literature
abstract
Human neuroimaging research aims to find mappings between brain activity and broad cognitive states. In particular, Functional Magnetic Resonance Imaging (fMRI) allows collecting information about activity in the brain in a non-invasive way. In this paper, we tackle the task of linking brain activity information from fMRI data with named entities expressed in functional neuroimaging literature. For the automatic extraction of those links, we focus on Named Entity Recognition (NER) and compare different methods to recognize relevant entities from fMRI literature. We selected 15 entity categories to describe cognitive states, anatomical areas, stimuli and responses. To cope with the lack of relevant training data, we proposed rule-based methods relying on noun-phrase detection and filtering. We also developed machine learning methods based on Conditional Random Fields (CRF) with morpho-syntactic and semantic features. We constructed a gold standard corpus to evaluate these different NER methods. A comparison of the obtained F1 scores showed that the proposed approaches significantly outperform three state-of-the-art methods in open and specific domains with a best result of 78.79% F1 score in exact span evaluation and 98.40% F1 in inexact span evaluation.
Asma Ben Abacha, Alba Garcia Seco de Herrera, L. Rodney Long, Sameer K. Antani, Dina Demner-Fushman
BIBM6
2017 MetaMap Lite: an evaluation of a new Java implementation of MetaMap
abstract
MetaMap is a widely used named entity recognition tool that identifies concepts from the Unified Medical Language System Metathesaurus in text. This study presents MetaMap Lite, an implementation of some of the basic MetaMap functions in Java. On several collections of biomedical literature and clinical text, MetaMap Lite demonstrated real-time speed and precision, recall, and F1 scores comparable to or exceeding those of MetaMap and other popular biomedical text processing tools, clinical Text Analysis and Knowledge Extraction System (cTAKES) and DNorm.
Dina Demner-Fushman, Willie J. Rogers, Alan R. Aronson
J. Am. Medical Informatics Assoc.1
2017 Automated classification of eligibility criteria in clinical trials to facilitate patient-trial matching for specific patient populations
abstract
OBJECTIVE: To develop automated classification methods for eligibility criteria in ClinicalTrials.gov to facilitate patient-trial matching for specific populations such as persons living with HIV or pregnant women. MATERIALS AND METHODS: We annotated 891 interventional cancer trials from ClinicalTrials.gov based on their eligibility for human immunodeficiency virus (HIV)-positive patients using their eligibility criteria. These annotations were used to develop classifiers based on regular expressions and machine learning (ML). After evaluating classification of cancer trials for eligibility of HIV-positive patients, we sought to evaluate the generalizability of our approach to more general diseases and conditions. We annotated the eligibility criteria for 1570 of the most recent interventional trials from ClinicalTrials.gov for HIV-positive and pregnancy eligibility, and the classifiers were retrained and reevaluated using these data. RESULTS: On the cancer-HIV dataset, the baseline regex model, the bag-of-words ML classifier, and the ML classifier with named entity recognition (NER) achieved macro-averaged F2 scores of 0.77, 0.87, and 0.87, respectively; the addition of NER did not result in a significant performance improvement. On the general dataset, ML + NER achieved macro-averaged F2 scores of 0.91 and 0.85 for HIV and pregnancy, respectively. DISCUSSION AND CONCLUSION: The eligibility status of specific patient populations, such as persons living with HIV and pregnant women, for clinical trials is of interest to both patients and clinicians. We show that it is feasible to develop a high-performing, automated trial classification system for eligibility status that can be integrated into consumer-facing search engines as well as patient-trial matching systems.
Dina Demner-Fushman
J. Am. Medical Informatics Assoc.2
2017 A protocol-driven approach to automatically finding authoritative answers to consumer health questions in online resources
abstract
The purpose of this research was to establish an upper bound on finding answers to health‐related questions in MedlinePlus and other online resources. Seven reference librarians tested a set of protocols to determine whether it was possible to use the types and foci of the questions extracted from customer requests submitted to the National Library of Medicine to find authoritative answers to these questions. Librarians tested the protocols manually to determine if the process was sufficiently robust and accurate to later automate. Results indicated that the extracted terms provide enough information to find authoritative answers for about 60% of questions and that certain question types are more likely to result in authoritative answers than others. The question corpus and analysis performed for this project will inform automatic question answering systems, and could lead to suggestions for new content to include in MedlinePlus. This approach can serve as an example to researchers interested in methods of evaluating question answering tools and the contents of online databases.
Ariel Deardorff, Kate Masterton, Kirk Roberts, Halil Kilicoglu, Dina Demner-Fushman
J. Assoc. Inf. Sci. Technol.5
2016 Recognizing Question Entailment for Medical Question Answering
Asma Ben Abacha, Dina Demner-Fushman
AMIA2
2016 Natural Language Processing Working Group Pre-Symposium: Graduate Student Consortium and 'Hackathon'
Stéphane M. Meystre, Sivaram Arabandi, Kavishwar B. Wagholikar, Jon D. Patrick, Guergana K. Savova, Chunhua Weng, Pierre Zweigenbaum, Dina Demner-Fushman, Özlem Uzuner, Hua Xu 0001
AMIA10
2016 Resolving Hierarchical Ambiguity in Indexing Recommendations
James G. Mork, Dina Demner-Fushman
AMIA2
2016 Combining Open-domain and Biomedical Knowledge for Topic Recognition in Consumer Health Questions
Yassine Mrabet, Halil Kilicoglu, Kirk Roberts, Dina Demner-Fushman
AMIA4
2016 Resource Classification for Medical Questions
Kirk Roberts, Laritza Rodriguez, Sonya E. Shooshan, Dina Demner-Fushman
AMIA4
2016 Towards automatic discovery of Genes related to Human Placenta
Laritza Rodriguez, Stephanie M. Morrison, Kathleen Greenberg, Dina Demner-Fushman
AMIA4
2016 Modality Classification for Searching Figures in Biomedical Literature
abstract
Image modality classification categorizes images according to their type. It is an important module in the Open-iSM multimodal (text+image) search engine that retrieves figures from biomedical articles. It is a hierarchical classification where on the top level the input figures are classified into two general categories: regular images (X-ray, CT, MRI, photographs, etc.) vs. illustration images (cartoon sketch, charts, graphs, etc.). This binary classification task is challenged by the vast diversity of visual material (image type), and the way it is organized (simple or compound figures). We present two methods for this binary classification: (i) Support Vector Machines (SVM) with manually-selected features, including a feature based on semantic concepts, and, (ii) Deep Learning method which avoids the process of feature handcrafting. Both methods were tested and compared on a dataset of 16400 figures. Both methods achieved good performance (above 95% accuracy). The slightly better performance of the feature-based method demonstrates the effectiveness of the features we chose.
Zhiyun Xue, Sameer K. Antani, L. Rodney Long, Dina Demner-Fushman, George R. Thoma
CBMS5
2016 A Hybrid Approach to Generation of Missing Abstracts in Biomedical Literature
abstract
Readers usually rely on abstracts to identify relevant medical information from scientific articles. Abstracts are also essential to advanced information retrieval methods. More than 50 thousand scientific publications in PubMed lack author-generated abstracts, and the relevancy judgements for these papers have to be based on their titles alone. In this paper, we propose a hybrid summarization technique that aims to select the most pertinent sentences from articles to generate an extractive summary in lieu of a missing abstract. We combine i) health outcome detection, ii) keyphrase extraction, and iii) textual entailment recognition between sentences. We evaluate our hybrid approach and analyze the improvements of multi-factor summarization over techniques that rely on a single method, using a collection of 295 manually generated reference summaries. The obtained results show that the hybrid approach outperforms the baseline techniques with an improvement of 13% in recall and 4% in F1 score.
Suchet K. Chachra, Asma Ben Abacha, Sonya E. Shooshan, Laritza Rodriguez, Dina Demner-Fushman
COLING5
2016 Learning to Read Chest X-Rays: Recurrent Neural Cascade Model for Automated Image Annotation
abstract
Despite the recent advances in automatically describing image contents, their applications have been mostly limited to image caption datasets containing natural images (e.g., Flickr 30k, MSCOCO). In this paper, we present a deep learning model to efficiently detect a disease from an image and annotate its contexts (e.g., location, severity and the affected organs). We employ a publicly available radiology dataset of chest x-rays and their reports, and use its image annotations to mine disease names to train convolutional neural networks (CNNs). In doing so, we adopt various regularization techniques to circumvent the large normalvs-diseased cases bias. Recurrent neural networks (RNNs) are then trained to describe the contexts of a detected disease, based on the deep CNN features. Moreover, we introduce a novel approach to use the weights of the already trained pair of CNN/RNN on the domain-specific image/text dataset, to infer the joint image/text contexts for composite image labeling. Significantly improved image annotation results are demonstrated using the recurrent neural cascade model by taking the joint image/text contexts into account.
Hoo-Chang Shin, Kirk Roberts, Le Lu 0001, Dina Demner-Fushman, Jianhua Yao 0001, Ronald M. Summers
CVPR4
2016 Unsupervised Ranking of Knowledge Bases for Named Entity Recognition
abstract
With the continuous growth of freely accessible knowledge bases and the heterogeneity of textual corpora, selecting the most adequate knowledge base for named entity recognition is becoming a challenge in itself. In this paper, we propose an unsupervised method to rank knowledge bases according to their adequacy for the recognition of named entities in a given corpus. Building on a state-of-the-art, unsupervised entity linking approach, we propose several evaluation metrics to measure the lexical and structural adequacy of a knowledge base for a given corpus. We study the correlation between these metrics and three standard performance measures: precision, recall and F1 score. Our multi-domain experiments on 9 different corpora with 6 knowledge bases show that three of the proposed metrics are strong performance predictors having 0.62 to 0.76 Pearson correlation with precision and 0.96 correlation with both recall and F1 score.
Yassine Mrabet, Halil Kilicoglu, Dina Demner-Fushman
ECAI3
2016 Annotating Named Entities in Consumer Health Questions
Halil Kilicoglu, Asma Ben Abacha, Yassine Mrabet, Kirk Roberts, Laritza Rodriguez, Sonya E. Shooshan, Dina Demner-Fushman
LREC7
2016 Annotating Logical Forms for EHR Questions
Kirk Roberts, Dina Demner-Fushman
LREC2
2016 State-of-the-art in biomedical literature retrieval for clinical cases: a survey of the TREC 2014 CDS track
Kirk Roberts, Matthew S. Simpson, Dina Demner-Fushman, Ellen M. Voorhees, William R. Hersh
Inf. Retr. J.3
2016 Preparing a collection of radiology examinations for distribution and retrieval
abstract
OBJECTIVE: Clinical documents made available for secondary use play an increasingly important role in discovery of clinical knowledge, development of research methods, and education. An important step in facilitating secondary use of clinical document collections is easy access to descriptions and samples that represent the content of the collections. This paper presents an approach to developing a collection of radiology examinations, including both the images and radiologist narrative reports, and making them publicly available in a searchable database. MATERIALS AND METHODS: The authors collected 3996 radiology reports from the Indiana Network for Patient Care and 8121 associated images from the hospitals' picture archiving systems. The images and reports were de-identified automatically and then the automatic de-identification was manually verified. The authors coded the key findings of the reports and empirically assessed the benefits of manual coding on retrieval. RESULTS: The automatic de-identification of the narrative was aggressive and achieved 100% precision at the cost of rendering a few findings uninterpretable. Automatic de-identification of images was not quite as perfect. Images for two of 3996 patients (0.05%) showed protected health information. Manual encoding of findings improved retrieval precision. CONCLUSION: Stringent de-identification methods can remove all identifiers from text radiology reports. DICOM de-identification of images does not remove all identifying information and needs special attention to images scanned from film. Adding manual coding to the radiologist narrative reports significantly improved relevancy of the retrieved clinical documents. The de-identified Indiana chest X-ray collection is available for searching and downloading from the National Library of Medicine (http://openi.nlm.nih.gov/).
Dina Demner-Fushman, Marc D. Kohli, Marc B. Rosenman, Sonya E. Shooshan, Laritza Rodriguez, Sameer K. Antani, George R. Thoma, Clement J. McDonald
J. Am. Medical Informatics Assoc.1
2016 Interactive use of online health resources: a comparison of consumer and professional questions
abstract
OBJECTIVE: To understand how consumer questions on online resources differ from questions asked by professionals, and how such consumer questions differ across resources. MATERIALS AND METHODS: Ten online question corpora, 5 consumer and 5 professional, with a combined total of over 40 000 questions, were analyzed using a variety of natural language processing techniques. These techniques analyze questions at the lexical, syntactic, and semantic levels, exposing differences in both form and content. RESULTS: Consumer questions tend to be longer than professional questions, more closely resemble open-domain language, and focus far more on medical problems. Consumers ask more sub-questions, provide far more background information, and ask different types of questions than professionals. Furthermore, there is substantial variance of these factors between the different consumer corpora. DISCUSSION: The form of consumer questions is highly dependent upon the individual online resource, especially in the amount of background information provided. Professionals, on the other hand, provide very little background information and often ask much shorter questions. The content of consumer questions is also highly dependent upon the resource. While professional questions commonly discuss treatments and tests, consumer questions focus disproportionately on symptoms and diseases. Further, consumers place far more emphasis on certain types of health problems (eg, sexual health). CONCLUSION: Websites for consumers to submit health questions are a popular online resource filling important gaps in consumer health information. By analyzing how consumers write questions on these resources, we can better understand these gaps and create solutions for improving information access.This article is part of the Special Focus on Person-Generated Health and Wellness Data, which published in the May 2016 issue, Volume 23, Issue 3.
Kirk Roberts, Dina Demner-Fushman
J. Am. Medical Informatics Assoc.2
2015 Automated searches for personalized evidence to prevent hospital acquired infection
Amos Cahan, Sonya E. Shooshan, Laritza Rodriguez, Dina Demner-Fushman
AMIA4
2015 Extracting Characteristics of the Study Subjects from Full-Text Articles
Dina Demner-Fushman, James G. Mork
AMIA1
2015 An Ensemble Method for Spelling Correction in Consumer Health Questions
Halil Kilicoglu, Marcelo Fiszman, Kirk Roberts, Dina Demner-Fushman
AMIA4
2015 Automatic Extraction and Post-coordination of Spatial Relations in Consumer Language
Kirk Roberts, Laritza Rodriguez, Sonya E. Shooshan, Dina Demner-Fushman
AMIA4
2015 Automatic Classification of Structured Product Labels for Pregnancy Risk Drug Categories, a Machine Learning Approach
Laritza Rodriguez, Dina Demner-Fushman
AMIA2
2015 Foreign object detection in chest X-rays
abstract
Automatic analysis of chest X-ray images is one important approach for screening/identifying pulmonary diseases. The existence of foreign objects in the images hinders the performance of such processing. In this paper, we focus on one type of foreign objects that is often shown in the images of a large dataset of chest X-rays we are working on-the buttons on the gown that the patient is wearing. The method we propose involves four major steps: intensity normalization, low contrast image identification and enhancement, segmentation of lung regions, and button object extraction. Based on the characteristics of the button objects, we applied two methods for the step of button object extraction. One was based on the circular Hough transform; the other was based on the Viola-Jones algorithm. We tested and compared both methods using a ground truth dataset containing 505 button objects. The results demonstrate the effectiveness of the proposed method.
Zhiyun Xue, Sema Candemir, Sameer K. Antani, L. Rodney Long, Stefan Jäger 0001, Dina Demner-Fushman, George R. Thoma
BIBM6
2015 The role of fine-grained annotations in supervised recognition of risk factors for heart disease from EHRs
abstract
This paper describes a supervised machine learning approach for identifying heart disease risk factors in clinical text, and assessing the impact of annotation granularity and quality on the system's ability to recognize these risk factors. We utilize a series of support vector machine models in conjunction with manually built lexicons to classify triggers specific to each risk factor. The features used for classification were quite simple, utilizing only lexical information and ignoring higher-level linguistic information such as syntax and semantics. Instead, we incorporated high-quality data to train the models by annotating additional information on top of a standard corpus. Despite the relative simplicity of the system, it achieves the highest scores (micro- and macro-F1, and micro- and macro-recall) out of the 20 participants in the 2014 i2b2/UTHealth Shared Task. This system obtains a micro- (macro-) precision of 0.8951 (0.8965), recall of 0.9625 (0.9611), and F1-measure of 0.9276 (0.9277). Additionally, we perform a series of experiments to assess the value of the annotated data we created. These experiments show how manually-labeled negative annotations can improve information extraction performance, demonstrating the importance of high-quality, fine-grained natural language annotations.
Kirk Roberts, Sonya E. Shooshan, Laritza Rodriguez, Swapna Abhyankar, Halil Kilicoglu, Dina Demner-Fushman
J. Biomed. Informatics6
2014 Addressing some statistical challenges of using EHR data for clinical research
Fiona M. Callaghan, Dina Demner-Fushman, Swapna Abhyankar, Matthew T. Jackson, Mallika Mundkur, Clement J. McDonald
AMIA2
2014 Vocabulary Density Method for Customized Indexing of MEDLINE Journals
James G. Mork, Dina Demner-Fushman, Susan Schmidt, Alan R. Aronson
AMIA2
2014 Error Propagation in EHRs via Copy/Paste: An Analysis of Relative Dates
Kirk Roberts, Amos Cahan, Dina Demner-Fushman
AMIA3
2014 Automatically Classifying Question Types for Consumer Health Questions
Kirk Roberts, Halil Kilicoglu, Marcelo Fiszman, Dina Demner-Fushman
AMIA4
2014 Facilitating Reconciliation of Inter-Annotator Disagreements
Johann Stan, Olivier Bodenreider, Kin Wah Fung, Dina Demner-Fushman
AMIA4
2014 Biomedical image segmentation for semantic visual feature extraction
abstract
Biomedical photographs comprise diverse optically acquired images. Accurate classification into meaningful subclasses is valuable in biomedical image retrieval systems. Conventional visual descriptors are limited in their ability to assign semantic labels to images for meaningful retrieval. In this paper we propose a Markov random field (MRF)-based biomedical image segmentation method to segment images into meaningful regions that can be associated with semantic labels. We focus on several tissue image types and develop two MRF models: (i) for tissue image detection from large photograph collection; and, (ii) for region segmentation and semantic labeling. Experimental results demonstrate that our method can detect tissue images in about 82% precision, and our proposed visual descriptors computed from the segmentation results outperform existing visual descriptors. This latter result can be effectively used in biomedical image retrieval systems for retrieving tissue images.
Daekeun You, Sameer K. Antani, Dina Demner-Fushman, George R. Thoma
BIBM3
2014 Body Segment Classification for Visible Human Cross Section Slices
abstract
Visible human data has been widely used in various medical research and computer science applications. We present a new application for this data: a method to classify which body segment a transverse cross section image belongs to. The labeling of the data is created with the guidance of an online body cross section tutorial. The visual properties of the images are represented using a variety of feature descriptors. To avoid problems that arise from the large dimensionality of features, feature selection is applied. The multi-class SVM is employed as the classifier. Both the CT scans and the color photographs of cryosections of the whole body (male and female) are used to test the proposed method. The high performance with overall accuracy above 98% on both the 2160 CT dataset and the 1870 cryosectional photos show the method is very promising. Because of its observed effectiveness on visible human data, we will extend our approach to classify figures in biomedical articles.
Zhiyun Xue, Sameer K. Antani, L. Rodney Long, Dina Demner-Fushman, George R. Thoma
CBMS4
2014 Does Figure-Text Improve Biomedical Article Retrieval? A Pilot Study
abstract
Graphical illustrations are often used in biomedical articles to convey statistical results, schematics, etc. They are frequently accompanied with superimposed text annotations, or figure-text. It is generally assumed that this figure-text provides information that complements associated textual metadata, such as the captions, or bibliographic citations for indexing the figures and enhanced retrieval quality. However, to the best of our knowledge nothing in the literature adequately supports the assumption of information gain. In this article, we report the results of a pilot study that compares image retrieval based on figure captions alone to that using figure-text in addition to captions. In the blind study two judges evaluated a set of figures retrieved on the topic of "lung cancer". We find that figure-text could help improve retrieval performance for specific queries.
Daekeun You, Sameer K. Antani, Dina Demner-Fushman, George R. Thoma
CBMS3
2014 An MRF Model for Biomedical Image Segmentation
abstract
We propose a Markov random field (MRF)-based method to segment photographic biomedical images into three image sub-regions, viz., tissue, photo, and background. Segmentation results are then used to extract local and global visual features to separate images with tissue, such as endoscopic images, from general photographs.
Daekeun You, Sameer K. Antani, Dina Demner-Fushman, George R. Thoma
CBMS3
2014 Annotating Question Decomposition on Complex Medical Questions
Kirk Roberts, Kate Masterton, Marcelo Fiszman, Halil Kilicoglu, Dina Demner-Fushman
LREC5
2014 Integrating visual words as bunch of n-grams for effective biomedical image classification
abstract
The Bag-of-Visual-Words (BoVW) has been frequently used in the classification of image data. However, this modeling approach does not take into consideration the spatial relationships of these words, which is important for similarity measurement between images. We have developed a novel technique to incorporate spatial information of visual words based on the n-grams representation. The method encodes regional layout with a 2-gram representation in the local keypoint neighborhood. The region is divided in two zones to capture the relative orientations of pair-wise visual words. In turn, each image is described by an accumulated vector of 2-grams. Then, we compute the Shannon entropy over a random “bunch” of 2-grams to reduce the dimensionality of the feature vector. We discovered that this reduction technique creates a more discriminative feature vector as well as presents a considerable dimensionality reduction of up to 99%. The final representation is a compact and efficient local image descriptor that encodes frequency and arrangement of visual words. The proposed approach was tested by classifying a standard biomedical image dataset into categories defined by image modality and body part. The experimental results demonstrate the importance of contextual relations of visual words. Our proposed approach improved the classification accuracy compared to the traditional BoVW by 6.03%.
Glauco Vitor Pedrosa, Sameer K. Antani, Dina Demner-Fushman, L. Rodney Long, Agma J. M. Traina
WACV4
2014 Multimodal biomedical image indexing and retrieval using descriptive text and global feature mapping
Matthew S. Simpson, Dina Demner-Fushman, Sameer K. Antani, George R. Thoma
Inf. Retr.2
2014 Research and applications: Combining structured and unstructured data to identify a cohort of ICU patients who received dialysis
abstract
OBJECTIVE: To develop a generalizable method for identifying patient cohorts from electronic health record (EHR) data-in this case, patients having dialysis-that uses simple information retrieval (IR) tools. METHODS: We used the coded data and clinical notes from the 24,506 adult patients in the Multiparameter Intelligent Monitoring in Intensive Care database to identify patients who had dialysis. We used SQL queries to search the procedure, diagnosis, and coded nursing observations tables based on ICD-9 and local codes. We used a domain-specific search engine to find clinical notes containing terms related to dialysis. We manually validated the available records for a 10% random sample of patients who potentially had dialysis and a random sample of 200 patients who were not identified as having dialysis based on any of the sources. RESULTS: We identified 1844 patients that potentially had dialysis: 1481 from the three coded sources and 1624 from the clinical notes. Precision for identifying dialysis patients based on available data was estimated to be 78.4% (95% CI 71.9% to 84.2%) and recall was 100% (95% CI 86% to 100%). CONCLUSIONS: Combining structured EHR data with information from clinical notes using simple queries increases the utility of both types of data for cohort identification. Patients identified by more than one source are more likely to meet the inclusion criteria; however, including patients found in any of the sources increases recall. This method is attractive because it is available to researchers with access to EHR data and off-the-shelf IR tools.
Swapna Abhyankar, Dina Demner-Fushman, Fiona M. Callaghan, Clement J. McDonald
J. Am. Medical Informatics Assoc.2
2013 A simple method to extract key maternal data from neonatal clinical notes
Swapna Abhyankar, Dina Demner-Fushman
AMIA2
2013 Mining MEDLINE for problems associated with vitamin D
Dina Demner-Fushman, James G. Mork, Alan R. Aronson
AMIA1
2013 Comparison and combination of several MeSH indexing approaches
Antonio Jimeno-Yepes, James G. Mork, Dina Demner-Fushman, Alan R. Aronson
AMIA3
2013 Extracting drug indication information from structured product labels using natural language processing
abstract
OBJECTIVE: To extract drug indications from structured drug labels and represent the information using codes from standard medical terminologies. MATERIALS AND METHODS: We used MetaMap and other publicly available resources to extract information from the indications section of drug labels. Drugs and indications were encoded by RxNorm and UMLS identifiers respectively. A sample was manually reviewed. We also compared the results with two independent information sources: National Drug File-Reference Terminology and the Semantic Medline project. RESULTS: A total of 6797 drug labels were processed, resulting in 19 473 unique drug-indication pairs. Manual review of 298 most frequently prescribed drugs by seven physicians showed a recall of 0.95 and precision of 0.77. Inter-rater agreement (Fleiss κ) was 0.713. The precision of the subset of results corroborated by Semantic Medline extractions increased to 0.93. DISCUSSION: Correlation of a patient's medical problems and drugs in an electronic health record has been used to improve data quality and reduce medication errors. Authoritative drug indication information is available from drug labels, but not in a format readily usable by computer applications. Our study shows that it is feasible to use publicly available natural language processing resources to extract drug indications from drug labels. The same method can be applied to other sections of the drug label-for example, adverse effects, contraindications. CONCLUSIONS: It is feasible to use publicly available natural language processing tools to extract indication information from freely available drug labels. Named entity recognition sources (eg, MetaMap) provide reasonable recall. Combination with other data sources provides higher precision.
Kin Wah Fung, Chiang S. Jao, Dina Demner-Fushman
J. Am. Medical Informatics Assoc.3
2013 Image retrieval from scientific publications: Text and image content processing to separate multipanel figures
abstract
Images contained in scientific publications are widely considered useful for educational and research purposes, and their accurate indexing is critical for efficient and effective retrieval. Such image retrieval is complicated by the fact that figures in the scientific literature often combine multiple individual subfigures (panels). Multipanel figures are in fact the predominant pattern in certain types of scientific publications. The goal of this work is to automatically segment multipanel figures—a necessary step for automatic semantic indexing and in the development of image retrieval systems targeting the scientific literature. We have developed a method that uses the image content as well as the associated figure caption to: (1) automatically detect panel boundaries; (2) detect panel labels in the images and convert them to text; and (3) detect the labels and textual descriptions of each panel within the captions. Our approach combines the output of image‐content and text‐based processing steps to split the multipanel figures into individual subfigures and assign to each subfigure its corresponding section of the caption. The developed system achieved precision of 81% and recall of 73% on the task of automatic segmentation of multipanel figures.
Emilia Apostolova, Daekeun You, Zhiyun Xue, Sameer K. Antani, Dina Demner-Fushman, George R. Thoma
J. Assoc. Inf. Sci. Technol.5
2012 Research Designs in Mesh and Emtree: A Comparative Study of Coverage
Tanja Bekhuis, Dina Demner-Fushman, Rebecca S. Jacobson
AMIA2
2012 Towards the Creation of a Visual Ontology of Biomedical Imaging Entities
Matthew S. Simpson, Daekeun You, Sameer K. Antani, George R. Thoma, Dina Demner-Fushman
AMIA6
2012 Window Classification of Brain CT Images in Biomedical Articles
Zhiyun Xue, Sameer K. Antani, L. Rodney Long, Dina Demner-Fushman, George R. Thoma
AMIA4
2012 Screening nonrandomized studies for medical systematic reviews: A comparative study of classifiers
Tanja Bekhuis, Dina Demner-Fushman
Artif. Intell. Medicine2
2012 Standardizing clinical laboratory data for secondary use
Swapna Abhyankar, Dina Demner-Fushman, Clement J. McDonald
J. Biomed. Informatics2
2012 A mutation-centric approach to identifying pharmacogenomic relations in text
Bastien Rance, Emily Doughty, Dina Demner-Fushman, Maricel G. Kann, Olivier Bodenreider
J. Biomed. Informatics3
2011 Detecting Figure-Panel Labels in Medical Journal Articles Using MRF
abstract
We present a method for figure-panel (subfigure) label detection and recognition in multi-panel figures extracted from biomedical articles. Figures in biomedical articles often comprise several subfigures that are identified by superimposed panel labels ('A', 'B', ...) which are referenced in the figure caption and discussion in the article body. Splitting such multi-panel figures into individual subfigures is a necessary step for improved multimodal biomedical information retrieval. Prior to feature extraction for indexing and retrieval of biomedical figures it is necessary to classify image content in each subfigure by its modality (X-ray, MRI, CT, etc.) and other relevant criteria. Subfigure labels are valuable in associating individual panels with relevant text in captions and discussion. We propose a 4-step panel label detection method based on Markov Random Field (MRF). Experiments on 515 multi-panel figures and analysis of the results show promising results. We present the successes and identify critical challenges.
Daekeun You, Sameer K. Antani, Dina Demner-Fushman, Venu Govindaraju, George R. Thoma
ICDAR3
2010 Extracting Rx information from clinical narrative
abstract
OBJECTIVE: The authors used the i2b2 Medication Extraction Challenge to evaluate their entity extraction methods, contribute to the generation of a publicly available collection of annotated clinical notes, and start developing methods for ontology-based reasoning using structured information generated from the unstructured clinical narrative. DESIGN: Extraction of salient features of medication orders from the text of de-identified hospital discharge summaries was addressed with a knowledge-based approach using simple rules and lookup lists. The entity recognition tool, MetaMap, was combined with dose, frequency, and duration modules specifically developed for the Challenge as well as a prototype module for reason identification. MEASUREMENTS: Evaluation metrics and corresponding results were provided by the Challenge organizers. RESULTS: The results indicate that robust rule-based tools achieve satisfactory results in extraction of simple elements of medication orders, but more sophisticated methods are needed for identification of reasons for the orders and durations. LIMITATIONS: Owing to the time constraints and nature of the Challenge, some obvious follow-on analysis has not been completed yet. CONCLUSIONS: The authors plan to integrate the new modules with MetaMap to enhance its accuracy. This integration effort will provide guidance in retargeting existing tools for better processing of clinical text.
James G. Mork, Olivier Bodenreider, Dina Demner-Fushman, Rezarta Islamaj Dogan, François-Michel Lang, Zhiyong Lu, Aurélie Névéol, Lee B. Peters, Sonya E. Shooshan, Alan R. Aronson
J. Am. Medical Informatics Assoc.3
2010 UMLS content views appropriate for NLP processing of the biomedical literature vs. clinical text
Dina Demner-Fushman, James G. Mork, Sonya E. Shooshan, Alan R. Aronson
J. Biomed. Informatics1
2010 Interactive publication: The document as a research tool
George R. Thoma, Glenn Ford, Sameer K. Antani, Dina Demner-Fushman, Matthew S. Simpson
J. Web Semant.4
2009 Using Non-Lexical Features to Identify Effective Indexing Terms for Biomedical Illustrations
Matthew S. Simpson, Dina Demner-Fushman, Charles Sneiderman, Sameer K. Antani, George R. Thoma
EACL2
2009 The potential for automated question answering in the context of genomic medicine: an assessment of existing resources and properties of answers
abstract
Knowledge gained in studies of genetic disorders is reported in a growing body of biomedical literature containing reports of genetic variation in individuals that map to medical conditions and/or response to therapy. These scientific discoveries need to be translated into practical applications to optimize patient care. Translating research into practice can be facilitated by supplying clinicians with research evidence. We assessed the role of existing tools in extracting answers to translational research questions in the area of genomic medicine. We: evaluate the coverage of translational research terms in the Unified Medical Language Systems (UMLS) Metathesaurus; determine where answers are most often found in full-text articles; and determine common answer patterns. Findings suggest that we will be able to leverage the UMLS in development of natural language processing algorithms for automated extraction of answers to translational research questions from biomedical text in the area of genomic medicine.
Casey Overby Taylor, Peter Tarczy-Hornoch, Dina Demner-Fushman
BMC Bioinform.3
2009 Viewpoint Paper: Towards Automatic Recognition of Scientifically Rigorous Clinical Research Evidence
abstract
The growing numbers of topically relevant biomedical publications readily available due to advances in document retrieval methods pose a challenge to clinicians practicing evidence-based medicine. It is increasingly time consuming to acquire and critically appraise the available evidence. This problem could be addressed in part if methods were available to automatically recognize rigorous studies immediately applicable in a specific clinical situation. We approach the problem of recognizing studies containing useable clinical advice from retrieved topically relevant articles as a binary classification problem. The gold standard used in the development of PubMed clinical query filters forms the basis of our approach. We identify scientifically rigorous studies using supervised machine learning techniques (Naïve Bayes, support vector machine (SVM), and boosting) trained on high-level semantic features. We combine these methods using an ensemble learning method (stacking). The performance of learning methods is evaluated using precision, recall and F(1) score, in addition to area under the receiver operating characteristic (ROC) curve (AUC). Using a training set of 10,000 manually annotated MEDLINE citations, and a test set of an additional 2,000 citations, we achieve 73.7% precision and 61.5% recall in identifying rigorous, clinically relevant studies, with stacking over five feature-classifier combinations and 82.5% precision and 84.3% recall in recognizing rigorous studies with treatment focus using stacking over word + metadata feature vector. Our results demonstrate that a high quality gold standard and advanced classification methods can help clinicians acquire best evidence from the medical literature.
Halil Kilicoglu, Dina Demner-Fushman, Thomas C. Rindflesch, Nancy L. Wilczynski, Robert Brian Haynes
J. Am. Medical Informatics Assoc.2
2009 What can natural language processing do for clinical decision support?
Dina Demner-Fushman, Wendy W. Chapman, Clement J. McDonald
J. Biomed. Informatics1
2009 Automatic summarization of MEDLINE citations for evidence-based medical treatment: A topic-oriented evaluation
Marcelo Fiszman, Dina Demner-Fushman, Halil Kilicoglu, Thomas C. Rindflesch
J. Biomed. Informatics2
2008 Methodology for Creating UMLS Content Views Appropriate for Biomedical Natural Language Processing
Alan R. Aronson, James G. Mork, Aurélie Névéol, Sonya E. Shooshan, Dina Demner-Fushman
AMIA5
2008 A Prototype System to Support Evidence-based Practice
Dina Demner-Fushman, Charlotte A. Seckman, Cheryl Fisher, Susan E. Hauser, Jennifer Clayton, George R. Thoma
AMIA1
2008 Toward Automatic Recognition of High Quality Clinical Evidence
Halil Kilicoglu, Dina Demner-Fushman, Thomas C. Rindflesch, Nancy L. Wilczynski, Robert Brian Haynes
AMIA2
2008 Themes in biomedical natural language processing: BioNLP08
abstract
A recent posting to the BioNLP mailing list notes that the past few months of 2008 have seen the appearance of over fifty papers on biomedical natural language processing/text mining (BioNLP). This number (which included medical, as well as genomic work) represents about as many papers on genomic language processing as existed in all of PubMed at the end of 2003 [1] – just five years ago, and the current supplement in BMC Bioinformatics presents another ten! These papers have in common the fact that they are follow-on work to papers originally published in the proceedings of the BioNLP 2008 workshop at the annual meeting of the Association for Computational Linguistics (ACL). All have gone through a separate rigorous review process and represent an advance beyond the work originally presented at the workshop. Like the annual BioNLP workshop itself, they represent a wide cross-section of the type of work that goes on in BioNLP today.
Dina Demner-Fushman, Sophia Ananiadou, Kevin Cohen 0001, John Pestian, Jun'ichi Tsujii, Bonnie L. Webber
BMC Bioinform.1
2007 Semantic Clustering of Answers to Clinical Questions
Jimmy Lin, Dina Demner-Fushman
AMIA2
2007 Frontiers of biomedical text mining: current progress
abstract
It is now almost 15 years since the publication of the first paper on text mining in the genomics domain, and decades since the first paper on text mining in the medical domain. Enormous progress has been made in the areas of information retrieval, evaluation methodologies and resource construction. Some problems, such as abbreviation-handling, can essentially be considered solved problems, and others, such as identification of gene mentions in text, seem likely to be solved soon. However, a number of problems at the frontiers of biomedical text mining continue to present interesting challenges and opportunities for great improvements and interesting research. In this article we review the current state of the art in biomedical text mining or 'BioNLP' in general, focusing primarily on papers published within the past year.
Pierre Zweigenbaum, Dina Demner-Fushman, Hong Yu 0001, Kevin Cohen 0001
Briefings Bioinform.2
2007 Answering Clinical Questions with Knowledge-Based and Statistical Techniques
abstract
The combination of recent developments in question-answering research and the availability of unparalleled resources developed specifically for automatic semantic processing of text in the medical domain provides a unique opportunity to explore complex question answering in the domain of clinical medicine. This article presents a system designed to satisfy the information needs of physicians practicing evidence-based medicine. We have developed a series of knowledge extractors, which employ a combination of knowledge-based and statistical techniques, for automatically identifying clinically relevant aspects of MEDLINE abstracts. These extracted elements serve as the input to an algorithm that scores the relevance of citations with respect to structured representations of information needs, in accordance with the principles of evidence-based medicine. Starting with an initial list of citations retrieved by PubMed, our system can bring relevant abstracts into higher ranking positions, and from these abstracts generate responses that directly answer physicians' questions. We describe three separate evaluations: one focused on the accuracy of the knowledge extractors, one conceptualized as a document reranking task, and finally, an evaluation of answers by two physicians. Experiments on a collection of real-world clinical questions show that our approach significantly outperforms the already competitive PubMed baseline.
Dina Demner-Fushman, Jimmy Lin
Comput. Linguistics1
2007 Research Paper: Using Wireless Handheld Computers to Seek Information at the Point of Care: An Evaluation by Clinicians
abstract
OBJECTIVE: To evaluate: (1) the effectiveness of wireless handheld computers for online information retrieval in clinical settings; (2) the role of MEDLINE in answering clinical questions raised at the point of care. DESIGN: A prospective single-cohort study: accompanying medical teams on teaching rounds, five internal medicine residents used and evaluated MD on Tap, an application for handheld computers, to seek answers in real time to clinical questions arising at the point of care. MEASUREMENTS: All transactions were stored by an intermediate server. Evaluators recorded clinical scenarios and questions, identified MEDLINE citations that answered the questions, and submitted daily and summative reports of their experience. A senior medical librarian corroborated the relevance of the selected citation to each scenario and question. RESULTS: Evaluators answered 68% of 363 background and foreground clinical questions during rounding sessions using a variety of MD on Tap features in an average session length of less than four minutes. The evaluator, the number and quality of query terms, the total number of citations found for a query, and the use of auto-spellcheck significantly contributed to the probability of query success. CONCLUSION: Handheld computers with Internet access are useful tools for healthcare providers to access MEDLINE in real time. MEDLINE citations can answer specific clinical questions when several medical terms are used to form a query. The MD on Tap application is an effective interface to MEDLINE in clinical settings, allowing clinicians to quickly find relevant citations.
Susan E. Hauser, Dina Demner-Fushman, Joshua L. Jacobs, Susanne M. Humphrey, Glenn Ford, George R. Thoma
J. Am. Medical Informatics Assoc.2
2007 Application of Information Technology: Essie: A Concept-based Search Engine for Structured Biomedical Text
abstract
This article describes the algorithms implemented in the Essie search engine that is currently serving several Web sites at the National Library of Medicine. Essie is a phrase-based search engine with term and concept query expansion and probabilistic relevancy ranking. Essie's design is motivated by an observation that query terms are often conceptually related to terms in a document, without actually occurring in the document text. Essie's performance was evaluated using data and standard evaluation methods from the 2003 and 2006 Text REtrieval Conference (TREC) Genomics track. Essie was the best-performing search engine in the 2003 TREC Genomics track and achieved results comparable to those of the highest-ranking systems on the 2006 TREC Genomics track task. Essie shows that a judicious combination of exploiting document structure, phrase searching, and concept based query expansion is a useful approach for information retrieval in the biomedical domain.
Nicholas C. Ide, Russell F. Loane, Dina Demner-Fushman
J. Am. Medical Informatics Assoc.3
2007 Research Paper: Knowledge-based Methods to Help Clinicians Find Answers in MEDLINE
abstract
OBJECTIVES: Large databases of published medical research can support clinical decision making by providing physicians with the best available evidence. The time required to obtain optimal results from these databases using traditional systems often makes accessing the databases impractical for clinicians. This article explores whether a hybrid approach of augmenting traditional information retrieval with knowledge-based methods facilitates finding practical clinical advice in the research literature. DESIGN: Three experimental systems were evaluated for their ability to find MEDLINE citations providing answers to clinical questions of different complexity. The systems (SemRep, Essie, and CQA-1.0), which rely on domain knowledge and semantic processing to varying extents, were evaluated separately and in combination. Fifteen therapy and prevention questions in three categories (general, intermediate, and specific questions) were searched. The first 10 citations retrieved by each system were randomized, anonymized, and evaluated on a three-point scale. The reasons for ratings were documented. MEASUREMENTS: Metrics evaluating the overall performance of a system (mean average precision, binary preference) and metrics evaluating the number of relevant documents in the first several presented to a physician were used. RESULTS: Scores (mean average precision = 0.57, binary preference = 0.71) for fusion of the retrieval results of the three systems are significantly (p < 0.01) better than those for any individual system. All three systems present three to four relevant citations in the first five for any question type. CONCLUSION: The improvements in finding relevant MEDLINE citations due to knowledge-based processing show promise in assisting physicians to answer questions in clinical practice.
Charles Sneiderman, Dina Demner-Fushman, Marcelo Fiszman, Nicholas C. Ide, Thomas C. Rindflesch
J. Am. Medical Informatics Assoc.2
2006 Answer Extraction, Semantic Clustering, and Extractive Summarization for Clinical Question Answering
abstract
This paper presents a hybrid approach to question answering in the clinical domain that combines techniques from summarization and information retrieval. We tackle a frequently-occurring class of questions that takes the form "What is the best drug treatment for X?" Starting from an initial set of MEDLINE citations, our system first identifies the drugs under study. Abstracts are then clustered using semantic classes from the UMLS ontology. Finally, a short extractive summary is generated for each abstract to populate the clusters. Two evaluations---a manual one focused on short answers and an automatic one focused on the supporting abstract---demonstrate that our system compares favorably to PubMed, the search system most widely used by physicians today.
Dina Demner-Fushman, Jimmy Lin
ACL1
2006 MEDLINE as a Source of Just-in-Time Answers to Clinical Questions
Dina Demner-Fushman, Susan E. Hauser, Susanne M. Humphrey, Glenn Ford, Joshua L. Jacobs, George R. Thoma
AMIA1
2006 Preliminary Comparison of Three Search Engines for Point of Care Access to MEDLINE® Citations
Susan E. Hauser, Dina Demner-Fushman, Glenn Ford, Joshua L. Jacobs, George R. Thoma
AMIA2
2006 Evaluation of PICO as a Knowledge Representation for Clinical Questions
Jimmy Lin, Dina Demner-Fushman
AMIA3
2006 Semantic Processing to Enhance Retrieval of Diagnosis Citations from Medline
Charles Sneiderman, Dina Demner-Fushman, Marcelo Fiszman, Graciela Rosemblat, François-Michel Lang, Daphne Norwood, Thomas C. Rindflesch
AMIA2
2006 Will Pyramids Built of Nuggets Topple Over?
Jimmy Lin, Dina Demner-Fushman
HLT-NAACL2
2006 The role of knowledge in conceptual retrieval: a study in the domain of clinical medicine
abstract
Despite its intuitive appeal, the hypothesis that retrieval at the level of "concepts" should outperform purely term-based approaches remains unverified empirically. In addition, the use of "knowledge" has not consistently resulted in performance gains. After identifying possible reasons for previous negative results, we present a novel framework for "conceptual retrieval" that articulates the types of knowledge that are important for information seeking. We instantiate this general framework in the domain of clinical medicine based on the principles of evidence-based medicine (EBM). Experiments show that an EBM-based scoring algorithm dramatically outperforms a state-of-the-art baseline that employs only term statistics. Ablation studies further yield a better understanding of the performance contributions of different components. Finally, we discuss how other domains can benefit from knowledge-based approaches.
Jimmy Lin, Dina Demner-Fushman
SIGIR2
2006 Exploring the limits of single-iteration clarification dialogs
abstract
Single-iteration clarification dialogs, as implemented in the TREC HARD track, represent an attempt to introduce interaction into ad hoc retrieval, while preserving the many benefits of large-scale evaluations. Although previous experiments have not conclusively demonstrated performance gains resulting from such interactions, it is unclear whether these findings speak to the nature of clarification dialogs, or simply the limitations of current systems. To probe the limits of such interactions, we employed a human intermediary to formulate clarification questions and exploit user responses. In addition to establishing a plausible upper bound on performance, we were also able to induce an "ontology of clarifications" to characterize human behavior. This ontology, in turn, serves as the input to a regression model that attempts to determine which types of clarification questions are most helpful. Our work can serve to inform the design of interactive systems that initiate user dialogs.
Jimmy Lin, Philip Fei Wu, Dina Demner-Fushman, Eileen G. Abels
SIGIR3
2006 Methods for automatically evaluating answers to complex questions
Jimmy Lin, Dina Demner-Fushman
Inf. Retr.2
2006 Research Paper: Automatically Identifying Health Outcome Information in MEDLINE Records
abstract
OBJECTIVE: Understanding the effect of a given intervention on the patient's health outcome is one of the key elements in providing optimal patient care. This study presents a methodology for automatic identification of outcomes-related information in medical text and evaluates its potential in satisfying clinical information needs related to health care outcomes. DESIGN: An annotation scheme based on an evidence-based medicine model for critical appraisal of evidence was developed and used to annotate 633 MEDLINE citations. Textual, structural, and meta-information features essential to outcome identification were learned from the created collection and used to develop an automatic system. Accuracy of automatic outcome identification was assessed in an intrinsic evaluation and in an extrinsic evaluation, in which ranking of MEDLINE search results obtained using PubMed Clinical Queries relied on identified outcome statements. MEASUREMENTS: The accuracy and positive predictive value of outcome identification were calculated. Effectiveness of the outcome-based ranking was measured using mean average precision and precision at rank 10. RESULTS: Automatic outcome identification achieved 88% to 93% accuracy. The positive predictive value of individual sentences identified as outcomes ranged from 30% to 37%. Outcome-based ranking improved retrieval accuracy, tripling mean average precision and achieving 389% improvement in precision at rank 10. CONCLUSION: Preliminary results in outcome-based document ranking show potential validity of the evidence-based medicine-model approach in timely delivery of information critical to clinical decision support at the point of service.
Dina Demner-Fushman, Barbara Few, Susan E. Hauser, George R. Thoma
J. Am. Medical Informatics Assoc.1
2006 Word sense disambiguation by selecting the best semantic type based on Journal Descriptor Indexing: Preliminary experiment
abstract
An experiment was performed at the National Library of Medicine((R)) (NLM((R))) in word sense disambiguation (WSD) using the Journal Descriptor Indexing (JDI) methodology. The motivation is the need to solve the ambiguity problem confronting NLM's MetaMap system, which maps free text to terms corresponding to concepts in NLM's Unified Medical Language System((R)) (UMLS((R))) Metathesaurus((R)). If the text maps to more than one Metathesaurus concept at the same high confidence score, MetaMap has no way of knowing which concept is the correct mapping. We describe the JDI methodology, which is ultimately based on statistical associations between words in a training set of MEDLINE((R)) citations and a small set of journal descriptors (assigned by humans to journals per se) assumed to be inherited by the citations. JDI is the basis for selecting the best meaning that is correlated to UMLS semantic types (STs) assigned to ambiguous concepts in the Metathesaurus. For example, the ambiguity transport has two meanings: "Biological Transport" assigned the ST Cell Function and "Patient transport" assigned the ST Health Care Activity. A JDI-based methodology can analyze text containing transport and determine which ST receives a higher score for that text, which then returns the associated meaning, presumed to apply to the ambiguity itself. We then present an experiment in which a baseline disambiguation method was compared to four versions of JDI in disambiguating 45 ambiguous strings from NLM's WSD Test Collection. Overall average precision for the highest-scoring JDI version was 0.7873 compared to 0.2492 for the baseline method, and average precision for individual ambiguities was greater than 0.90 for 23 of them (51%), greater than 0.85 for 24 (53%), and greater than 0.65 for 35 (79%). On the basis of these results, we hope to improve performance of JDI and test its use in applications.
Susanne M. Humphrey, Willie J. Rogers, Halil Kilicoglu, Dina Demner-Fushman, Thomas C. Rindflesch
J. Assoc. Inf. Sci. Technol.4
2005 The Role of Title, Metadata and Abstract in Identifying Clinically Relevant Journal Articles
Dina Demner-Fushman, Susan E. Hauser, George R. Thoma
AMIA1
2005 "Bag of Words" is not enough for Strength of Evidence Classification
Jimmy Lin, Dina Demner-Fushman
AMIA2
2005 Semantic characteristics of MEDLINE citations useful for therapeutic decision-making
Charles Sneiderman, Dina Demner-Fushman, Marcelo Fiszman, Thomas C. Rindflesch
AMIA2
2004 A Testbed System for Mobile Point-of-Care Information Delivery
abstract
PubMed on Tap is a testbed system that supports search of and retrieval from the National Library of Medicine's MEDLINE/spl reg/ database from a PDA. The goal of the PubMed on Tap project is to discover and implement design principles for point-of-care delivery of clinical support information. The project explores user interface issues, information content and organization, and system performance. Here we present our progress in these areas.
Susan E. Hauser, Dina Demner-Fushman, Glenn Ford, George R. Thoma
CBMS2
2003 Making MIRACLEs: Interactive translingual search for Cebuano and Hindi
abstract
Searching is inherently a user-centered process; people pose the questions for which machines seek answers, and ultimately people judge the degree to which retrieved documents meet their needs. Rapid development of interactive systems that use queries expressed in one language to search documents written in another poses five key challenges: (1) interaction design, (2) query formulation, (3) cross-language search, (4) construction of translated summaries, and (5) machine translation. This article describes the design of MIRACLE, an easily extensible system based on English queries that has previously been used to search French, German, and Spanish documents, and explains how the capabilities of MIRACLE were rapidly extended to accommodate Cebuano and Hindi. Evaluation results for the cross-language search component are presented for both languages, along with results from a brief full-system interactive experiment with Hindi. The article concludes with some observations on directions for further research on interactive cross-language information retrieval.
Daqing He, Douglas W. Oard, Jianqiang Wang 0002, Dina Demner-Fushman, Kareem Darwish, Philip Resnik, Sanjeev Khudanpur, Michael Nossal, Michael Subotin, Anton Leuski
ACM Trans. Asian Lang. Inf. Process.5