VLDB 2026 Research / reviewers in the wild / expert
Kirk Roberts
dblp:59/1629 · also Kirk E. Roberts
· DBLP profile ↗
78ranked-venue papers
25as first author
23since 2021 · last 2026
0000-0001-6525-5213ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 56 · 16 first-author · 18 since 2021Artificial intelligence and machine learning · 19 · 9 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Information extraction from clinical notes: are we ready to switch to large language models?abstractOBJECTIVES: To assess the performance, generalizability, and computational efficiency of instruction-tuned Large Language Model Meta AI (LLaMA)-2 and LLaMA-3 models compared to bidirectional encoder representations from transformers (BERT) for clinical information extraction (IE) tasks, specifically named entity recognition (NER) and relation extraction (RE). MATERIALS AND METHODS: We developed a comprehensive annotated corpus of 1588 clinical notes from 4 data sources-UT Physicians (UTP) (1342 notes), Transcribed Medical Transcription Sample Reports and Examples (MTSamples) (146), Medical Information Mart for Intensive Care (MIMIC)-III (50), and Informatics for Integrating Biology and the Bedside (i2b2) (50), capturing 4 clinical entities (problems, tests, medications, other treatments) and 16 modifiers (eg, negation, certainty). Large Language Model Meta AI-2 and LLaMA-3 were instruction-tuned for clinical NER and RE, and their performance was benchmarked against BERT. RESULTS: Large Language Model Meta AI models consistently outperformed BERT across datasets. In data-rich settings (eg, UTP), LLaMA achieved marginal gains (approximately 1% improvement for NER and 1.5%-3.7% for RE). Under limited data conditions (eg, MTSamples, MIMIC-III) and on the unseen i2b2 dataset, LLaMA-3-70B improved F1 scores by over 7% for NER and 4% for RE. However, performance gains came with increased computational costs, with LLaMA models requiring more memory and Graphics Processing Unit (GPU) hours and running up to 28 times slower than BERT. DISCUSSION: While LLaMA models offer enhanced performance, their higher computational demands and slower throughput highlight the need to balance performance with practical resource constraints. Application-specific considerations are essential when choosing between LLMs and BERT for clinical IE. CONCLUSION: Instruction-tuned LLaMA models show promise for clinical NER and RE tasks. However, the tradeoff between improved performance and increased computational cost must be carefully evaluated. We release our Kiwi package (https://kiwi.clinicalnlp.org/) to facilitate the application of both LLaMA and BERT models in clinical IE applications. Xu Zuo, Yujia Zhou 0003, Xueqing Peng, Jimin Huang, Vipina Kuttichi Keloth, Vincent J. Zhang, Ruey-Ling Weng, Cathy Shyr, Qingyu Chen 0001, Xiaoqian Jiang, Kirk Roberts, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 12 |
| 2026 | Clinical document metadata extraction: A scoping reviewabstractOBJECTIVES: Clinical document metadata, such as document type, structure, author role, medical specialty, and encounter setting, is essential for accurate interpretation of information captured in clinical documents. However, vast documentation heterogeneity and drift over time challenge harmonization of document metadata. Automated extraction methods have emerged to coalesce metadata from disparate practices into target schema. This scoping review aims to catalog research on clinical document metadata extraction, identify methodological trends and applications, and highlight gaps warranting further investigation. METHODS: We followed the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews) guidelines to identify articles from Ovid MEDLINE, Ovid EMBASE, Scopus, Web of Science and external sources that perform clinical document metadata extraction, either primarily as a methodology study, secondarily as a feature for a downstream application, or for analysis. We initially identified and screened 342 articles published between 2011 and 2025, then comprehensively reviewed 77 we deemed relevant to our study. RESULTS: Among the 77 articles included in our full text review, 49 were methodological, 22 used document metadata as features in a downstream application, and 6 analyzed document metadata composition. We observe myriad purposes for methodological study and application types. Available labelled public data remains sparse except for structural section datasets. Methods for extracting document metadata have progressed from largely rule-based and traditional machine learning with ample feature engineering to transformer-based architectures with minimal feature engineering. DISCUSSION AND CONCLUSION: Clinical document metadata extraction research has accelerated over recent years. The emergence of large language models has enabled broader exploration of generalizability across tasks and datasets, allowing the possibility of advanced clinical text processing systems. We anticipate that research will continue to expand into richer document metadata representations and integrate further into clinical applications and workflows. Kurt Miller, Qiuhao Lu, William R. Hersh, Kirk Roberts, Steven Bedrick, Andrew Wen |
J. Biomed. Informatics | 4 |
| 2025 | Quantitatively assessing the impact of the quality of SNOMED CT subtype hierarchy on cohort queriesabstractOBJECTIVE: SNOMED CT provides a standardized terminology for clinical concepts, allowing cohort queries over heterogeneous clinical data including Electronic Health Records (EHRs). While it is intuitive that missing and inaccurate subtype (or is-a) relations in SNOMED CT reduce the recall and precision of cohort queries, the extent of these impacts has not been formally assessed. This study fills this gap by developing quantitative metrics to measure these impacts and performing statistical analysis on their significance. MATERIAL AND METHODS: We used the Optum de-identified COVID-19 Electronic Health Record dataset. We defined micro-averaged and macro-averaged recall and precision metrics to assess the impact of missing and inaccurate is-a relations on cohort queries. Both practical and simulated analyses were performed. Practical analyses involved 407 missing and 48 inaccurate is-a relations confirmed by domain experts, with statistical testing using Wilcoxon signed-rank tests. Simulated analyses used two random sets of 400 is-a relations to simulate missing and inaccurate is-a relations. RESULTS: Wilcoxon signed-rank tests from both practical and simulated analyses (P-values < .001) showed that missing is-a relations significantly reduced the micro- and macro-averaged recall, and inaccurate is-a relations significantly reduced the micro- and macro-averaged precision. DISCUSSION: The introduced impact metrics can assist SNOMED CT maintainers in prioritizing critical hierarchical defects for quality enhancement. These metrics are generally applicable for assessing the quality impact of a terminology's subtype hierarchy on its cohort query applications. CONCLUSION: Our results indicate a significant impact of missing and inaccurate is-a relations in SNOMED CT on the recall and precision of cohort queries. Our work highlights the importance of high-quality terminology hierarchy for cohort queries over EHR data and provides valuable insights for prioritizing quality improvements of SNOMED CT's hierarchy. Xubing Hao, Yan Huang 0034, Jay Shi, Rashmie Abeysinghe, Cui Tao, Kirk Roberts, Guo-Qiang Zhang 0001, Licong Cui |
J. Am. Medical Informatics Assoc. | 7 |
| 2025 | Dynamic few-shot prompting for clinical note section classification using lightweight, open-source large language modelsabstractOBJECTIVE: Unlocking clinical information embedded in clinical notes has been hindered to a significant degree by domain-specific and context-sensitive language. Identification of note sections and structural document elements has been shown to improve information extraction and dependent downstream clinical natural language processing (NLP) tasks and applications. This study investigates the viability of a dynamic example selection prompting method to section classification using lightweight, open-source large language models (LLMs) as a practical solution for real-world healthcare clinical NLP systems. MATERIALS AND METHODS: We develop a dynamic few-shot prompting approach to classifying sections where section samples are first embedded using a transformer-based model and deposited in a vector store. During inference, the embedded samples with the most similar contextual embeddings to a given input section text are retrieved from the vector store and inserted into the LLM prompt. We evaluate this technique on two datasets comprising two section schemas, including varying levels of context. We compare the performance to baseline zero-shot and randomly selected few-shot scenarios. RESULTS: The dynamic few-shot prompting experiments yielded the highest F1 scores in each of the classification tasks and datasets for all seven of the LLMs included in the evaluation, averaging a macro F1 increase of 39.3% and 21.1% in our primary section classification task over the zero-shot and static few-shot baselines, respectively. DISCUSSION AND CONCLUSION: Our results showcase substantial performance improvements imparted by dynamically selecting examples for few-shot LLM prompting, and further improvement by including section context, demonstrating compelling potential for clinical applications. Kurt Miller, Steven Bedrick, Qiuhao Lu, Andrew Wen, William R. Hersh, Kirk Roberts |
J. Am. Medical Informatics Assoc. | 6 |
| 2025 | A comparative study of recent large language models on generating hospital discharge summaries for lung cancer patients
Fang Li 0011, Na Hong, Manqi Li, Kirk Roberts, Licong Cui, Cui Tao, Hua Xu 0001 |
J. Biomed. Informatics | 5 |
| 2024 | Improving large language models for clinical named entity recognition via prompt engineeringabstractIMPORTANCE: The study highlights the potential of large language models, specifically GPT-3.5 and GPT-4, in processing complex clinical data and extracting meaningful information with minimal training data. By developing and refining prompt-based strategies, we can significantly enhance the models' performance, making them viable tools for clinical NER tasks and possibly reducing the reliance on extensive annotated datasets. OBJECTIVES: This study quantifies the capabilities of GPT-3.5 and GPT-4 for clinical named entity recognition (NER) tasks and proposes task-specific prompts to improve their performance. MATERIALS AND METHODS: We evaluated these models on 2 clinical NER tasks: (1) to extract medical problems, treatments, and tests from clinical notes in the MTSamples corpus, following the 2010 i2b2 concept extraction shared task, and (2) to identify nervous system disorder-related adverse events from safety reports in the vaccine adverse event reporting system (VAERS). To improve the GPT models' performance, we developed a clinical task-specific prompt framework that includes (1) baseline prompts with task description and format specification, (2) annotation guideline-based prompts, (3) error analysis-based instructions, and (4) annotated samples for few-shot learning. We assessed each prompt's effectiveness and compared the models to BioClinicalBERT. RESULTS: Using baseline prompts, GPT-3.5 and GPT-4 achieved relaxed F1 scores of 0.634, 0.804 for MTSamples and 0.301, 0.593 for VAERS. Additional prompt components consistently improved model performance. When all 4 components were used, GPT-3.5 and GPT-4 achieved relaxed F1 socres of 0.794, 0.861 for MTSamples and 0.676, 0.736 for VAERS, demonstrating the effectiveness of our prompt framework. Although these results trail BioClinicalBERT (F1 of 0.901 for the MTSamples dataset and 0.802 for the VAERS), it is very promising considering few training samples are needed. DISCUSSION: The study's findings suggest a promising direction in leveraging LLMs for clinical NER tasks. However, while the performance of GPT models improved with task-specific prompts, there's a need for further development and refinement. LLMs like GPT-4 show potential in achieving close performance to state-of-the-art models like BioClinicalBERT, but they still require careful prompt engineering and understanding of task-specific knowledge. The study also underscores the importance of evaluation schemas that accurately reflect the capabilities and performance of LLMs in clinical settings. CONCLUSION: While direct application of GPT models to clinical NER tasks falls short of optimal performance, our task-specific prompt framework, incorporating medical knowledge and training samples, significantly enhances GPT models' feasibility for potential clinical applications. Qingyu Chen 0001, Jingcheng Du, Xueqing Peng, Vipina Kuttichi Keloth, Xu Zuo, Yujia Zhou 0003, Zehan Li, Xiaoqian Jiang, Zhiyong Lu, Kirk Roberts, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 11 |
| 2024 | Ensemble pretrained language models to extract biomedical knowledge from literatureabstractOBJECTIVES: The rapid expansion of biomedical literature necessitates automated techniques to discern relationships between biomedical concepts from extensive free text. Such techniques facilitate the development of detailed knowledge bases and highlight research deficiencies. The LitCoin Natural Language Processing (NLP) challenge, organized by the National Center for Advancing Translational Science, aims to evaluate such potential and provides a manually annotated corpus for methodology development and benchmarking. MATERIALS AND METHODS: For the named entity recognition (NER) task, we utilized ensemble learning to merge predictions from three domain-specific models, namely BioBERT, PubMedBERT, and BioM-ELECTRA, devised a rule-driven detection method for cell line and taxonomy names and annotated 70 more abstracts as additional corpus. We further finetuned the T0pp model, with 11 billion parameters, to boost the performance on relation extraction and leveraged entites' location information (eg, title, background) to enhance novelty prediction performance in relation extraction (RE). RESULTS: Our pioneering NLP system designed for this challenge secured first place in Phase I-NER and second place in Phase II-relation extraction and novelty prediction, outpacing over 200 teams. We tested OpenAI ChatGPT 3.5 and ChatGPT 4 in a Zero-Shot setting using the same test set, revealing that our finetuned model considerably surpasses these broad-spectrum large language models. DISCUSSION AND CONCLUSION: Our outcomes depict a robust NLP system excelling in NER and RE across various biomedical entities, emphasizing that task-specific models remain superior to generic large ones. Such insights are valuable for endeavors like knowledge graph development and hypothesis formulation in biomedical research. Qiang Wei 0002, Liang-Chin Huang, Jianfu Li, Yao-Shun Chuang, Jianping He 0002, Avisha Das, Vipina Kuttichi Keloth, Yuntao Yang, Chiamaka S. Diala, Kirk Roberts, Cui Tao, Xiaoqian Jiang, W. Jim Zheng, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 12 |
| 2023 | quEHRy: a question answering system to query electronic health recordsabstractOBJECTIVE: We propose a system, quEHRy, to retrieve precise, interpretable answers to natural language questions from structured data in electronic health records (EHRs). MATERIALS AND METHODS: We develop/synthesize the main components of quEHRy: concept normalization (MetaMap), time frame classification (new), semantic parsing (existing), visualization with question understanding (new), and query module for FHIR mapping/processing (new). We evaluate quEHRy on 2 clinical question answering (QA) datasets. We evaluate each component separately as well as holistically to gain deeper insights. We also conduct a thorough error analysis for a crucial subcomponent, medical concept normalization. RESULTS: Using gold concepts, the precision of quEHRy is 98.33% and 90.91% for the 2 datasets, while the overall accuracy was 97.41% and 87.75%. Precision was 94.03% and 87.79% even after employing an automated medical concept extraction system (MetaMap). Most incorrectly predicted medical concepts were broader in nature than gold-annotated concepts (representative of the ones present in EHRs), eg, Diabetes versus Diabetes Mellitus, Non-Insulin-Dependent. DISCUSSION: The primary performance barrier to deployment of the system is due to errors in medical concept extraction (a component not studied in this article), which affects the downstream generation of correct logical structures. This indicates the need to build QA-specific clinical concept normalizers that understand EHR context to extract the "relevant" medical concepts from questions. CONCLUSION: We present an end-to-end QA system that allows information access from EHRs using natural language and returns an exact, verifiable answer. Our proposed system is high-precision and interpretable, checking off the requirements for clinical use. Sarvesh Soni, Surabhi Datta, Kirk Roberts |
J. Am. Medical Informatics Assoc. | 3 |
| 2023 | Discerning conversational context in online health communities for personalized digital behavior change solutions using Pragmatics to Reveal Intent in Social Media (PRISM) frameworkabstractBACKGROUND: Online health communities (OHCs) have emerged as prominent platforms for behavior modification, and the digitization of online peer interactions has afforded researchers with unique opportunities to model multilevel mechanisms that drive behavior change. Existing studies, however, have been limited by a lack of methods that allow the capture of conversational context and socio-behavioral dynamics at scale, as manifested in these digital platforms. OBJECTIVE: We develop, evaluate, and apply a novel methodological framework, Pragmatics to Reveal Intent in Social Media (PRISM), to facilitate granular characterization of peer interactions by combining multidimensional facets of human communication. METHODS: We developed and applied PRISM to analyze peer interactions (N = 2.23 million) in QuitNet, an OHC for tobacco cessation. First, we generated a labeled set of peer interactions (n = 2,005) through manual annotation along three dimensions: communication themes (CTs), behavior change techniques (BCTs), and speech acts (SAs). Second, we used deep learning models to apply our qualitative codes at scale. Third, we applied our validated model to perform a retrospective analysis. Finally, using social network analysis (SNA), we portrayed large-scale patterns and relationships among the aforementioned communication dimensions embedded in peer interactions in QuitNet. RESULTS: Qualitative analysis showed that the themes of social support and behavioral progress were common. The most used BCTs were feedback and monitoring and comparison of behavior, and users most commonly expressed their intentions using SAs-expressive and emotion. With additional in-domain pre-training, bidirectional encoder representations from Transformers (BERT) outperformed other deep learning models on the classification tasks. Content-specific SNA revealed that users' engagement or abstinence status is associated with the prevalence of various categories of BCTs and SAs, which also was evident from the visualization of network structures. CONCLUSIONS: Our study describes the interplay of multilevel characteristics of online communication and their association with individual health behaviors. Tavleen Singh, Kirk Roberts, Trevor Cohen, Nathan K. Cobb, Amy Franklin, Sahiti Myneni |
J. Biomed. Informatics | 2 |
| 2023 | Scholarly recommendation systems: a literature surveyabstractAbstract A scholarly recommendation system is an important tool for identifying prior and related resources such as literature, datasets, grants, and collaborators. A well-designed scholarly recommender significantly saves the time of researchers and can provide information that would not otherwise be considered. The usefulness of scholarly recommendations, especially literature recommendations, has been established by the widespread acceptance of web search engines such as CiteSeerX, Google Scholar, and Semantic Scholar. This article discusses different aspects and developments of scholarly recommendation systems. We searched the ACM Digital Library, DBLP, IEEE Explorer, and Scopus for publications in the domain of scholarly recommendations for literature, collaborators, reviewers, conferences and journals, datasets, and grant funding. In total, 225 publications were identified in these areas. We discuss methodologies used to develop scholarly recommender systems. Content-based filtering is the most commonly applied technique, whereas collaborative filtering is more popular among conference recommenders. The implementation of deep learning algorithms in scholarly recommendation systems is rare among the screened publications. We found fewer publications in the areas of the dataset and grant funding recommenders than in other areas. Furthermore, studies analyzing users’ feedback to improve scholarly recommendation systems are rare for recommenders. This survey provides background knowledge regarding existing research on scholarly recommenders and aids in developing future recommendation systems in this domain. Braja Gopal Patra, Ashraf Yaseen, Rachit Sabharwal, Kirk Roberts, Tru Cao, Hulin Wu |
Knowl. Inf. Syst. | 6 |
| 2022 | Eye-SpatialNet: Spatial Information Extraction from Ophthalmology Notes
Surabhi Datta, Tasneem Kaochar, Hio Cheng Lam, Nelly Nwosu, Alice Z. Chuang, Robert M. Feldman, Kirk Roberts |
AMIA | 7 |
| 2022 | Toward a Neural Semantic Parsing System for EHR Question Answering
Sarvesh Soni, Kirk Roberts |
AMIA | 2 |
| 2022 | DrugEHRQA: A Question Answering Dataset on Structured and Unstructured Electronic Health Records For Medicine Related QueriesabstractThis paper develops the first question answering dataset (DrugEHRQA) containing question-answer pairs from both structured tables and unstructured notes from a publicly available Electronic Health Record (EHR). EHRs contain patient records, stored in structured tables and unstructured clinical notes. The information in structured and unstructured EHRs is not strictly disjoint: information may be duplicated, contradictory, or provide additional context between these sources. Our dataset has medication-related queries, containing over 70,000 question-answer pairs. To provide a baseline model and help analyze the dataset, we have used a simple model (MultimodalEHRQA) which uses the predictions of a modality selection network to choose between EHR tables and clinical notes to answer the questions. This is used to direct the questions to the table-based or text-based state-of-the-art QA model. In order to address the problem arising from complex, nested queries, this is the first time Relation-Aware Schema Encoding and Linking for Text-to-SQL Parsers (RAT-SQL) has been used to test the structure of query templates in EHR data. Our goal is to provide a benchmark dataset for multi-modal QA systems, and to open up new avenues of research in improving question answering over EHR structured data by using context from unstructured clinical data. Jayetri Bardhan, Anthony M. Colas, Kirk Roberts, Daisy Zhe Wang |
LREC | 3 |
| 2022 | A Cross-document Coreference Dataset for Longitudinal Tracking across Radiology ReportsabstractThis paper proposes a new cross-document coreference resolution (CDCR) dataset for identifying co-referring radiological findings and medical devices across a patient’s radiology reports. Our annotated corpus contains 5872 mentions (findings and devices) spanning 638 MIMIC-III radiology reports across 60 patients, covering multiple imaging modalities and anatomies. There are a total of 2292 mention chains. We describe the annotation process in detail, highlighting the complexities involved in creating a sizable and realistic dataset for radiology CDCR. We apply two baseline methods–string matching and transformer language models (BERT)–to identify cross-report coreferences. Our results indicate the requirement of further model development targeting better understanding of domain language and context to address this challenging and unexplored task. This dataset can serve as a resource to develop more advanced natural language processing CDCR methods in the future. This is one of the first attempts focusing on CDCR in the clinical domain and holds potential in benefiting physicians and clinical research through long-term tracking of radiology findings. Surabhi Datta, Hio Cheng Lam, Atieh Pajouhi, Sunitha Mogalla, Kirk Roberts |
LREC | 5 |
| 2022 | RadQA: A Question Answering Dataset to Improve Comprehension of Radiology ReportsabstractWe present a radiology question answering dataset, RadQA, with 3074 questions posed against radiology reports and annotated with their corresponding answer spans (resulting in a total of 6148 question-answer evidence pairs) by physicians. The questions are manually created using the clinical referral section of the reports that take into account the actual information needs of ordering physicians and eliminate bias from seeing the answer context (and, further, organically create unanswerable questions). The answer spans are marked within the Findings and Impressions sections of a report. The dataset aims to satisfy the complex clinical requirements by including complete (yet concise) answer phrases (which are not just entities) that can span multiple lines. We conduct a thorough analysis of the proposed dataset by examining the broad categories of disagreement in annotation (providing insights on the errors made by humans) and the reasoning requirements to answer a question (uncovering the huge dependence on medical knowledge for answering the questions). The advanced transformer language models achieve the best F1 score of 63.55 on the test set, however, the best human performance is 90.31 (with an average of 84.52). This demonstrates the challenging nature of RadQA that leaves ample scope for future method research. Sarvesh Soni, Meghana Gudala, Atieh Pajouhi, Kirk Roberts |
LREC | 4 |
| 2022 | Toward a standard formal semantic representation of the model card reportabstractBACKGROUND: Model card reports aim to provide informative and transparent description of machine learning models to stakeholders. This report document is of interest to the National Institutes of Health's Bridge2AI initiative to address the FAIR challenges with artificial intelligence-based machine learning models for biomedical research. We present our early undertaking in developing an ontology for capturing the conceptual-level information embedded in model card reports. RESULTS: Sourcing from existing ontologies and developing the core framework, we generated the Model Card Report Ontology. Our development efforts yielded an OWL2-based artifact that represents and formalizes model card report information. The current release of this ontology utilizes standard concepts and properties from OBO Foundry ontologies. Also, the software reasoner indicated no logical inconsistencies with the ontology. With sample model cards of machine learning models for bioinformatics research (HIV social networks and adverse outcome prediction for stent implantation), we showed the coverage and usefulness of our model in transforming static model card reports to a computable format for machine-based processing. CONCLUSIONS: The benefit of our work is that it utilizes expansive and standard terminologies and scientific rigor promoted by biomedical ontologists, as well as, generating an avenue to make model cards machine-readable using semantic web technology. Our future goal is to assess the veracity of our model and later expand the model to include additional concepts to address terminological gaps. We discuss tools and software that will utilize our ontology for potential application services. Muhammad Amith, Licong Cui, Degui Zhi, Kirk Roberts, Xiaoqian Jiang, Fang Li 0011, Evan Yu, Cui Tao |
BMC Bioinform. | 4 |
| 2022 | Closing the loop: automatically identifying abnormal imaging results in scanned documentsabstractOBJECTIVES: Scanned documents (SDs), while common in electronic health records and potentially rich in clinically relevant information, rarely fit well with clinician workflow. Here, we identify scanned imaging reports requiring follow-up with high recall and practically useful precision. MATERIALS AND METHODS: We focused on identifying imaging findings for 3 common causes of malpractice claims: (1) potentially malignant breast (mammography) and (2) lung (chest computed tomography [CT]) lesions and (3) long-bone fracture (X-ray) reports. We train our ClinicalBERT-based pipeline on existing typed/dictated reports classified manually or using ICD-10 codes, evaluate using a test set of manually classified SDs, and compare against string-matching (baseline approach). RESULTS: A total of 393 mammograms, 305 chest CT, and 683 bone X-ray reports were manually reviewed. The string-matching approach had an F1 of 0.667. For mammograms, chest CTs, and bone X-rays, respectively: models trained on manually classified training data and optimized for F1 reached an F1 of 0.900, 0.905, and 0.817, while separate models optimized for recall achieved a recall of 1.000 with precisions of 0.727, 0.518, and 0.275. Models trained on ICD-10-labelled data and optimized for F1 achieved F1 scores of 0.647, 0.830, and 0.643, while those optimized for recall achieved a recall of 1.0 with precisions of 0.407, 0.683, and 0.358. DISCUSSION: Our pipeline can identify abnormal reports with potentially useful performance and so decrease the manual effort required to screen for abnormal findings that require follow-up. CONCLUSION: It is possible to automatically identify clinically significant abnormalities in SDs with high recall and practically useful precision in a generalizable and minimally laborious way. Akshat Kumar, Heath Goodrum, Ashley Kim, Carly Stender, Kirk Roberts, Elmer V. Bernstam |
J. Am. Medical Informatics Assoc. | 5 |
| 2021 | Dental EHR-infused Persona Ontologies to Enrich Dental Dialogue Interaction of AgentsabstractThe quality of patient-provider communication can predict the healthcare outcomes in patients, and therefore, training dental providers to handle the communication effort with patients is crucial. In our previous work, we developed an ontology model that can standardize and represent patient-provider communication, which can later be integrated in conversational agents as tools for dental communication training. In this study, we embark on enriching our previous model with an ontology of patient personas to portray and express types of dental patient archetypes. The Ontology of Patient Personas that we developed was rooted in terminologies from an OBO Foundry ontology and dental electronic health record data elements. We discuss how this ontology aims to enhance the aforementioned dialogue ontology and future direction in executing our model in software agents to train dental students. Patricia Ngantcha, Muhammad Amith, Kirk Roberts, John A. Valenza, Muhammad F. Walji, Cui Tao |
BIBM | 3 |
| 2021 | On the Quality of the TREC-COVID IR Test CollectionsabstractShared text collections continue to be vital infrastructure for IR research. The COVID-19 pandemic offered an opportunity to create a test collection that captured the rapidly changing information space during a pandemic, and the TREC-COVID effort was created to build such a collection using the TREC framework. This paper examines the quality of the resulting TREC-COVID test collections, and in doing so, offers a critique of the state-of-the-art in building reusable IR test collections. The largest of the collections--called 'TREC-COVID Complete'--is found to be on par with previous TREC ad~hoc collections with existing quality tests uncovering no apparent problems. Yet the lack of any way to definitively demonstrate the collection's quality and its violation of previously used quality heuristics suggest much work remains to be done to understand the factors affecting collection quality. Ellen M. Voorhees, Kirk Roberts |
SIGIR | 2 |
| 2021 | An evaluation of two commercial deep learning-based information retrieval systems for COVID-19 literatureabstractThe COVID-19 pandemic has resulted in a tremendous need for access to the latest scientific information, leading to both corpora for COVID-19 literature and search engines to query such data. While most search engine research is performed in academia with rigorous evaluation, major commercial companies dominate the web search market. Thus, it is expected that commercial pandemic-specific search engines will gain much higher traction than academic alternatives, leading to questions about the empirical performance of these tools. This paper seeks to empirically evaluate two commercial search engines for COVID-19 (Google and Amazon) in comparison with academic prototypes evaluated in the TREC-COVID task. We performed several steps to reduce bias in the manual judgments to ensure a fair comparison of all systems. We find the commercial search engines sizably underperformed those evaluated under TREC-COVID. This has implications for trust in popular health search engines and developing biomedical search engines for future health crises. Sarvesh Soni, Kirk Roberts |
J. Am. Medical Informatics Assoc. | 2 |
| 2021 | Searching for scientific evidence in a pandemic: An overview of TREC-COVIDabstractWe present an overview of the TREC-COVID Challenge, an information retrieval (IR) shared task to evaluate search on scientific literature related to COVID-19. The goals of TREC-COVID include the construction of a pandemic search test collection and the evaluation of IR methods for COVID-19. The challenge was conducted over five rounds from April to July 2020, with participation from 92 unique teams and 556 individual submissions. A total of 50 topics (sets of related queries) were used in the evaluation, starting at 30 topics for Round 1 and adding 5 new topics per round to target emerging topics at that state of the still-emerging pandemic. This paper provides a comprehensive overview of the structure and results of TREC-COVID. Specifically, the paper provides details on the background, task structure, topic structure, corpus, participation, pooling, assessment, judgments, results, top-performing systems, lessons learned, and benchmark datasets. Kirk Roberts, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, Kyle Lo, Ian Soboroff, Ellen M. Voorhees, Lucy Lu Wang, William R. Hersh |
J. Biomed. Informatics | 1 |
| 2021 | Generalized and transferable patient language representation for phenotyping with limited data
Yuqi Si, Elmer V. Bernstam, Kirk Roberts |
J. Biomed. Informatics | 3 |
| 2021 | Deep representation learning of patient data from Electronic Health Records (EHR): A systematic review
Yuqi Si, Jingcheng Du, Xiaoqian Jiang, Timothy A. Miller, Fei Wang 0001, W. Jim Zheng, Kirk Roberts |
J. Biomed. Informatics | 8 |
| 2020 | RadLex Normalization in Radiology Reports
Surabhi Datta, Jordan Godfrey-Stovall, Kirk Roberts |
AMIA | 3 |
| 2020 | Deep Representation Learning of Patient Data from Electronic Health Records: A Systematic Review
Yuqi Si, Jingcheng Du, Xiaoqian Jiang, Timothy A. Miller, Fei Wang 0001, W. Jim Zheng, Kirk Roberts |
AMIA | 8 |
| 2020 | Patient Cohort Retrieval using Transformer Language Models
Sarvesh Soni, Kirk Roberts |
AMIA | 2 |
| 2020 | A health consumer ontology of fast food informationabstractA variety of severe health issues can be attributed to poor nutrition and poor eating behaviors. Research has explored the impact of nutritional knowledge on an individual's inclination to purchase and consume certain foods. This paper introduces the Ontology of Fast Food Facts, a knowledge base that models consumer nutritional data from major fast food establishments. This artifact serves as an aggregate knowledge base to centralize nutritional information for consumers. As a semantically-linked data source, the Ontology of Fast Food Facts could engender methods and tools to further the research and impact the health consumers' diet and behavior, which is a factor in many severe health outcomes. We describe the initial development of this ontology and future directions we plan with this knowledge base. Muhammad Amith, Grace Xiong, Kirk Roberts, Cui Tao |
BIBM | 4 |
| 2020 | Extracting Adherence Information from Electronic Health RecordsabstractJordan Sanders, Meghana Gudala, Kathleen Hamilton, Nishtha Prasad, Jordan Stovall, Eduardo Blanco, Jane E Hamilton, Kirk Roberts. Proceedings of the 28th International Conference on Computational Linguistics. 2020. Jordan Sanders, Meghana Gudala, Kathleen E. Hamilton, Nishtha Prasad, Jordan Godfrey-Stovall, Eduardo Blanco 0002, Jane Elizabeth Hamilton, Kirk Roberts |
COLING | 8 |
| 2020 | Rad-SpatialNet: A Frame-based Resource for Fine-Grained Spatial Relations in Radiology ReportsabstractThis paper proposes a representation framework for encoding spatial language in radiology based on frame semantics. The framework is adopted from the existing SpatialNet representation in the general domain with the aim to generate more accurate representations of spatial language used by radiologists. We describe Rad-SpatialNet in detail along with illustrating the importance of incorporating domain knowledge in understanding the varied linguistic expressions involved in different radiological spatial relations. This work also constructs a corpus of 400 radiology reports of three examination types (chest X-rays, brain MRIs, and babygrams) annotated with fine-grained contextual information according to this schema. Spatial trigger expressions and elements corresponding to a spatial frame are annotated. We apply BERT-based models (BERT-Base and BERT- Large) to first extract the trigger terms (lexical units for a spatial frame) and then to identify the related frame elements. The results of BERT- Large are decent, with F1 of 77.89 for spatial trigger extraction and an overall F1 of 81.61 and 66.25 across all frame elements using gold and predicted spatial triggers respectively. This frame-based resource can be used to develop and evaluate more advanced natural language processing (NLP) methods for extracting fine-grained spatial information from radiology text in the future. Surabhi Datta, Morgan Ulinski, Jordan Godfrey-Stovall, Shekhar Khanpara, Roy Riascos-Castaneda, Kirk Roberts |
LREC | 6 |
| 2020 | Evaluation of Dataset Selection for Pre-Training and Fine-Tuning Transformer Language Models for Clinical Question AnsweringabstractWe evaluate the performance of various Transformer language models, when pre-trained and fine-tuned on different combinations of open-domain, biomedical, and clinical corpora on two clinical question answering (QA) datasets (CliCR and emrQA). We perform our evaluations on the task of machine reading comprehension, which involves training the model to answer a question given an unstructured context paragraph. We conduct a total of 48 experiments on different combinations of the large open-domain and domain-specific corpora. We found that an initial fine-tuning on an open-domain dataset, SQuAD, consistently improves the clinical QA performance across all the model variants. Sarvesh Soni, Kirk Roberts |
LREC | 2 |
| 2020 | TREC-COVID: rationale and structure of an information retrieval shared task for COVID-19abstractTREC-COVID is an information retrieval (IR) shared task initiated to support clinicians and clinical research during the COVID-19 pandemic. IR for pandemics breaks many normal assumptions, which can be seen by examining 9 important basic IR research questions related to pandemic situations. TREC-COVID differs from traditional IR shared task evaluations with special considerations for the expected users, IR modality considerations, topic development, participant requirements, assessment process, relevance criteria, evaluation metrics, iteration process, projected timeline, and the implications of data use as a post-task test collection. This article describes how all these were addressed for the particular requirements of developing IR systems under a pandemic situation. Finally, initial participation numbers are also provided, which demonstrate the tremendous interest the IR community has in this effort. Kirk Roberts, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, Kyle Lo, Ian Soboroff, Ellen M. Voorhees, Lucy Lu Wang, William R. Hersh |
J. Am. Medical Informatics Assoc. | 1 |
| 2020 | Deep learning in clinical natural language processing: a methodical reviewabstractOBJECTIVE: This article methodically reviews the literature on deep learning (DL) for natural language processing (NLP) in the clinical domain, providing quantitative analysis to answer 3 research questions concerning methods, scope, and context of current research. MATERIALS AND METHODS: We searched MEDLINE, EMBASE, Scopus, the Association for Computing Machinery Digital Library, and the Association for Computational Linguistics Anthology for articles using DL-based approaches to NLP problems in electronic health records. After screening 1,737 articles, we collected data on 25 variables across 212 papers. RESULTS: DL in clinical NLP publications more than doubled each year, through 2018. Recurrent neural networks (60.8%) and word2vec embeddings (74.1%) were the most popular methods; the information extraction tasks of text classification, named entity recognition, and relation extraction were dominant (89.2%). However, there was a "long tail" of other methods and specific tasks. Most contributions were methodological variants or applications, but 20.8% were new methods of some kind. The earliest adopters were in the NLP community, but the medical informatics community was the most prolific. DISCUSSION: Our analysis shows growing acceptance of deep learning as a baseline for NLP research, and of DL-based NLP in the medical community. A number of common associations were substantiated (eg, the preference of recurrent neural networks for sequence-labeling named entity recognition), while others were surprisingly nuanced (eg, the scarcity of French language clinical NLP with deep learning). CONCLUSION: Deep learning has not yet fully penetrated clinical NLP and is growing rapidly. This review highlighted both the popular and unique trends in this active field. Stephen Wu 0004, Kirk Roberts, Surabhi Datta, Jingcheng Du, Zongcheng Ji, Yuqi Si, Sarvesh Soni, Qiang Wei 0002, Yang Xiang 0003, Bo Zhao 0001, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | Understanding spatial language in radiology: Representation framework, annotation, and spatial relation extraction from chest X-ray reports using deep learning
Surabhi Datta, Yuqi Si, Laritza Rodriguez, Sonya E. Shooshan, Dina Demner-Fushman, Kirk Roberts |
J. Biomed. Informatics | 6 |
| 2020 | A content-based literature recommendation system for datasets to improve data reusability - A case study on Gene Expression Omnibus (GEO) datasets
Braja Gopal Patra, Vahed Maroufy, Babak Soltanalizadeh, W. Jim Zheng, Kirk Roberts, Hulin Wu |
J. Biomed. Informatics | 6 |
| 2019 | A frame semantic overview of NLP-based information extraction for cancer-related EHR notes: a scoping review
Surabhi Datta, Elmer V. Bernstam, Kirk Roberts |
AMIA | 3 |
| 2019 | A Dataset Recommendation System for Researchers based on Publications
Braja Gopal Patra, Kirk Roberts, Hulin Wu |
AMIA | 2 |
| 2019 | Using FHIR to Construct a Corpus of Clinical Questions Annotated with Logical Forms and Answers
Sarvesh Soni, Meghana Gudala, Daisy Zhe Wang, Kirk Roberts |
AMIA | 4 |
| 2019 | Relation Extraction from Clinical Narratives Using Pre-trained Language Models
Qiang Wei 0002, Zongcheng Ji, Yuqi Si, Jingcheng Du, Firat Tiryaki, Stephen Wu 0004, Cui Tao, Kirk Roberts, Hua Xu 0001 |
AMIA | 9 |
| 2019 | Ontology of Consumer Health Vocabulary: providing a formal and interoperable semantic resource for linking lay language and medical terminologyabstractThe Consumer Health Vocabulary has been an important contribution to the health informatics field since its introduction in 2006. Many studies have utilized the vocabulary for various scientific research to bridge the gap between consumers and health experts. Given the flat file format of the Consumer Health Vocabulary dataset, we developed a SKOS-based ontology of the dataset. As an ontology, this dataset can be semantically linked to other resources to provide consumer-level meaning. In addition with this artifact, we plan to further expand the terminology. Muhammad Amith, Licong Cui, Kirk Roberts, Hua Xu 0001, Cui Tao |
BIBM | 3 |
| 2019 | Conceiving an application ontology to model patient human papillomavirus vaccine counseling for dialogue managementabstractBACKGROUND: In the United States and parts of the world, the human papillomavirus vaccine uptake is below the prescribed coverage rate for the population. Some research have noted that dialogue that communicates the risks and benefits, as well as patient concerns, can improve the uptake levels. In this paper, we introduce an application ontology for health information dialogue called Patient Health Information Dialogue Ontology for patient-level human papillomavirus vaccine counseling and potentially for any health-related counseling. RESULTS: The ontology's class level hierarchy is segmented into 4 basic levels - Discussion, Goal, Utterance, and Speech Task. The ontology also defines core low-level utterance interaction for communicating human papillomavirus health information. We discuss the design of the ontology and the execution of the utterance interaction. CONCLUSION: With an ontology that represents patient-centric dialogue to communicate health information, we have an application-driven model that formalizes the structure for the communication of health information, and a reusable scaffold that can be integrated for software agents. Our next step will to be develop the software engine that will utilize the ontology and automate the dialogue interaction of a software agent. Muhammad Amith, Kirk Roberts, Cui Tao |
BMC Bioinform. | 2 |
| 2019 | Enhancing clinical concept extraction with contextual embeddingsabstractOBJECTIVE: Neural network-based representations ("embeddings") have dramatically advanced natural language processing (NLP) tasks, including clinical NLP tasks such as concept extraction. Recently, however, more advanced embedding methods and representations (eg, ELMo, BERT) have further pushed the state of the art in NLP, yet there are no common best practices for how to integrate these representations into clinical tasks. The purpose of this study, then, is to explore the space of possible options in utilizing these new models for clinical concept extraction, including comparing these to traditional word embedding methods (word2vec, GloVe, fastText). MATERIALS AND METHODS: Both off-the-shelf, open-domain embeddings and pretrained clinical embeddings from MIMIC-III (Medical Information Mart for Intensive Care III) are evaluated. We explore a battery of embedding methods consisting of traditional word embeddings and contextual embeddings and compare these on 4 concept extraction corpora: i2b2 2010, i2b2 2012, SemEval 2014, and SemEval 2015. We also analyze the impact of the pretraining time of a large language model like ELMo or BERT on the extraction performance. Last, we present an intuitive way to understand the semantic information encoded by contextual embeddings. RESULTS: Contextual embeddings pretrained on a large clinical corpus achieves new state-of-the-art performances across all concept extraction tasks. The best-performing model outperforms all state-of-the-art methods with respective F1-measures of 90.25, 93.18 (partial), 80.74, and 81.65. CONCLUSIONS: We demonstrate the potential of contextual embeddings through the state-of-the-art performance these methods achieve on clinical concept extraction. Additionally, we demonstrate that contextual embeddings encode valuable semantic information not accounted for in traditional word representations. Yuqi Si, Hua Xu 0001, Kirk Roberts |
J. Am. Medical Informatics Assoc. | 4 |
| 2019 | A frame semantic overview of NLP-based information extraction for cancer-related EHR notes
Surabhi Datta, Elmer V. Bernstam, Kirk Roberts |
J. Biomed. Informatics | 3 |
| 2019 | Distributed learning from multiple EHR databases: Contextual embedding models for medical events
Ziyi Li 0001, Kirk Roberts, Xiaoqian Jiang, Qi Long |
J. Biomed. Informatics | 2 |
| 2018 | Adverse Reactions and Drug-Drug Interaction Extraction tracks at the Text Analysis Conference (TAC)
Dina Demner-Fushman, Joseph M. Tonning, Kin Wah Fung, Phong Do, Richard D. Boyce, Kirk Roberts |
AMIA | 6 |
| 2018 | Benchmarking Information Retrieval for Precision Oncology: the TREC Precision Medicine Track
Kirk Roberts, Dina Demner-Fushman, Ellen M. Voorhees, William R. Hersh, Steven Bedrick, Alexander J. Lazar, Shubham Pant |
AMIA | 1 |
| 2018 | A Frame-Based NLP System for Cancer-Related Information Extraction
Yuqi Si, Kirk Roberts |
AMIA | 2 |
| 2018 | A FrameNet for Cancer Information in Clinical Narratives: Schema and Annotation
Kirk Roberts, Yuqi Si, Anshul Gandhi, Elmer V. Bernstam |
LREC | 1 |
| 2017 | Leveraging existing corpora for de-identification of psychiatric notes using domain adaptation
Hee-Jin Lee, Yaoyun Zhang, Kirk Roberts, Hua Xu 0001 |
AMIA | 3 |
| 2017 | Information Retrieval for Biomedical Datasets: The 2016 bioCADDIE Challenge
Kirk Roberts, Anupama E. Gururaj, Saeid Pournejati, Trevor Cohen, William R. Hersh, Dina Demner-Fushman, Lucila Ohno-Machado, Hua Xu 0001 |
AMIA | 1 |
| 2017 | A Semantic Parsing Method for Mapping Clinical Questions to Logical Forms
Kirk Roberts, Braja Gopal Patra |
AMIA | 1 |
| 2017 | Interweaving Domain Knowledge and Unsupervised Learning for Psychiatric Stressor Extraction from Clinical Notes
Olivia R. Zhang, Yaoyun Zhang, Jun Xu 0007, Kirk Roberts, Xiang Y. Zhang, Hua Xu 0001 |
IEA/AIE (2) | 4 |
| 2017 | Biomedical informatics advancing the national health agenda: the AMIA 2015 year-in-review in clinical and consumer informaticsabstractThe field of biomedical informatics experienced a productive 2015 in terms of research. In order to highlight the accomplishments of that research, elicit trends, and identify shortcomings at a macro level, a 19-person team conducted an extensive review of the literature in clinical and consumer informatics. The result of this process included a year-in-review presentation at the American Medical Informatics Association Annual Symposium and a written report (see supplemental data). Key findings are detailed in the report and summarized here. This article organizes the clinical and consumer health informatics research from 2015 under 3 themes: the electronic health record (EHR), the learning health system (LHS), and consumer engagement. Key findings include the following: (1) There are significant advances in establishing policies for EHR feature implementation, but increased interoperability is necessary for these to gain traction. (2) Decision support systems improve practice behaviors, but evidence of their impact on clinical outcomes is still lacking. (3) Progress in natural language processing (NLP) suggests that we are approaching but have not yet achieved truly interactive NLP systems. (4) Prediction models are becoming more robust but remain hampered by the lack of interoperable clinical data records. (5) Consumers can and will use mobile applications for improved engagement, yet EHR integration remains elusive. Kirk Roberts, Mary Regina Boland, Lisiane Pruinelli, Jina J. Dcruz, Andrew B. L. Berry, Mattias Georgsson, Rebecca Hazen, Raymond Francis Sarmiento, Uba Backonja, Kun-Hsing Yu, Patricia Flatley Brennan |
J. Am. Medical Informatics Assoc. | 1 |
| 2017 | A protocol-driven approach to automatically finding authoritative answers to consumer health questions in online resourcesabstractThe purpose of this research was to establish an upper bound on finding answers to health‐related questions in MedlinePlus and other online resources. Seven reference librarians tested a set of protocols to determine whether it was possible to use the types and foci of the questions extracted from customer requests submitted to the National Library of Medicine to find authoritative answers to these questions. Librarians tested the protocols manually to determine if the process was sufficiently robust and accurate to later automate. Results indicated that the extracted terms provide enough information to find authoritative answers for about 60% of questions and that certain question types are more likely to result in authoritative answers than others. The question corpus and analysis performed for this project will inform automatic question answering systems, and could lead to suggestions for new content to include in MedlinePlus. This approach can serve as an example to researchers interested in methods of evaluating question answering tools and the contents of online databases. Ariel Deardorff, Kate Masterton, Kirk Roberts, Halil Kilicoglu, Dina Demner-Fushman |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2016 | Combining Open-domain and Biomedical Knowledge for Topic Recognition in Consumer Health Questions
Yassine Mrabet, Halil Kilicoglu, Kirk Roberts, Dina Demner-Fushman |
AMIA | 3 |
| 2016 | Resource Classification for Medical Questions
Kirk Roberts, Laritza Rodriguez, Sonya E. Shooshan, Dina Demner-Fushman |
AMIA | 1 |
| 2016 | Learning to Read Chest X-Rays: Recurrent Neural Cascade Model for Automated Image AnnotationabstractDespite the recent advances in automatically describing image contents, their applications have been mostly limited to image caption datasets containing natural images (e.g., Flickr 30k, MSCOCO). In this paper, we present a deep learning model to efficiently detect a disease from an image and annotate its contexts (e.g., location, severity and the affected organs). We employ a publicly available radiology dataset of chest x-rays and their reports, and use its image annotations to mine disease names to train convolutional neural networks (CNNs). In doing so, we adopt various regularization techniques to circumvent the large normalvs-diseased cases bias. Recurrent neural networks (RNNs) are then trained to describe the contexts of a detected disease, based on the deep CNN features. Moreover, we introduce a novel approach to use the weights of the already trained pair of CNN/RNN on the domain-specific image/text dataset, to infer the joint image/text contexts for composite image labeling. Significantly improved image annotation results are demonstrated using the recurrent neural cascade model by taking the joint image/text contexts into account. Hoo-Chang Shin, Kirk Roberts, Le Lu 0001, Dina Demner-Fushman, Jianhua Yao 0001, Ronald M. Summers |
CVPR | 2 |
| 2016 | Annotating Named Entities in Consumer Health Questions
Halil Kilicoglu, Asma Ben Abacha, Yassine Mrabet, Kirk Roberts, Laritza Rodriguez, Sonya E. Shooshan, Dina Demner-Fushman |
LREC | 4 |
| 2016 | Annotating Logical Forms for EHR Questions
Kirk Roberts, Dina Demner-Fushman |
LREC | 1 |
| 2016 | State-of-the-art in biomedical literature retrieval for clinical cases: a survey of the TREC 2014 CDS track
Kirk Roberts, Matthew S. Simpson, Dina Demner-Fushman, Ellen M. Voorhees, William R. Hersh |
Inf. Retr. J. | 1 |
| 2016 | Interactive use of online health resources: a comparison of consumer and professional questionsabstractOBJECTIVE: To understand how consumer questions on online resources differ from questions asked by professionals, and how such consumer questions differ across resources. MATERIALS AND METHODS: Ten online question corpora, 5 consumer and 5 professional, with a combined total of over 40 000 questions, were analyzed using a variety of natural language processing techniques. These techniques analyze questions at the lexical, syntactic, and semantic levels, exposing differences in both form and content. RESULTS: Consumer questions tend to be longer than professional questions, more closely resemble open-domain language, and focus far more on medical problems. Consumers ask more sub-questions, provide far more background information, and ask different types of questions than professionals. Furthermore, there is substantial variance of these factors between the different consumer corpora. DISCUSSION: The form of consumer questions is highly dependent upon the individual online resource, especially in the amount of background information provided. Professionals, on the other hand, provide very little background information and often ask much shorter questions. The content of consumer questions is also highly dependent upon the resource. While professional questions commonly discuss treatments and tests, consumer questions focus disproportionately on symptoms and diseases. Further, consumers place far more emphasis on certain types of health problems (eg, sexual health). CONCLUSION: Websites for consumers to submit health questions are a popular online resource filling important gaps in consumer health information. By analyzing how consumers write questions on these resources, we can better understand these gaps and create solutions for improving information access.This article is part of the Special Focus on Person-Generated Health and Wellness Data, which published in the May 2016 issue, Volume 23, Issue 3. Kirk Roberts, Dina Demner-Fushman |
J. Am. Medical Informatics Assoc. | 1 |
| 2015 | An Ensemble Method for Spelling Correction in Consumer Health Questions
Halil Kilicoglu, Marcelo Fiszman, Kirk Roberts, Dina Demner-Fushman |
AMIA | 3 |
| 2015 | Automatic Extraction and Post-coordination of Spatial Relations in Consumer Language
Kirk Roberts, Laritza Rodriguez, Sonya E. Shooshan, Dina Demner-Fushman |
AMIA | 1 |
| 2015 | The role of fine-grained annotations in supervised recognition of risk factors for heart disease from EHRsabstractThis paper describes a supervised machine learning approach for identifying heart disease risk factors in clinical text, and assessing the impact of annotation granularity and quality on the system's ability to recognize these risk factors. We utilize a series of support vector machine models in conjunction with manually built lexicons to classify triggers specific to each risk factor. The features used for classification were quite simple, utilizing only lexical information and ignoring higher-level linguistic information such as syntax and semantics. Instead, we incorporated high-quality data to train the models by annotating additional information on top of a standard corpus. Despite the relative simplicity of the system, it achieves the highest scores (micro- and macro-F1, and micro- and macro-recall) out of the 20 participants in the 2014 i2b2/UTHealth Shared Task. This system obtains a micro- (macro-) precision of 0.8951 (0.8965), recall of 0.9625 (0.9611), and F1-measure of 0.9276 (0.9277). Additionally, we perform a series of experiments to assess the value of the annotated data we created. These experiments show how manually-labeled negative annotations can improve information extraction performance, demonstrating the importance of high-quality, fine-grained natural language annotations. Kirk Roberts, Sonya E. Shooshan, Laritza Rodriguez, Swapna Abhyankar, Halil Kilicoglu, Dina Demner-Fushman |
J. Biomed. Informatics | 1 |
| 2014 | Error Propagation in EHRs via Copy/Paste: An Analysis of Relative Dates
Kirk Roberts, Amos Cahan, Dina Demner-Fushman |
AMIA | 1 |
| 2014 | Automatically Classifying Question Types for Consumer Health Questions
Kirk Roberts, Halil Kilicoglu, Marcelo Fiszman, Dina Demner-Fushman |
AMIA | 1 |
| 2014 | Annotating Question Decomposition on Complex Medical Questions
Kirk Roberts, Kate Masterton, Marcelo Fiszman, Halil Kilicoglu, Dina Demner-Fushman |
LREC | 1 |
| 2013 | A flexible framework for recognizing events, temporal expressions, and temporal relations in clinical textabstractOBJECTIVE: To provide a natural language processing method for the automatic recognition of events, temporal expressions, and temporal relations in clinical records. MATERIALS AND METHODS: A combination of supervised, unsupervised, and rule-based methods were used. Supervised methods include conditional random fields and support vector machines. A flexible automated feature selection technique was used to select the best subset of features for each supervised task. Unsupervised methods include Brown clustering on several corpora, which result in our method being considered semisupervised. RESULTS: On the 2012 Informatics for Integrating Biology and the Bedside (i2b2) shared task data, we achieved an overall event F1-measure of 0.8045, an overall temporal expression F1-measure of 0.6154, an overall temporal link detection F1-measure of 0.5594, and an end-to-end temporal link detection F1-measure of 0.5258. The most competitive system was our event recognition method, which ranked third out of the 14 participants in the event task. DISCUSSION: Analysis reveals the event recognition method has difficulty determining which modifiers to include/exclude in the event span. The temporal expression recognition method requires significantly more normalization rules, although many of these rules apply only to a small number of cases. Finally, the temporal relation recognition method requires more advanced medical knowledge and could be improved by separating the single discourse relation classifier into multiple, more targeted component classifiers. CONCLUSIONS: Recognizing events and temporal expressions can be achieved accurately by combining supervised and unsupervised methods, even when only minimal medical knowledge is available. Temporal normalization and temporal relation recognition, however, are far more dependent on the modeling of medical knowledge. Kirk Roberts, Bryan Rink, Sanda M. Harabagiu |
J. Am. Medical Informatics Assoc. | 1 |
| 2012 | A Machine Learning Approach for Identifying Anatomical Locations of Actionable Findings in Radiology Reports
Kirk Roberts, Bryan Rink, Sanda M. Harabagiu, Richard H. Scheuermann, Seth M. Toomay, Travis Browning, Teresa Bosler, Ronald M. Peshock |
AMIA | 1 |
| 2012 | Locational relativity and domain constraints in spatial questionsabstractSpatial queries in the form of natural language questions have typically been assumed to have unconstrained geographic answers. However, analysis of prototypical spatial questions reveals two important types of constraints that must be considered by spatial question answering systems. First, locational relativity constraints limit answers to a particular location or the user's implied location. Second, domain constraints specify non-geographic locations such as web pages or anatomical sites. In order to detect these constraints, we have conducted a crowd-sourced annotation effort for a set of over 1,200 questions gathered from a community question answering website. We utilize machine learning techniques trained on this data to automatically classify these two types of constraints. We report results nearing 90% accuracy at locational relativity detection and 76% accuracy at domain classification using this approach. Kirk Roberts, Sanda M. Harabagiu |
SIGSPATIAL/GIS | 1 |
| 2012 | Annotating Spatial Containment Relations Between Events
Kirk Roberts, Travis R. Goodwin, Sanda M. Harabagiu |
LREC | 1 |
| 2012 | EmpaTweet: Annotating and Detecting Emotions on Twitter
Kirk Roberts, Michael A. Roach, Josh Guthrie, Sanda M. Harabagiu |
LREC | 1 |
| 2012 | A supervised framework for resolving coreference in clinical recordsabstractOBJECTIVE: A method for the automatic resolution of coreference between medical concepts in clinical records. MATERIALS AND METHODS: A multiple pass sieve approach utilizing support vector machines (SVMs) at each pass was used to resolve coreference. Information such as lexical similarity, recency of a concept mention, synonymy based on Wikipedia redirects, and local lexical context were used to inform the method. Results were evaluated using an unweighted average of MUC, CEAF, and B(3) coreference evaluation metrics. The datasets used in these research experiments were made available through the 2011 i2b2/VA Shared Task on Coreference. RESULTS: The method achieved an average F score of 0.821 on the ODIE dataset, with a precision of 0.802 and a recall of 0.845. These results compare favorably to the best-performing system with a reported F score of 0.827 on the dataset and the median system F score of 0.800 among the eight teams that participated in the 2011 i2b2/VA Shared Task on Coreference. On the i2b2 dataset, the method achieved an average F score of 0.906, with a precision of 0.895 and a recall of 0.918 compared to the best F score of 0.915 and the median of 0.859 among the 16 participating teams. DISCUSSION: Post hoc analysis revealed significant performance degradation on pathology reports. The pathology reports were characterized by complex synonymy and very few patient mentions. CONCLUSION: The use of several simple lexical matching methods had the most impact on achieving competitive performance on the task of coreference resolution. Moreover, the ability to detect patients in electronic medical records helped to improve coreference resolution more than other linguistic analysis. Bryan Rink, Kirk Roberts, Sanda M. Harabagiu |
J. Am. Medical Informatics Assoc. | 2 |
| 2011 | Unsupervised Learning of Selectional Restrictions and Detection of Argument Coercions
Kirk Roberts, Sanda M. Harabagiu |
EMNLP | 1 |
| 2011 | Automatic extraction of relations between medical concepts in clinical textsabstractOBJECTIVE: A supervised machine learning approach to discover relations between medical problems, treatments, and tests mentioned in electronic medical records. MATERIALS AND METHODS: A single support vector machine classifier was used to identify relations between concepts and to assign their semantic type. Several resources such as Wikipedia, WordNet, General Inquirer, and a relation similarity metric inform the classifier. RESULTS: The techniques reported in this paper were evaluated in the 2010 i2b2 Challenge and obtained the highest F1 score for the relation extraction task. When gold standard data for concepts and assertions were available, F1 was 73.7, precision was 72.0, and recall was 75.3. F1 is defined as 2*Precision*Recall/(Precision+Recall). Alternatively, when concepts and assertions were discovered automatically, F1 was 48.4, precision was 57.6, and recall was 41.7. DISCUSSION: Although a rich set of features was developed for the classifiers presented in this paper, little knowledge mining was performed from medical ontologies such as those found in UMLS. Future studies should incorporate features extracted from such knowledge sources, which we expect to further improve the results. Moreover, each relation discovery was treated independently. Joint classification of relations may further improve the quality of results. Also, joint learning of the discovery of concepts, assertions, and relations may also improve the results of automatic relation extraction. CONCLUSION: Lexical and contextual features proved to be very important in relation extraction from medical texts. When they are not available to the classifier, the F1 score decreases by 3.7%. In addition, features based on similarity contribute to a decrease of 1.1% when they are not available. Bryan Rink, Sanda M. Harabagiu, Kirk Roberts |
J. Am. Medical Informatics Assoc. | 3 |
| 2011 | A flexible framework for deriving assertions from electronic medical recordsabstractOBJECTIVE: This paper describes natural-language-processing techniques for two tasks: identification of medical concepts in clinical text, and classification of assertions, which indicate the existence, absence, or uncertainty of a medical problem. Because so many resources are available for processing clinical texts, there is interest in developing a framework in which features derived from these resources can be optimally selected for the two tasks of interest. MATERIALS AND METHODS: The authors used two machine-learning (ML) classifiers: support vector machines (SVMs) and conditional random fields (CRFs). Because SVMs and CRFs can operate on a large set of features extracted from both clinical texts and external resources, the authors address the following research question: Which features need to be selected for obtaining optimal results? To this end, the authors devise feature-selection techniques which greatly reduce the amount of manual experimentation and improve performance. RESULTS: The authors evaluated their approaches on the 2010 i2b2/VA challenge data. Concept extraction achieves 79.59 micro F-measure. Assertion classification achieves 93.94 micro F-measure. DISCUSSION: Approaching medical concept extraction and assertion classification through ML-based techniques has the advantage of easily adapting to new data sets and new medical informatics tasks. However, ML-based techniques perform best when optimal features are selected. By devising promising feature-selection techniques, the authors obtain results that outperform the current state of the art. CONCLUSION: This paper presents two ML-based approaches for processing language in the clinical texts evaluated in the 2010 i2b2/VA challenge. By using novel feature-selection methods, the techniques presented in this paper are unique among the i2b2 participants. Kirk Roberts, Sanda M. Harabagiu |
J. Am. Medical Informatics Assoc. | 1 |
| 2010 | C-3: Coherence and Coreference Corpus
Cristina Nicolae, Gabriel Nicolae, Kirk Roberts |
LREC | 3 |
| 2010 | A Linguistic Resource for Semantic Parsing of Motion Events
Kirk Roberts, Srikanth Gullapalli, Cosmin Adrian Bejan, Sanda M. Harabagiu |
LREC | 1 |
| 2008 | Scaling Answer Type Detection to Large Hierarchies
Kirk Roberts, Andrew Hickl |
LREC | 1 |