EDBT 2026 Demo / reviewers in the wild / expert
Hong Yu 0001
dblp:55/6749-1
· DBLP profile ↗
98ranked-venue papers
12as first author
29since 2021 · last 2026
0000-0001-9263-5035ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 55 · 10 first-author · 5 since 2021Artificial intelligence and machine learning · 42 · 2 first-author · 24 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PRIME: Planning and Retrieval-Integrated Memory for Enhanced ReasoningabstractInspired by the dual-process theory of human cognition from Thinking, Fast and Slow, we introduce PRIME (Planning and Retrieval-Integrated Memory for Enhanced Reasoning), a multi-agent reasoning framework that dynamically integrates System 1 (fast, intuitive thinking) and System 2 (slow, deliberate thinking). PRIME first employs a Quick Thinking Agent to generate a rapid answer; if uncertainty is detected, it then triggers a structured System 2 reasoning pipeline composed of specialized agents for planning, hypothesis generation, retrieval, information integration, and decision-making. This multi-agent design mimics human cognitive processes faithfully and enhances both efficiency and accuracy. Experimental results with LLaMA 3 models demonstrate that PRIME enables open-source LLMs to perform competitively with state-of-the-art closed-source models like GPT-4 and GPT-4o on benchmarks requiring multi-hop and knowledge-grounded reasoning. This research establishes PRIME as a scalable solution for improving LLMs in domains requiring complex, knowledge-intensive reasoning. Hieu Tran, Zonghai Yao, Nguyen Luong Tran, Zhichao Yang 0001, Feiyun Ouyang, Razieh Rahimi, Hong Yu 0001 |
AAAI | 8 |
| 2026 | ChatCLIDS: Simulating Persuasive AI Dialogues to Promote Closed-Loop Insulin Adoption in Type 1 Diabetes CareabstractReal-world adoption of closed-loop insulin delivery systems (CLIDS) in type 1 diabetes remains low, driven not by technical failure, but by diverse behavioral, psychosocial, and social barriers. We introduce ChatCLIDS, the first benchmark to rigorously evaluate LLM–driven persuasive dialogue for health behavior change. Our framework features a library of expert-validated virtual patients, each with clinically grounded, heterogeneous profiles and realistic adoption barriers, and simulates multi-turn interactions with nurse agents equipped with a diverse set of evidence-based persuasive strategies. ChatCLIDS uniquely supports longitudinal counseling and adversarial social influence scenarios, enabling robust, multi-dimensional evaluation. Our findings reveal that while larger and more reflective LLMs adapt strategies over time, all models struggle to overcome resistance, especially under realistic social pressure. These results highlight critical limitations of current LLMs for behavior change, and offer a high-fidelity, scalable testbed for advancing trustworthy persuasive AI in healthcare and beyond. Zonghai Yao, Talha Chafekar, Junda Wang, Feiyun Ouyang, Junhui Qian, Hong Yu 0001 |
AAAI | 8 |
| 2026 | LLM-Based Multi-Agent Systems for Clinical Workflows: A Survey of AI HospitalsabstractThis survey reviews LLM-based multi-agent systems for clinical and healthcare workflows, including diagnosis, triage, consultation, discharge, mental health, and EHR-linked decision support.We define AI hospitals as workflow-level clinical systems in which agents take explicit roles, hand off shared state, use EHR-or guideline-grounded tools, and operate with safety gates and audit-ready logs.We argue that these systems should be compared at the workflow level, rather than only by model components or end-task accuracy, because clinical action, evidence, and accountability are expressed through state transitions and handoffs.We organize the literature through a workflowlevel taxonomy covering roles and handoffs, memory and evidence, tools, and reasoning, control, and escalation.We further synthesize major workflow settings and task families, introduce a four-layer evaluation stack spanning safety, process, outcome, and operations, and connect model capabilities to workflow observables relevant to deployment.Finally, we present Integration Readiness Levels (IRL1-IRL6), task-level instrumentation requirements, and recurring workflow failure modes as a practical framework for comparing, evaluating, and deploying clinical LLM agents and AI hospitals.1 Zonghai Yao, Hong Yu 0001 |
ACL (1) | 2 |
| 2025 | RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Modelsabstract, which leverages information retrieval specifically for generated sub-questions and re-answers these sub-questions with the relevant contextual information. Additionally, a Retrieval-Augmented Factuality Scorer is proposed to replace the original discriminator, prioritizing reasoning paths that meet high standards of factuality. Experimental results with LLaMA 3.1 show that RARE enables open-source LLMs to achieve competitive performance with top closed-source models like GPT-4 and GPT-4o. This research establishes RARE as a scalable solution for improving LLMs in domains where logical coherence and factual integrity are critical. Hieu Tran, Zonghai Yao, Zhichao Yang 0001, Junda Wang, Feiyun Ouyang, Hong Yu 0001 |
ACL (1) | 8 |
| 2025 | Improving Rare and Common ICD Coding via a Multi-Agent LLM-Based ApproachabstractLarge Language Models (LLMs) have shown strong performance in tasks such as zero- and few-shot information extraction from clinical text without domain-specific training. However, in the ICD coding task, LLMs often hallucinate key details and produce high-recall but low-precision outputs due to the high-dimensional and imbalanced nature of ICD code distributions. Existing LLM-based approaches typically fail to capture the complex, dynamic interactions among human agents involved in real-world coding workflows-such as patients, physicians, and coders-and often lack interpretability and reliability. To address these challenges, we propose a novel multi-agent framework for ICD coding that simulates the real-world process using five role-specific LLM agents-patient, physician, coder, reviewer, and adjuster-and integrates the Subjective, Objective, Assessment, and Plan (SOAP) structure from Electronic Health Records to enhance performance. Evaluated on the MIMIC-III dataset, our method significantly outperforms zero-shot Chain-of-Thought prompting, self-consistency strategies, and LLM-designed agent baselines, particularly for rare codes. Ablation studies confirm the contribution of each agent role, and the system achieves competitive performance with state-of-the-art fine-tuned models, while offering better explainability and requiring no task-specific pre-training. Rumeng Li, Hong Yu 0001 |
CIKM | 3 |
| 2025 | Synth-SBDH: A Synthetic Dataset of Social and Behavioral Determinants of Health for Clinical TextabstractSocial and behavioral determinants of health (SBDH) play a crucial role in health outcomes and are frequently documented in clinical text. Automatically extracting SBDH information from clinical text relies on publicly available good-quality datasets. However, existing SBDH datasets exhibit substantial limitations in their availability and coverage. In this study, we introduce Synth-SBDH, a novel synthetic dataset with detailed SBDH annotations, encompassing status, temporal information, and rationale across 15 SBDH categories. We showcase the utility of Synth-SBDH on three tasks using real-world clinical datasets from two distinct hospital settings, highlighting its versatility, generalizability, and distillation capabilities. Models trained on Synth-SBDH consistently outperform counterparts with no Synth-SBDH training, achieving up to 63.75% macro-F improvements. Additionally, Synth-SBDH proves effective for rare SBDH categories and under-resource constraints while being substantially cheaper than expert-annotated real-world data. Human evaluation reveals a 71.06% Human-LLM alignment and uncovers areas for future refinements. Avijit Mitra, Zhichao Yang 0001, Emily Druhl, Raelene Goodwin, Hong Yu 0001 |
EMNLP | 5 |
| 2025 | From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical CalculationsabstractLarge language models (LLMs) have demonstrated promising performance on medical benchmarks; however, their ability to perform medical calculations, a crucial aspect of clinical decision-making, remains underexplored and poorly evaluated. Existing benchmarks often assess only the final answer with a wide numerical tolerance, overlooking systematic reasoning failures and potentially causing serious clinical misjudgments. In this work, we revisit medical calculation evaluation with a stronger focus on clinical trustworthiness. First, we clean and restructure the MedCalc-Bench dataset and propose a new step-by-step evaluation pipeline that independently assesses formula selection, entity extraction, and arithmetic computation. Under this granular framework, the accuracy of GPT-4o drops from 62.7% to 43.6%, revealing errors masked by prior evaluations. Second, we introduce an automatic error analysis framework that generates structured attribution for each failure mode. Human evaluation confirms its alignment with expert judgment, enabling scalable and explainable diagnostics. Finally, we propose a modular agentic pipeline, MedRaC, that combines retrieval-augmented generation and Python-based code execution. Without any fine-tuning, MedRaC improves the accuracy of different LLMs from 16.35% up to 53.19%. Our work highlights the limitations of current benchmark practices and proposes a more clinically faithful methodology. By enabling transparent and transferable reasoning evaluation, we move closer to making LLM-based systems trustworthy for real-world medical applications. Benlu Wang, Iris Xia, Junda Wang, Feiyun Ouyang, Arman Cohan, Hong Yu 0001, Zonghai Yao |
EMNLP | 8 |
| 2025 | DischargeSim: A Simulation Benchmark for Educational Doctor-Patient Communication at DischargeabstractDischarge communication is a critical yet underexplored component of patient care, where the goal shifts from diagnosis to education.While recent large language model (LLM) benchmarks emphasize in-visit diagnostic reasoning, they fail to evaluate models' ability to support patients after the visit.We introduce DischargeSim, a novel benchmark that evaluates LLMs on their ability to act as personalized discharge educators.DischargeSim simulates post-visit, multi-turn conversations between LLM-driven DoctorAgents and Pa-tientAgents with diverse psychosocial profiles (e.g., health literacy, education, emotion).Interactions are structured across six clinically grounded discharge topics and assessed along three axes: (1) dialogue quality via automatic and LLM-as-judge evaluation, (2) personalized document generation including free-text summaries and structured AHRQ checklists, and ( 3) patient comprehension through a downstream multiple-choice exam.Experiments across 18 LLMs reveal significant gaps in discharge education capability, with performance varying widely across patient profiles.Notably, model size does not always yield better education outcomes, highlighting trade-offs in strategy use and content prioritization.DischargeSim offers a first step toward benchmarking LLMs in post-visit clinical education and promoting equitable, personalized patient support. 1 .* indicates equal contribution 1 The source code is released at: https://github.com/ michaels6060/DischargeSim with CC-BY-NC 4.0 license.Discharge notes …… -Discharge Diagnoses -Discharge Medications -Discharge Condition -Discharge Instructions -Follow-up Plan : You'll be taking furosemide 40 mg by mouth every morning-it helps reduce fluid overload and also lowers your blood pressure.But because it can make you urinate more and lower potassium, you're also on potassium chloride 10 mEq daily.I took the 40 mg furosemide today and felt lightheaded.Should I be worried?: It's not uncommon.Make sure to take the 40 mg dose after breakfast, not on an empty stomach.Also, rise slowly from sitting or lying down.If the dizziness continues or worsens, call us-we may lower the dose or adjust your schedule.… Zonghai Yao, Michael Sun, Won Seok Jang, Sunjae Kwon, Soie Kwon, Hong Yu 0001 |
EMNLP | 6 |
| 2025 | MCQG-SRefine: Multiple Choice Question Generation and Evaluation with Iterative Self-Critique, Correction, and Comparison FeedbackabstractAutomatic question generation (QG) is essential for AI and NLP, particularly in intelligent tutoring, dialogue systems, and fact verification. Generating multiple-choice questions (MCQG) for professional exams, like the United States Medical Licensing Examination (USMLE), is particularly challenging, requiring domain expertise and complex multi-hop reasoning for high-quality questions. However, current large language models (LLMs) like GPT-4 struggle with professional MCQG due to outdated knowledge, hallucination issues, and prompt sensitivity, resulting in unsatisfactory quality and difficulty. To address these challenges, we propose MCQG-SRefine, an LLM self-refine-based (Critique and Correction) framework for converting medical cases into high-quality USMLE-style questions. By integrating expert-driven prompt engineering with iterative self-critique and self-correction feedback, MCQG-SRefine significantly enhances human expert satisfaction regarding both the quality and difficulty of the questions. Furthermore, we introduce an LLM-as-Judge-based automatic metric to replace the complex and costly expert evaluation process, ensuring reliable and expert-aligned assessments. Zonghai Yao, Aditya Parashar, Huixue Zhou, Won Seok Jang, Feiyun Ouyang, Zhichao Yang 0001, Hong Yu 0001 |
NAACL (Long Papers) | 7 |
| 2025 | RADAR: Benchmarking Language Models on Imperfect Tabular DataabstractLanguage models (LMs) are increasingly being deployed to perform autonomous data analyses. However, their data awareness—the ability to recognize, reason over, and appropriately handle data artifacts such as missing values, outliers, and logical inconsistencies—remains underexplored. These artifacts are especially common in real-world tabular data and, if mishandled, can significantly compromise the validity of analytical conclusions. To address this gap, we present RADAR, a benchmark for systematically evaluating data-aware reasoning on tabular data. We develop a framework to simulate data artifacts via programmatic perturbations to enable targeted evaluation of model behavior. RADAR comprises 2,980 table-query pairs, grounded in real-world data spanning 9 domains and 5 data artifact types. In addition to evaluating artifact handling, RADAR systematically varies table size to study how reasoning performance holds when increasing table size. Our evaluation reveals that, despite decent performance on tables without data artifacts, frontier models degrade significantly when data artifacts are introduced, exposing critical gaps in their capacity for robust, data-aware analysis. Designed to be flexible and extensible, RADAR supports diverse perturbation types and controllable table sizes, offering a valuable resource for advancing tabular reasoning. Ken Gu, Zhihan Zhang 0002, Kate Lin, Yuwei Zhang 0001, Akshay Paruchuri, Hong Yu 0001, Mehran Kazemi, Kumar Ayush, A. Ali Heydari, Maxwell A. Xu, Yun Liu 0013, Ming-Zher Poh, Yuzhe Yang 0003, Mark Malhotra, Shwetak N. Patel, Hamid Palangi, Xuhai Xu, Daniel McDuff, Tim Althoff, Xin Liu 0034 |
NeurIPS | 6 |
| 2024 | LocalTweets to LocalHealth: A Mental Health Surveillance Framework Based on Twitter DataabstractPrior research on Twitter (now X) data has provided positive evidence of its utility in developing supplementary health surveillance systems. In this study, we present a new framework to surveil public health, focusing on mental health (MH) outcomes. We hypothesize that locally posted tweets are indicative of local MH outcomes and collect tweets posted from 765 neighborhoods (census block groups) in the USA. We pair these tweets from each neighborhood with the corresponding MH outcome reported by the Center for Disease Control (CDC) to create a benchmark dataset, LocalTweets. With LocalTweets, we present the first population-level evaluation task for Twitter-based MH surveillance systems. We then develop an efficient and effective method, LocalHealth, for predicting MH outcomes based on LocalTweets. When used with GPT3.5, LocalHealth achieves the highest F1-score and accuracy of 0.7429 and 79.78%, respectively, a 59% improvement in F1-score over the GPT3.5 in zero-shot setting. We also utilize LocalHealth to extrapolate CDC’s estimates to proxy unreported neighborhoods, achieving an F1-score of 0.7291. Our work suggests that Twitter data can be effectively leveraged to simulate neighborhood-level MH outcomes. Vijeta Deshpande, Minhwa Lee, Zonghai Yao, Zihao Zhang 0001, Jason Brian Gibbons, Hong Yu 0001 |
LREC/COLING | 6 |
| 2024 | LlamaCare: An Instruction Fine-Tuned Large Language Model for Clinical NLPabstractLarge language models (LLMs) have shown remarkable abilities in generating natural texts for various tasks across different domains. However, applying LLMs to clinical settings still poses significant challenges, as it requires specialized knowledge, vocabulary, as well as reliability. In this work, we propose a novel method of instruction fine-tuning for adapting LLMs to the clinical domain, which leverages the instruction-following capabilities of LLMs and the availability of diverse real-world data sources. We generate instructions, inputs, and outputs covering a wide spectrum of clinical services, from primary cares to nursing, radiology, physician, and social work, and use them to fine-tune LLMs. We evaluated the fine-tuned LLM, LlamaCare, on various clinical tasks, such as generating discharge summaries, predicting mortality and length of stay, and more. Using both automatic and human metrics, we demonstrated that LlamaCare surpasses other LLM baselines in predicting clinical outcomes and producing more accurate and coherent clinical texts. We also discuss the challenges and limitations of LLMs that need to be addressed before they can be widely adopted in clinical settings. Rumeng Li, Hong Yu 0001 |
LREC/COLING | 3 |
| 2024 | SYNFAC-EDIT: Synthetic Imitation Edit Feedback for Factual Alignment in Clinical SummarizationabstractLarge Language Models (LLMs) such as GPT & Llama have demonstrated significant achievements in summarization tasks but struggle with factual inaccuracies, a critical issue in clinical NLP applications where errors could lead to serious consequences.To counter the high costs and limited availability of expert-annotated data for factual alignment, this study introduces an innovative pipeline that utilizes >100B parameter GPT variants like GPT-3.5 & GPT-4 to act as synthetic experts to generate high-quality synthetics feedback aimed at enhancing factual consistency in clinical note summarization.Our research primarily focuses on edit feedback generated by these synthetic feedback experts without additional human annotations, mirroring and optimizing the practical scenario in which medical professionals refine AI system outputs.Although such 100B+ parameter GPT variants have proven to demonstrate expertise in various clinical NLP tasks, such as the Medical Licensing Examination, there is scant research on their capacity to act as synthetic feedback experts and deliver expert-level edit feedback for improving the generation quality of weaker (<10B parameter) LLMs like GPT-2 (1.5B) & Llama 2 (7B) in clinical domain.So in this work, we leverage 100B+ GPT variants to act as synthetic feedback experts offering expert-level edit feedback, that is used to reduce hallucinations and align weaker (<10B parameter) LLMs with medical facts using two distinct alignment algorithms (DPO & SALT), endeavoring to narrow the divide between AIgenerated content and factual accuracy.This highlights the substantial potential of LLMbased synthetic edits in enhancing the alignment of clinical factuality 1 . * indicates equal contribution † Presently in AMD AI Prakamya Mishra, Zonghai Yao, Parth Vashisht, Feiyun Ouyang, Beining Wang, Vidhi Dhaval Mody, Hong Yu 0001 |
EMNLP | 7 |
| 2024 | ODD: A Benchmark Dataset for the Natural Language Processing Based Opioid Related Aberrant Behavior DetectionabstractOpioid related aberrant behaviors (ORABs) present novel risk factors for opioid overdose. This paper introduces a novel biomedical natural language processing benchmark dataset named ODD, for ORAB Detection Dataset. ODD is an expert-annotated dataset designed to identify ORABs from patients' EHR notes and classify them into nine categories; 1) Confirmed Aberrant Behavior, 2) Suggested Aberrant Behavior, 3) Opioids, 4) Indication, 5) Diagnosed opioid dependency, 6) Benzodiazepines, 7) Medication Changes, 8) Central Nervous System-related, and 9) Social Determinants of Health. We explored two state-of-the-art natural language processing models (fine-tuning and prompt-tuning approaches) to identify ORAB. Experimental results show that the prompt-tuning models outperformed the fine-tuning models in most categories and the gains were especially higher among uncommon categories (Suggested Aberrant Behavior, Confirmed Aberrant Behaviors, Diagnosed Opioid Dependence, and Medication Change). Although the best model achieved the highest 88.17% on macro average area under precision recall curve, uncommon classes still have a large room for performance improvement. ODD is publicly available. Sunjae Kwon, Weisong Liu, Emily Druhl, Minhee L. Sung, Joel I. Reisman, Robert D. Kerns, William Becker, Hong Yu 0001 |
NAACL-HLT | 10 |
| 2024 | BioInstruct: instruction tuning of large language models for biomedical natural language processingabstractOBJECTIVES: To enhance the performance of large language models (LLMs) in biomedical natural language processing (BioNLP) by introducing a domain-specific instruction dataset and examining its impact when combined with multi-task learning principles. MATERIALS AND METHODS: We created the BioInstruct, comprising 25 005 instructions to instruction-tune LLMs (LLaMA 1 and 2, 7B and 13B version). The instructions were created by prompting the GPT-4 language model with 3-seed samples randomly drawn from an 80 human curated instructions. We employed Low-Rank Adaptation (LoRA) for parameter-efficient fine-tuning. We then evaluated these instruction-tuned LLMs on several BioNLP tasks, which can be grouped into 3 major categories: question answering (QA), information extraction (IE), and text generation (GEN). We also examined whether categories (eg, QA, IE, and generation) of instructions impact model performance. RESULTS AND DISCUSSION: Comparing with LLMs without instruction-tuned, our instruction-tuned LLMs demonstrated marked performance gains: 17.3% in QA on average accuracy metric, 5.7% in IE on average F1 metric, and 96% in Generation tasks on average GPT-4 score metric. Our 7B-parameter instruction-tuned LLaMA 1 model was competitive or even surpassed other LLMs in the biomedical domain that were also fine-tuned from LLaMA 1 with vast domain-specific data or a variety of tasks. Our results also show that the performance gain is significantly higher when instruction fine-tuning is conducted with closely related tasks. Our findings align with the observations of multi-task learning, suggesting the synergies between 2 tasks. CONCLUSION: The BioInstruct dataset serves as a valuable resource and instruction tuned LLMs lead to the best performing BioNLP applications. Hieu Tran, Zhichao Yang 0001, Zonghai Yao, Hong Yu 0001 |
J. Am. Medical Informatics Assoc. | 4 |
| 2023 | Multi-Label Few-Shot ICD Coding as Autoregressive Generation with PromptabstractAutomatic International Classification of Diseases (ICD) coding aims to assign multiple ICD codes to a medical note with an average of 3,000+ tokens. This task is challenging due to the high-dimensional space of multi-label assignment (155,000+ ICD code candidates) and the long-tail challenge - Many ICD codes are infrequently assigned yet infrequent ICD codes are important clinically. This study addresses the long-tail challenge by transforming this multi-label classification task into an autoregressive generation task. Specifically, we first introduce a novel pretraining objective to generate free text diagnosis and procedure descriptions using the SOAP structure, the medical logic physicians use for note documentation. Second, instead of directly predicting the high dimensional space of ICD codes, our model generates the lower dimension of text descriptions, which then infer ICD codes. Third, we designed a novel prompt template for multi-label classification. We evaluate our Generation with Prompt (GP) model with the benchmark of all code assignment (MIMIC-III-full) and few shot ICD code assignment evaluation benchmark (MIMIC-III-few). Experiments on MIMIC-III-few show that our model performs with a marco F1 30.2, which substantially outperforms the previous MIMIC-III-full SOTA model (marco F1 4.3) and the model specifically designed for few/zero shot setting (marco F1 18.7). Finally, we design a novel ensemble learner, a cross attention reranker with prompts, to integrate previous SOTA and our best few-shot coding predictions. Experiments on MIMIC-III-full show that our ensemble learner substantially improves both macro and micro F1, from 10.4 to 14.6 and from 58.2 to 59.1, respectively. Zhichao Yang 0001, Sunjae Kwon, Zonghai Yao, Hong Yu 0001 |
AAAI | 4 |
| 2023 | Generating User-Engaging News HeadlinesabstractPengshan Cai, Kaiqiang Song, Sangwoo Cho, Hongwei Wang, Xiaoyang Wang, Hong Yu, Fei Liu, Dong Yu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Pengshan Cai, Kaiqiang Song, Sangwoo Cho, Hongwei Wang 0010, Xiaoyang Wang 0001, Hong Yu 0001, Fei Liu 0004, Dong Yu 0001 |
ACL (1) | 6 |
| 2023 | Vision Meets Definitions: Unsupervised Visual Word Sense Disambiguation Incorporating Gloss InformationabstractVisual Word Sense Disambiguation (VWSD) is a task to find the image that most accurately depicts the correct sense of the target word for the given context.Previously, image-text matching models often suffered from recognizing polysemous words.This paper introduces an unsupervised VWSD approach that uses gloss information of an external lexical knowledge-base, especially the sense definitions.Specifically, we suggest employing Bayesian inference to incorporate the sense definitions when sense information of the answer is not provided.In addition, to ameliorate the out-of-vocabulary (OOV) issue, we propose a context-aware definition generation with GPT-3.Experimental results show that VWSD performance increased significantly with our Bayesian inference-based approach.In addition, our context-aware definition generation achieved prominent performance improvement in OOV examples exhibiting better performance than the existing definition generation method. Sunjae Kwon, Rishabh Garodia, Minhwa Lee, Zhichao Yang 0001, Hong Yu 0001 |
ACL (1) | 5 |
| 2023 | Web Information Extraction for Social Good: Food Pantry Answering As an ExampleabstractSocial determinants of health (SDH) are the conditions in which we are born, live, work, and age. Food insecurity (FI) is an important domain of SDH. FI is associated with poor health outcomes. Food bank/pantry (food pantry) directly addresses FI. Improving the availability and quality of food from food pantries could reduce FI, leading to improved health outcomes. However, it is difficult for a client to access food pantry information. In this study, we built a food pantry answering framework by combining location-aware information retrieval, web information extraction and domain-specific answering. Our proposed framework first retrieves pantry candidates based on geolocation of the client, and utilizes structural information from markup language to extract semantic chunks related to six common client requests. We use BERT and RoBERTa as information extraction models and compare three different web page segmentation methods in the experiments. Huan-Yuan Chen, Hong Yu 0001 |
WWW | 2 |
| 2023 | Automated identification of eviction status from electronic health record notesabstractOBJECTIVE: Evictions are important social and behavioral determinants of health. Evictions are associated with a cascade of negative events that can lead to unemployment, housing insecurity/homelessness, long-term poverty, and mental health problems. In this study, we developed a natural language processing system to automatically detect eviction status from electronic health record (EHR) notes. MATERIALS AND METHODS: We first defined eviction status (eviction presence and eviction period) and then annotated eviction status in 5000 EHR notes from the Veterans Health Administration (VHA). We developed a novel model, KIRESH, that has shown to substantially outperform other state-of-the-art models such as fine-tuning pretrained language models like BioBERT and Bio_ClinicalBERT. Moreover, we designed a novel prompt to further improve the model performance by using the intrinsic connection between the 2 subtasks of eviction presence and period prediction. Finally, we used the Temperature Scaling-based Calibration on our KIRESH-Prompt method to avoid overconfidence issues arising from the imbalance dataset. RESULTS: KIRESH-Prompt substantially outperformed strong baseline models including fine-tuning the Bio_ClinicalBERT model to achieve 0.74672 MCC, 0.71153 Macro-F1, and 0.83396 Micro-F1 in predicting eviction period and 0.66827 MCC, 0.62734 Macro-F1, and 0.7863 Micro-F1 in predicting eviction presence. We also conducted additional experiments on a benchmark social determinants of health (SBDH) dataset to demonstrate the generalizability of our methods. CONCLUSION AND FUTURE WORK: KIRESH-Prompt has substantially improved eviction status classification. We plan to deploy KIRESH-Prompt to the VHA EHRs as an eviction surveillance system to help address the US Veterans' housing insecurity. Zonghai Yao, Jack Tsai, Weisong Liu, David A. Levy, Emily Druhl, Joel I. Reisman, Hong Yu 0001 |
J. Am. Medical Informatics Assoc. | 7 |
| 2023 | PaniniQA: Enhancing Patient Education Through Interactive Question AnsweringabstractAbstract A patient portal allows discharged patients to access their personalized discharge instructions in electronic health records (EHRs). However, many patients have difficulty understanding or memorizing their discharge instructions (Zhao et al., 2017). In this paper, we present PaniniQA, a patient-centric interactive question answering system designed to help patients understand their discharge instructions. PaniniQA first identifies important clinical content from patients’ discharge instructions and then formulates patient-specific educational questions. In addition, PaniniQA is also equipped with answer verification functionality to provide timely feedback to correct patients’ misunderstandings. Our comprehensive automatic & human evaluation results demonstrate our PaniniQA is capable of improving patients’ mastery of their medical instructions through effective interactions.1 Pengshan Cai, Zonghai Yao, Fei Liu 0004, Dakuo Wang, Meghan Reilly, Huixue Zhou, Alok Kapoor, Adarsha Bajracharya, Dan Berlowitz, Hong Yu 0001 |
Trans. Assoc. Comput. Linguistics | 12 |
| 2022 | An Investigation of the Representation of Social Determinants of Health in the UMLS
Bhanu Pratap Singh Rawat, Hong Yu 0001, Heather Keating, Raelene Goodwin, Emily Druhl |
AMIA | 2 |
| 2022 | Extracting Biomedical Factual Knowledge Using Pretrained Language Model and Electronic Health Record Context
Zonghai Yao, Zhichao Yang 0001, Vijeta Deshpande, Hong Yu 0001 |
AMIA | 5 |
| 2022 | Generation of Patient After-Visit Summaries to Support PhysiciansabstractAn after-visit summary (AVS) is a summary note given to patients after their clinical visit. It recaps what happened during their clinical visit and guides patients’ disease self-management. Studies have shown that a majority of patients found after-visit summaries useful. However, many physicians face excessive workloads and do not have time to write clear and informative summaries. In this paper, we study the problem of automatic generation of after-visit summaries and examine whether those summaries can convey the gist of clinical visits. We report our findings on a new clinical dataset that contains a large number of electronic health record (EHR) notes and their associated summaries. Our results suggest that generation of lay language after-visit summaries remains a challenging task. Crucially, we introduce a feedback mechanism that alerts physicians when an automatic summary fails to capture the important details of the clinical notes or when it contains hallucinated facts that are potentially detrimental to the summary quality. Automatic and human evaluation demonstrates the effectiveness of our approach in providing writing feedback and supporting physicians. Pengshan Cai, Fei Liu 0004, Adarsha Bajracharya, Joe Sills, Alok Kapoor, Weisong Liu, Dan Berlowitz, David A. Levy, Richeek Pradhan, Hong Yu 0001 |
COLING | 10 |
| 2022 | MedJEx: A Medical Jargon Extraction Model with Wiki's Hyperlink Span and Contextualized Masked Language Model ScoreabstractThis paper proposes a new natural language processing (NLP) application for identifying medical jargon terms potentially difficult for patients to comprehend from electronic health record (EHR) notes.We first present a novel and publicly available dataset with expertannotated medical jargon terms from 18K+ EHR note sentences (M edJ).Then, we introduce a novel medical jargon extraction (M edJEx) model which has been shown to outperform existing state-of-the-art NLP models.First, MedJEx improved the overall performance when it was trained on an auxiliary Wikipedia hyperlink span dataset, where hyperlink spans provide additional Wikipedia articles to explain the spans (or terms), and then fine-tuned on the annotated MedJ data.Secondly, we found that a contextualized masked language model score was beneficial for detecting domain-specific unfamiliar jargon terms.Moreover, our results show that training on the auxiliary Wikipedia hyperlink span datasets improved six out of eight biomedical named entity recognition benchmark datasets.MedJEx is publicly available 1 . Sunjae Kwon, Zonghai Yao, Harmon S. Jordan, David A. Levy, Brian Corner, Hong Yu 0001 |
EMNLP | 6 |
| 2022 | Learning as Conversation: Dialogue Systems Reinforced for Information AcquisitionabstractPengshan Cai, Hui Wan, Fei Liu, Mo Yu, Hong Yu, Sachindra Joshi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Pengshan Cai, Hui Wan 0001, Fei Liu 0004, Mo Yu, Hong Yu 0001, Sachindra Joshi |
NAACL-HLT | 5 |
| 2022 | ScAN: Suicide Attempt and Ideation Events Datasetabstractto identify the type of suicidal behavior (SA and SI) concluded during the patient's stay at the hospital. ScANER achieved a macro-weighted F1-score of 0.83 for identifying suicidal behavioral evidences and a macro F1-score of 0.78 and 0.60 for classification of SA and SI for the patient's hospital-stay, respectively. ScAN and ScANER are publicly available. Bhanu Pratap Singh Rawat, Samuel Kovaly, Hong Yu 0001, Wilfred R. Pigeon |
NAACL-HLT | 3 |
| 2021 | Improving Formality Style Transfer with Context-Aware Rule InjectionabstractModels pre-trained on large-scale regular text corpora often do not work well for user-generated data where the language styles differ significantly from the mainstream text. Here we present Context-Aware Rule Injection (CARI), an innovative method for formality style transfer (FST). CARI injects multiple rules into an end-to-end BERT-based encoder and decoder model. It learns to select optimal rules based on context. The intrinsic evaluation showed that CARI achieved the new highest performance on the FST benchmark dataset. Our extrinsic evaluation showed that CARI can greatly improve the regular pre-trained models' performance on several tweet sentiment analysis tasks. Zonghai Yao, Hong Yu 0001 |
ACL/IJCNLP (1) | 2 |
| 2021 | Neural data-to-text generation with dynamic content planning
Kai Chen 0020, Fayuan Li, Baotian Hu, Weihua Peng, Qingcai Chen, Hong Yu 0001, Yang Xiang 0003 |
Knowl. Based Syst. | 6 |
| 2020 | ICD Coding from Clinical Text Using Multi-Filter Residual Convolutional Neural NetworkabstractAutomated ICD coding, which assigns the International Classification of Disease codes to patient visits, has attracted much research attention since it can save time and labor for billing. The previous state-of-the-art model utilized one convolutional layer to build document representations for predicting ICD codes. However, the lengths and grammar of text fragments, which are closely related to ICD coding, vary a lot in different documents. Therefore, a flat and fixed-length convolutional architecture may not be capable of learning good document representations. In this paper, we proposed a Multi-Filter Residual Convolutional Neural Network (MultiResCNN) for ICD coding. The innovations of our model are two-folds: it utilizes a multi-filter convolutional layer to capture various text patterns with different lengths and a residual convolutional layer to enlarge the receptive field. We evaluated the effectiveness of our model on the widely-used MIMIC dataset. On the full code set of MIMIC-III, our model outperformed the state-of-the-art model in 4 out of 6 evaluation metrics. On the top-50 code set of MIMIC-III and the full code set of MIMIC-II, our model outperformed all the existing and state-of-the-art models in all evaluation metrics. The code is available at https://github.com/foxlf823/Multi-Filter-Residual-Convolutional-Neural-Network. Fei Li 0021, Hong Yu 0001 |
AAAI | 2 |
| 2020 | MetaMT, a Meta Learning Method Leveraging Multiple Domain Data for Low Resource Machine TranslationabstractNeural machine translation (NMT) models have achieved state-of-the-art translation quality with a large quantity of parallel corpora available. However, their performance suffers significantly when it comes to domain-specific translations, in which training data are usually scarce. In this paper, we present a novel NMT model with a new word embedding transition technique for fast domain adaption. We propose to split parameters in the model into two groups: model parameters and meta parameters. The former are used to model the translation while the latter are used to adjust the representational space to generalize the model to different domains. We mimic the domain adaptation of the machine translation model to low-resource domains using multiple translation tasks on different domains. A new training strategy based on meta-learning is developed along with the proposed model to update the model parameters and meta parameters alternately. Experiments on datasets of different domains showed substantial improvements of NMT performances on a limited amount of data. Rumeng Li, Hong Yu 0001 |
AAAI | 3 |
| 2020 | Calibrating Structured Output Predictors for Natural Language ProcessingabstractWe address the problem of calibrating prediction confidence for output entities of interest in natural language processing (NLP) applications. It is important that NLP applications such as named entity recognition and question answering produce calibrated confidence scores for their predictions, especially if the applications are to be deployed in a safety-critical domain such as healthcare. However, the output space of such structured prediction models is often too large to adapt binary or multi-class calibration methods directly. In this study, we propose a general calibration scheme for output entities of interest in neural network based structured prediction models. Our proposed method can be used with any binary class calibration scheme and a neural network model. Additionally, we show that our calibration method can also be used as an uncertainty-aware, entity-specific decoding step to improve the performance of the underlying model at no additional training cost or data requirements. We show that our method outperforms current calibration techniques for named-entity-recognition, part-of-speech and question answering. We also improve our model's performance from our decoding step across several tasks and benchmark datasets. Our method improves the calibration and model performance on out-ofdomain test scenarios as well. Abhyuday Jagannatha, Hong Yu 0001 |
ACL | 2 |
| 2020 | Neural Multi-Task Learning for Adverse Drug Reaction Extraction
Hong Yu 0001, Jennifer Tjia |
AMIA | 3 |
| 2020 | Bleeding Entity Recognition in Electronic Health Records: A Comprehensive Analysis of End-to-End Systems
Avijit Mitra, Bhanu Pratap Singh Rawat, David McManus, Alok Kapoor, Hong Yu 0001 |
AMIA | 5 |
| 2020 | Inferring ADR causality by predicting the Naranjo Score from Clinical Notes
Bhanu Pratap Singh Rawat, Abhyuday Jagannatha, Hong Yu 0001 |
AMIA | 4 |
| 2020 | Conversational Machine Comprehension: a Literature ReviewabstractConversational Machine Comprehension (CMC), a research track in conversational AI, expects the machine to understand an open-domain natural language text and thereafter engage in a multi-turn conversation to answer questions related to the text.While most of the research in Machine Reading Comprehension (MRC) revolves around single-turn question answering (QA), multi-turn CMC has recently gained prominence, thanks to the advancement in natural language understanding via neural language models such as BERT and the introduction of large-scale conversational datasets such as CoQA and QuAC.The rise in interest has, however, led to a flurry of concurrent publications, each with a different yet structurally similar modeling approach and an inconsistent view of the surrounding literature.With the volume of model submissions to conversational datasets increasing every year, there exists a need to consolidate the scattered knowledge in this domain to streamline future research.This literature review attempts at providing a holistic overview of CMC with an emphasis on the common trends across recently published models, specifically in their approach to tackling conversational history.The review synthesizes a generic framework for CMC models while highlighting the differences in recent approaches and intends to serve as a compendium of CMC for future researchers. Somil Gupta, Bhanu Pratap Singh Rawat, Hong Yu 0001 |
COLING | 3 |
| 2019 | Extracting Drug-drug Interactions with a Dependency-based Graph Convolution Neural NetworkabstractDrug-drug interactions (DDIs) play a key role in various applications of biomedicine such as pharmacovigilance. DDIs are frequently reported in biomedical publications, making them an effective source for DDI extraction. Although neural networks have achieved competitive performances in DDI extraction, previous work depended on dependency paths to remove noise from the sentences of biomedical publications. However, such method may ignore crucial information about DDIs. Effectively exploiting much dependency information can improve DDI extraction. In this article, we proposed a model that combines the graph convolution neural network (GCNN) and bidirectional long short-term memory (BiLSTM) to extract DDI interactions from entire dependency graphs of sentences. We evaluated our model in the benchmark corpus for this domain, namely the DDIExtraction 2013 corpus. Our model achieved the state-of-the-art result (77.0% in F1), which are superior to the results reported in previous work. The code is available at https://github.com/woodyXwt/DDI_extraction. Wuti Xiong, Fei Li 0021, Hong Yu 0001, Donghong Ji |
BIBM | 3 |
| 2019 | Learning Latent Parameters without Human Response Patterns: Item Response Theory with Artificial CrowdsabstractIncorporating Item Response Theory (IRT) into NLP tasks can provide valuable information about model performance and behavior. Traditionally, IRT models are learned using human response pattern (RP) data, presenting a significant bottleneck for large data sets like those required for training deep neural networks (DNNs). In this work we propose learning IRT models using RPs generated from artificial crowds of DNN models. We demonstrate the effectiveness of learning IRT models using DNN-generated data through quantitative and qualitative analyses for two NLP tasks. Parameters learned from human and machine RPs for natural language inference and sentiment analysis exhibit medium to large positive correlations. We demonstrate a use-case for latent difficulty item parameters, namely training set filtering, and show that using difficulty to sample training data outperforms baseline methods. Finally, we highlight cases where human expectation about item difficulty does not match difficulty as estimated from the machine RPs. John Lalor, Hao Wu 0055, Hong Yu 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Generating Classical Chinese Poems from Vernacular ChineseabstractClassical Chinese poetry is a jewel in the treasure house of Chinese culture. Previous poem generation models only allow users to employ keywords to interfere the meaning of generated poems, leaving the dominion of generation to the model. In this paper, we propose a novel task of generating classical Chinese poems from vernacular, which allows users to have more control over the semantic of generated poems. We adapt the approach of unsupervised machine translation (UMT) to our task. We use segmentation-based padding and reinforcement learning to address under-translation and over-translation respectively. According to experiments, our approach significantly improve the perplexity and BLEU compared with typical UMT models. Furthermore, we explored guidelines on how to write the input vernacular to generate better poems. Human evaluation showed our approach can generate high-quality poems which are comparable to amateur poems. Zhichao Yang 0001, Pengshan Cai, Yansong Feng 0002, Fei Li 0021, Weijiang Feng, Elena Suet-Ying Chiu, Hong Yu 0001 |
EMNLP/IJCNLP (1) | 7 |
| 2019 | Naranjo Question Answering using End-to-End Multi-task Learning ModelabstractIn the clinical domain, it is important to understand whether an adverse drug reaction (ADR) is caused by a particular medication. Clinical judgement studies help judge the causal relation between a medication and its ADRs. In this study, we present the first attempt to automatically infer the causality between a drug and an ADR from electronic health records (EHRs) by answering the Naranjo questionnaire, the validated clinical question answering set used by domain experts for ADR causality assessment. Using physicians' annotation as the gold standard, our proposed joint model, which uses multi-task learning to predict the answers of a subset of the Naranjo questionnaire, significantly outperforms the baseline pipeline model with a good margin, achieving a macro-weighted f-score between 0.3652 - 0.5271 and micro-weighted f-score between 0.9523 - 0.9918. Bhanu Pratap Singh Rawat, Fei Li 0021, Hong Yu 0001 |
KDD | 3 |
| 2019 | An investigation of single-domain and multidomain medication and adverse drug event relation extraction from electronic health record notes using advanced deep learning modelsabstractOBJECTIVE: We aim to evaluate the effectiveness of advanced deep learning models (eg, capsule network [CapNet], adversarial training [ADV]) for single-domain and multidomain relation extraction from electronic health record (EHR) notes. MATERIALS AND METHODS: We built multiple deep learning models with increased complexity, namely a multilayer perceptron (MLP) model and a CapNet model for single-domain relation extraction and fully shared (FS), shared-private (SP), and adversarial training (ADV) modes for multidomain relation extraction. Our models were evaluated in 2 ways: first, we compared our models using our expert-annotated cancer (the MADE1.0 corpus) and cardio corpora; second, we compared our models with the systems in the MADE1.0 and i2b2 challenges. RESULTS: Multidomain models outperform single-domain models by 0.7%-1.4% in F1 (t test P < .05), but the results of FS, SP, and ADV modes are mixed. Our results show that the MLP model generally outperforms the CapNet model by 0.1%-1.0% in F1. In the comparisons with other systems, the CapNet model achieves the state-of-the-art result (87.2% in F1) in the cancer corpus and the MLP model generally outperforms MedEx in the cancer, cardiovascular diseases, and i2b2 corpora. CONCLUSIONS: Our MLP or CapNet model generally outperforms other state-of-the-art systems in medication and adverse drug event relation extraction. Multidomain models perform better than single-domain models. However, neither the SP nor the ADV mode can always outperform the FS mode significantly. Moreover, the CapNet model is not superior to the MLP model for our corpora. Fei Li 0021, Hong Yu 0001 |
J. Am. Medical Informatics Assoc. | 2 |
| 2019 | Learning to detect and understand drug discontinuation events from clinical narrativesabstractOBJECTIVE: Identifying drug discontinuation (DDC) events and understanding their reasons are important for medication management and drug safety surveillance. Structured data resources are often incomplete and lack reason information. In this article, we assessed the ability of natural language processing (NLP) systems to unlock DDC information from clinical narratives automatically. MATERIALS AND METHODS: We collected 1867 de-identified providers' notes from the University of Massachusetts Medical School hospital electronic health record system. Then 2 human experts chart reviewed those clinical notes to annotate DDC events and their reasons. Using the annotated data, we developed and evaluated NLP systems to automatically identify drug discontinuations and reasons at the sentence level using a novel semantic enrichment-based vector representation (SEVR) method for enhanced feature representation. RESULTS: Our SEVR-based NLP system achieved the best performance of 0.785 (AUC-ROC) for detecting discontinuation events and 0.745 (AUC-ROC) for identifying reasons when testing this highly imbalanced data, outperforming 2 state-of-the-art non-SEVR-based models. Compared with a rule-based baseline system for discontinuation detection, our system improved the sensitivity significantly (57.75% vs 18.31%, absolute value) while retaining a high specificity of 99.25%, leading to a significant improvement in AUC-ROC by 32.83% (absolute value). CONCLUSION: Experiments have shown that a high-performance NLP system can be developed to automatically identify DDCs and their reasons from providers' notes. The SEVR model effectively improved the system performance showing better generalization and robustness on unseen test data. Our work is an important step toward identifying reasons for drug discontinuation that will inform drug safety surveillance and pharmacovigilance. Richeek Pradhan, Emily Druhl, Elaine T. Freund, Weisong Liu, Brian C. Sauer, Fran Cunningham, Adam J. Gordon, Celena B. Peters, Hong Yu 0001 |
J. Am. Medical Informatics Assoc. | 10 |
| 2018 | Detecting Hypoglycemia Incidents from Patients' Secure Messages
Jinying Chen, John Lalor, Hong Yu 0001 |
AMIA | 3 |
| 2018 | Reference Standard Development to Train Natural Language Processing Algorithms to Detect Problematic Buprenorphine-Naloxone Therapy
Celena B. Peters, Fran Cunningham, Hong Yu 0001, Adam J. Gordon, Cedric L. Salone, Jessica Zacher, Ronald Carico, Jianwei Leng, Nikolh Smith-Nemard, Weisong Liu, Sophia Lu, Emily Druhl, Tina Huynh, Zachary Burningham, Brian C. Sauer |
AMIA | 3 |
| 2018 | Accuracy of International Classification of Disease Clinical Modification Codes for Detecting Bleeding Events in Electronic Health Records and When to Use Them
Victoria J. Wang, Richeek Pradhan, Harmon S. Jordan, Arlene S. Ash, David McManus, Hong Yu 0001 |
AMIA | 6 |
| 2018 | Understanding Deep Learning Performance through an Examination of Test Set Difficulty: A Psychometric Case StudyabstractInterpreting the performance of deep learning models beyond test set accuracy is challenging. Characteristics of individual data points are often not considered during evaluation, and each data point is treated equally. We examine the impact of a test set question's difficulty to determine if there is a relationship between difficulty and performance. We model difficulty using well-studied psychometric methods on human response patterns. Experiments on Natural Language Inference (NLI) and Sentiment Analysis (SA) show that the likelihood of answering a question correctly is impacted by the question's difficulty. As DNNs are trained with more data, easy examples are learned more quickly than hard examples. John Lalor, Hao Wu 0055, Tsendsuren Munkhdalai, Hong Yu 0001 |
EMNLP | 4 |
| 2017 | Generating a Test of Electronic Health Record Narrative Comprehension with Item Response Theory
John Lalor, Hao Wu 0055, Kathleen M. Mazor, Hong Yu 0001 |
AMIA | 5 |
| 2017 | A hybrid Neural Network Model for Joint Prediction of Medical Presence and Period Assertions in Clinical Notes
Rumeng Li, Abhyuday Jagannatha, Hong Yu 0001 |
AMIA | 3 |
| 2017 | Detecting Opioid-Related Aberrant Behavior using Natural Language Processing
Jesse M. Lingeman, Priscilla Wang, William Becker, Hong Yu 0001 |
AMIA | 4 |
| 2017 | Assessing Electronic Health Record Readability
Jiaping Zheng, Hong Yu 0001 |
AMIA | 2 |
| 2017 | Neural Tree Indexers for Text UnderstandingabstractRecurrent neural networks (RNNs) process input text sequentially and model the conditional transition between word tokens.In contrast, the advantages of recursive networks include that they explicitly model the compositionality and the recursive structure of natural language.However, the current recursive architecture is limited by its dependence on syntactic tree.In this paper, we introduce a robust syntactic parsing-independent tree structured model, Neural Tree Indexers (NTI) that provides a middle ground between the sequential RNNs and the syntactic treebased recursive models.NTI constructs a full n-ary tree by processing the input text with its node function in a bottom-up fashion.Attention mechanism can then be applied to both structure and node function.We implemented and evaluated a binarytree model of NTI, showing the model achieved the state-of-the-art performance on three different NLP tasks: natural language inference, answer sentence selection, and sentence classification, outperforming state-of-the-art recurrent and recursive neural networks 1 . Tsendsuren Munkhdalai, Hong Yu 0001 |
EACL (1) | 2 |
| 2017 | Neural Semantic EncodersabstractWe present a memory augmented neural network for natural language understanding: Neural Semantic Encoders.NSE is equipped with a novel memory update rule and has a variable sized encoding memory that evolves over time and maintains the understanding of input sequences through read, compose and write operations.NSE can also access 1 multiple and shared memories.In this paper, we demonstrated the effectiveness and the flexibility of NSE on five different natural language tasks: natural language inference, question answering, sentence classification, document sentiment analysis and machine translation where NSE achieved state-of-the-art performance when evaluated on publically available benchmarks.For example, our shared-memory model showed an encouraging result on neural machine translation, improving an attention-based baseline by approximately 1.0 BLEU. Tsendsuren Munkhdalai, Hong Yu 0001 |
EACL (1) | 2 |
| 2017 | Reasoning with Memory Augmented Neural Networks for Language Comprehension
Tsendsuren Munkhdalai, Hong Yu 0001 |
ICLR (Poster) | 2 |
| 2017 | Meta NetworksabstractNeural networks have been successfully applied in applications with a large amount of labeled data. However, the task of rapid generalization on new concepts with small training data while preserving performances on previously learned ones still presents a significant challenge to neural network models. In this work, we introduce a novel meta learning method, Meta Networks (MetaNet), that learns a meta-level knowledge across tasks and shifts its inductive biases via fast parameterization for rapid generalization. When evaluated on Omniglot and Mini-ImageNet benchmarks, our MetaNet models achieve a near human-level performance and outperform the baseline approaches by up to 6\% accuracy. We demonstrate several appealing properties of MetaNet relating to generalization and continual learning. Tsendsuren Munkhdalai, Hong Yu 0001 |
ICML | 2 |
| 2017 | Unsupervised ensemble ranking of terms in electronic health record notes based on their importance to patientsabstractBACKGROUND: Allowing patients to access their own electronic health record (EHR) notes through online patient portals has the potential to improve patient-centered care. However, EHR notes contain abundant medical jargon that can be difficult for patients to comprehend. One way to help patients is to reduce information overload and help them focus on medical terms that matter most to them. Targeted education can then be developed to improve patient EHR comprehension and the quality of care. OBJECTIVE: The aim of this work was to develop FIT (Finding Important Terms for patients), an unsupervised natural language processing (NLP) system that ranks medical terms in EHR notes based on their importance to patients. METHODS: We built FIT on a new unsupervised ensemble ranking model derived from the biased random walk algorithm to combine heterogeneous information resources for ranking candidate terms from each EHR note. Specifically, FIT integrates four single views (rankers) for term importance: patient use of medical concepts, document-level term salience, word co-occurrence based term relatedness, and topic coherence. It also incorporates partial information of term importance as conveyed by terms' unfamiliarity levels and semantic types. We evaluated FIT on 90 expert-annotated EHR notes and used the four single-view rankers as baselines. In addition, we implemented three benchmark unsupervised ensemble ranking methods as strong baselines. RESULTS: FIT achieved 0.885 AUC-ROC for ranking candidate terms from EHR notes to identify important terms. When including term identification, the performance of FIT for identifying important terms from EHR notes was 0.813 AUC-ROC. Both performance scores significantly exceeded the corresponding scores from the four single rankers (P<0.001). FIT also outperformed the three ensemble rankers for most metrics. Its performance is relatively insensitive to its parameter. CONCLUSIONS: FIT can automatically identify EHR terms important to patients. It may help develop future interventions to improve quality of care. By using unsupervised learning as well as a robust and flexible framework for information fusion, FIT can be readily applied to other domains and applications. Jinying Chen, Hong Yu 0001 |
J. Biomed. Informatics | 2 |
| 2016 | Structured prediction models for RNN based sequence labeling in clinical textabstractSequence labeling is a widely used method for named entity recognition and information extraction from unstructured natural language data. In clinical domain one major application of sequence labeling involves extraction of medical entities such as medication, indication, and side-effects from Electronic Health Record narratives. Sequence labeling in this domain, presents its own set of challenges and objectives. In this work we experimented with various CRF based structured learning models with Recurrent Neural Networks. We extend the previously studied LSTM-CRF models with explicit modeling of pairwise potentials. We also propose an approximate version of skip-chain CRF inference with RNN potentials. We use these methodologies for structured prediction in order to improve the exact phrase detection of various medical entities. Abhyuday Jagannatha, Hong Yu 0001 |
EMNLP | 2 |
| 2016 | Building an Evaluation Scale using Item Response TheoryabstractEvaluation of NLP methods requires testing against a previously vetted gold-standard test set and reporting standard metrics (accuracy/precision/recall/F1). The current assumption is that all items in a given test set are equal with regards to difficulty and discriminating power. We propose Item Response Theory (IRT) from psychometrics as an alternative means for gold-standard test-set generation and NLP system evaluation. IRT is able to describe characteristics of individual items - their difficulty and discriminating power - and can account for these characteristics in its estimation of human intelligence or ability for an NLP task. In this paper, we demonstrate IRT by generating a gold-standard test set for Recognizing Textual Entailment. By collecting a large number of human responses and fitting our IRT model, we show that our IRT model compares NLP systems with the performance in a human population and is able to provide more insight into system performance than standard evaluation metrics. We show that a high accuracy score does not always imply a high IRT score, which depends on the item characteristics and the response pattern. John Lalor, Hao Wu 0055, Hong Yu 0001 |
EMNLP | 3 |
| 2016 | Bidirectional RNN for Medical Event Detection in Electronic Health RecordsabstractSequence labeling for extraction of medical events and their attributes from unstructured text in Electronic Health Record (EHR) notes is a key step towards semantic understanding of EHRs. It has important applications in health informatics including pharmacovigilance and drug surveillance. The state of the art supervised machine learning models in this domain are based on Conditional Random Fields (CRFs) with features calculated from fixed context windows. In this application, we explored recurrent neural network frameworks and show that they significantly out-performed the CRF models. Abhyuday Jagannatha, Hong Yu 0001 |
HLT-NAACL | 2 |
| 2016 | Methods for linking EHR notes to education materials
Jiaping Zheng, Hong Yu 0001 |
Inf. Retr. J. | 2 |
| 2015 | Rethinking Document Retrieval for Scientific Literature: A Learning to Rank Approach
Jesse M. Lingeman, Hong Yu 0001 |
AMIA | 2 |
| 2015 | Key Concept Identification for Medical Information RetrievalabstractThe difficult language in Electronic Health Records (EHRs) presents a challenge to patients' understanding of their own conditions.One approach to lowering the barrier is to provide tailored patient education based on their own EHR notes.We are developing a system to retrieve EHR note-tailored online consumer oriented health education materials.We explored topic model and key concept identification methods to construct queries from the EHR notes.Our experiments show that queries using identified key concepts with pseudo-relevance feedback significantly outperform (over 10-fold improvement) the baseline system of using the full text note. Jiaping Zheng, Hong Yu 0001 |
EMNLP | 2 |
| 2014 | Automatically Detecting Acute Myocardial Infarction Events from EHR Text: A Preliminary Study
Jiaping Zheng, Jorge Yarzebski, Balaji Polepalli Ramesh, Robert Goldberg, Hong Yu 0001 |
AMIA | 5 |
| 2012 | MedTxting: Learning based and Knowledge Rich SMS-style Medical Text Contraction
Soheil Moosavinasab, Thomas K. Houston, Hong Yu 0001 |
AMIA | 4 |
| 2012 | Automatic discourse connective detection in biomedical textabstractOBJECTIVE: Relation extraction in biomedical text mining systems has largely focused on identifying clause-level relations, but increasing sophistication demands the recognition of relations at discourse level. A first step in identifying discourse relations involves the detection of discourse connectives: words or phrases used in text to express discourse relations. In this study supervised machine-learning approaches were developed and evaluated for automatically identifying discourse connectives in biomedical text. MATERIALS AND METHODS: Two supervised machine-learning models (support vector machines and conditional random fields) were explored for identifying discourse connectives in biomedical literature. In-domain supervised machine-learning classifiers were trained on the Biomedical Discourse Relation Bank, an annotated corpus of discourse relations over 24 full-text biomedical articles (~112,000 word tokens), a subset of the GENIA corpus. Novel domain adaptation techniques were also explored to leverage the larger open-domain Penn Discourse Treebank (~1 million word tokens). The models were evaluated using the standard evaluation metrics of precision, recall and F1 scores. RESULTS AND CONCLUSION: Supervised machine-learning approaches can automatically identify discourse connectives in biomedical text, and the novel domain adaptation techniques yielded the best performance: 0.761 F1 score. A demonstration version of the fully implemented classifier BioConn is available at: http://bioconn.askhermes.org. Balaji Polepalli Ramesh, Rashmi Prasad, Brian Harrington 0001, Hong Yu 0001 |
J. Am. Medical Informatics Assoc. | 5 |
| 2011 | Figure summarizer browser extensions for PubMed CentralabstractSUMMARY: Figures in biomedical articles present visual evidence for research facts and help readers understand the article better. However, when figures are taken out of context, it is difficult to understand their content. We developed a summarization algorithm to summarize the content of figures and used it in our figure search engine (http://figuresearch.askhermes.org/). In this article, we report on the development of web browser extensions for Mozilla Firefox, Google Chrome and Apple Safari to display summaries for figures in PubMed Central and NCBI Images. AVAILABILITY: The extensions can be downloaded from http://figuresearch.askhermes.org/articlesearch/extensions.php. Shashank Agarwal, Hong Yu 0001 |
Bioinform. | 2 |
| 2011 | Simple and efficient machine learning frameworks for identifying protein-protein interaction relevant articles and experimental methods used to study the interactionsabstractBACKGROUND: Protein-protein interaction (PPI) is an important biomedical phenomenon. Automatically detecting PPI-relevant articles and identifying methods that are used to study PPI are important text mining tasks. In this study, we have explored domain independent features to develop two open source machine learning frameworks. One performs binary classification to determine whether the given article is PPI relevant or not, named "Simple Classifier", and the other one maps the PPI relevant articles with corresponding interaction method nodes in a standardized PSI-MI (Proteomics Standards Initiative-Molecular Interactions) ontology, named "OntoNorm". RESULTS: We evaluated our system in the context of BioCreative challenge competition using the standardized data set. Our systems are amongst the top systems reported by the organizers, attaining 60.8% F1-score for identifying relevant documents, and 52.3% F1-score for mapping articles to interaction method ontology. CONCLUSION: Our results show that domain-independent machine learning frameworks can perform competitively well at the tasks of detecting PPI relevant articles and identifying the methods that were used to study the interaction in such articles. AVAILABILITY: Simple Classifier is available at http://sourceforge.net/p/simpleclassify/home/ and OntoNorm at http://sourceforge.net/p/ontonorm/home/. Shashank Agarwal, Hong Yu 0001 |
BMC Bioinform. | 3 |
| 2011 | BioNOT: A searchable database of biomedical negated sentencesabstractBACKGROUND: Negated biomedical events are often ignored by text-mining applications; however, such events carry scientific significance. We report on the development of BioNØT, a database of negated sentences that can be used to extract such negated events. DESCRIPTION: Currently BioNØT incorporates ≈32 million negated sentences, extracted from over 336 million biomedical sentences from three resources: ≈2 million full-text biomedical articles in Elsevier and the PubMed Central, as well as ≈20 million abstracts in PubMed. We evaluated BioNØT on three important genetic disorders: autism, Alzheimer's disease and Parkinson's disease, and found that BioNØT is able to capture negated events that may be ignored by experts. CONCLUSIONS: The BioNØT database can be a useful resource for biomedical researchers. BioNØT is freely available at http://bionot.askhermes.org/. In future work, we will develop semantic web related technologies to enrich BioNØT. Shashank Agarwal, Hong Yu 0001, Isaac S. Kohane |
BMC Bioinform. | 2 |
| 2011 | The Biomedical Discourse Relation BankabstractBACKGROUND: Identification of discourse relations, such as causal and contrastive relations, between situations mentioned in text is an important task for biomedical text-mining. A biomedical text corpus annotated with discourse relations would be very useful for developing and evaluating methods for biomedical discourse processing. However, little effort has been made to develop such an annotated resource. RESULTS: We have developed the Biomedical Discourse Relation Bank (BioDRB), in which we have annotated explicit and implicit discourse relations in 24 open-access full-text biomedical articles from the GENIA corpus. Guidelines for the annotation were adapted from the Penn Discourse TreeBank (PDTB), which has discourse relations annotated over open-domain news articles. We introduced new conventions and modifications to the sense classification. We report reliable inter-annotator agreement of over 80% for all sub-tasks. Experiments for identifying the sense of explicit discourse connectives show the connective itself as a highly reliable indicator for coarse sense classification (accuracy 90.9% and F1 score 0.89). These results are comparable to results obtained with the same classifier on the PDTB data. With more refined sense classification, there is degradation in performance (accuracy 69.2% and F1 score 0.28), mainly due to sparsity in the data. The size of the corpus was found to be sufficient for identifying the sense of explicit connectives, with classifier performance stabilizing at about 1900 training instances. Finally, the classifier performs poorly when trained on PDTB and tested on BioDRB (accuracy 54.5% and F1 score 0.57). CONCLUSION: Our work shows that discourse relations can be reliably annotated in biomedical text. Coarse sense disambiguation of explicit connectives can be done with high reliability by using just the connective as a feature, but more refined sense classification requires either richer features or more annotated data. The poor performance of a classifier trained in the open domain and tested in the biomedical domain suggests significant differences in the semantic usage of connectives across these domains, and provides robust evidence for a biomedical sublanguage for discourse and the need to develop a specialized biomedical discourse annotated corpus. The results of our cross-domain experiments are consistent with related work on identifying connectives in BioDRB. Rashmi Prasad, Susan McRoy, Nadya Frid, Aravind K. Joshi, Hong Yu 0001 |
BMC Bioinform. | 5 |
| 2011 | Towards spoken clinical-question answering: evaluating and adapting automatic speech-recognition systems for spoken clinical questionsabstractOBJECTIVE: To evaluate existing automatic speech-recognition (ASR) systems to measure their performance in interpreting spoken clinical questions and to adapt one ASR system to improve its performance on this task. DESIGN AND MEASUREMENTS: The authors evaluated two well-known ASR systems on spoken clinical questions: Nuance Dragon (both generic and medical versions: Nuance Gen and Nuance Med) and the SRI Decipher (the generic version SRI Gen). The authors also explored language model adaptation using more than 4000 clinical questions to improve the SRI system's performance, and profile training to improve the performance of the Nuance Med system. The authors reported the results with the NIST standard word error rate (WER) and further analyzed error patterns at the semantic level. RESULTS: Nuance Gen and Med systems resulted in a WER of 68.1% and 67.4% respectively. The SRI Gen system performed better, attaining a WER of 41.5%. After domain adaptation with a language model, the performance of the SRI system improved 36% to a final WER of 26.7%. CONCLUSION: Without modification, two well-known ASR systems do not perform well in interpreting spoken clinical questions. With a simple domain adaptation, one of the ASR systems improved significantly on the clinical question task, indicating the importance of developing domain/genre-specific ASR systems. Gökhan Tür, Dilek Hakkani-Tür, Hong Yu 0001 |
J. Am. Medical Informatics Assoc. | 4 |
| 2011 | AskHERMES: An online question answering system for complex clinical questionsabstractOBJECTIVE: Clinical questions are often long and complex and take many forms. We have built a clinical question answering system named AskHERMES to perform robust semantic analysis on complex clinical questions and output question-focused extractive summaries as answers. DESIGN: This paper describes the system architecture and a preliminary evaluation of AskHERMES, which implements innovative approaches in question analysis, summarization, and answer presentation. Five types of resources were indexed in this system: MEDLINE abstracts, PubMed Central full-text articles, eMedicine documents, clinical guidelines and Wikipedia articles. MEASUREMENT: We compared the AskHERMES system with Google (Google and Google Scholar) and UpToDate and asked physicians to score the three systems by ease of use, quality of answer, time spent, and overall performance. RESULTS: AskHERMES allows physicians to enter a question in a natural way with minimal query formulation and allows physicians to efficiently navigate among all the answer sentences to quickly meet their information needs. In contrast, physicians need to formulate queries to search for information in Google and UpToDate. The development of the AskHERMES system is still at an early stage, and the knowledge resource is limited compared with Google or UpToDate. Nevertheless, the evaluation results show that AskHERMES' performance is comparable to the other systems. In particular, when answering complex clinical questions, it demonstrates the potential to outperform both Google and UpToDate systems. CONCLUSIONS: AskHERMES, available at http://www.AskHERMES.org, has the potential to help physicians practice evidence-based medicine and improve the quality of patient care. Yonggang Cao, Pippa Simpson, Lamont D. Antieau, Andrew S. Bennett, James J. Cimino, John W. Ely, Hong Yu 0001 |
J. Biomed. Informatics | 8 |
| 2011 | Automatic figure classification in bioscience literature
Balaji Polepalli Ramesh, Hong Yu 0001 |
J. Biomed. Informatics | 3 |
| 2011 | Toward automated consumer question answering: Automatically separating consumer questions from professional questions in the healthcare domain
Lamont D. Antieau, Hong Yu 0001 |
J. Biomed. Informatics | 3 |
| 2010 | Biomedical negation scope detection with conditional random fieldsabstractOBJECTIVE: Negation is a linguistic phenomenon that marks the absence of an entity or event. Negated events are frequently reported in both biological literature and clinical notes. Text mining applications benefit from the detection of negation and its scope. However, due to the complexity of language, identifying the scope of negation in a sentence is not a trivial task. DESIGN: Conditional random fields (CRF), a supervised machine-learning algorithm, were used to train models to detect negation cue phrases and their scope in both biological literature and clinical notes. The models were trained on the publicly available BioScope corpus. MEASUREMENT: The performance of the CRF models was evaluated on identifying the negation cue phrases and their scope by calculating recall, precision and F1-score. The models were compared with four competitive baseline systems. RESULTS: The best CRF-based model performed statistically better than all baseline systems and NegEx, achieving an F1-score of 98% and 95% on detecting negation cue phrases and their scope in clinical notes, and an F1-score of 97% and 85% on detecting negation cue phrases and their scope in biological literature. CONCLUSIONS: This approach is robust, as it can identify negation scope in both biological and clinical text. To benefit text mining applications, the system is publicly available as a Java API and as an online application at http://negscope.askhermes.org. Shashank Agarwal, Hong Yu 0001 |
J. Am. Medical Informatics Assoc. | 2 |
| 2010 | Lancet: a high precision medication event extraction system for clinical textabstractOBJECTIVE: This paper presents Lancet, a supervised machine-learning system that automatically extracts medication events consisting of medication names and information pertaining to their prescribed use (dosage, mode, frequency, duration and reason) from lists or narrative text in medical discharge summaries. DESIGN: Lancet incorporates three supervised machine-learning models: a conditional random fields model for tagging individual medication names and associated fields, an AdaBoost model with decision stump algorithm for determining which medication names and fields belong to a single medication event, and a support vector machines disambiguation model for identifying the context style (narrative or list). MEASUREMENTS: The authors, from the University of Wisconsin-Milwaukee, participated in the third i2b2 shared-task for challenges in natural language processing for clinical data: medication extraction challenge. With the performance metrics provided by the i2b2 challenge, the micro F1 (precision/recall) scores are reported for both the horizontal and vertical level. RESULTS: Among the top 10 teams, Lancet achieved the highest precision at 90.4% with an overall F1 score of 76.4% (horizontal system level with exact match), a gain of 11.2% and 12%, respectively, compared with the rule-based baseline system jMerki. By combining the two systems, the hybrid system further increased the F1 score by 3.4% from 76.4% to 79.0%. CONCLUSIONS: Supervised machine-learning systems with minimal external knowledge resources can achieve a high precision with a competitive overall F1 score.Lancet based on this learning framework does not rely on expensive manually curated rules. The system is available online at http://code.google.com/p/lancet/. Zuofeng Li, Lamont D. Antieau, Yonggang Cao, Hong Yu 0001 |
J. Am. Medical Informatics Assoc. | 5 |
| 2010 | Detecting hedge cues and their scope in biomedical text with conditional random fields
Shashank Agarwal, Hong Yu 0001 |
J. Biomed. Informatics | 2 |
| 2010 | Automatically extracting information needs from complex clinical questions
Yonggang Cao, James J. Cimino, John W. Ely, Hong Yu 0001 |
J. Biomed. Informatics | 4 |
| 2010 | An IR-Aided Machine Learning Framework for the BioCreative II.5 ChallengeabstractThe team at the University of Wisconsin-Milwaukee developed an information retrieval and machine learning framework. Our framework requires only the standardized training data and depends upon minimal external knowledge resources and minimal parsing. Within the framework, we built our text mining systems and participated for the first time in all three BioCreative II.5 Challenge tasks. The results show that our systems performed among the top five teams for raw F1 scores in all three tasks and came in third place for the homonym ortholog F1 scores for the INT task. The results demonstrated that our IR-based framework is efficient, robust, and potentially scalable. Yonggang Cao, Zuofeng Li, Shashank Agarwal, Hong Yu 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2009 | FigSum: Automatically Generating Structured Text Summaries for Figures in Biomedical Literature
Shashank Agarwal, Hong Yu 0001 |
AMIA | 2 |
| 2009 | Hierarchical Image Classification in the Bioscience Literature
Hong Yu 0001 |
AMIA | 2 |
| 2009 | Automatically classifying sentences in full-text biomedical articles into Introduction, Methods, Results and DiscussionabstractBiomedical texts can be typically represented by four rhetorical categories: Introduction, Methods, Results and Discussion (IMRAD). Classifying sentences into these categories can benefit many other text-mining tasks. Although many studies have applied different approaches for automatically classifying sentences in MEDLINE abstracts into the IMRAD categories, few have explored the classification of sentences that appear in full-text biomedical articles. We first evaluated whether sentences in full-text biomedical articles could be reliably annotated into the IMRAD format and then explored different approaches for automatically classifying these sentences into the IMRAD categories. Our results show an overall annotation agreement of 82.14% with a Kappa score of 0.756. The best classification system is a multinomial naïve Bayes classifier trained on manually annotated data that achieved 91.95% accuracy and an average F-score of 91.55%, which is significantly higher than baseline systems. A web version of this system is available online at-http://wood.ims.uwm.edu/full_text_classifier/. Shashank Agarwal, Hong Yu 0001 |
Bioinform. | 2 |
| 2008 | Automatically Extracting Information Needs from Ad Hoc Clinical Questions
Hong Yu 0001, Yonggang Cao |
AMIA | 1 |
| 2007 | Frontiers of biomedical text mining: current progressabstractIt is now almost 15 years since the publication of the first paper on text mining in the genomics domain, and decades since the first paper on text mining in the medical domain. Enormous progress has been made in the areas of information retrieval, evaluation methodologies and resource construction. Some problems, such as abbreviation-handling, can essentially be considered solved problems, and others, such as identification of gene mentions in text, seem likely to be solved soon. However, a number of problems at the frontiers of biomedical text mining continue to present interesting challenges and opportunities for great improvements and interesting research. In this article we review the current state of the art in biomedical text mining or 'BioNLP' in general, focusing primarily on papers published within the past year. Pierre Zweigenbaum, Dina Demner-Fushman, Hong Yu 0001, Kevin Cohen 0001 |
Briefings Bioinform. | 3 |
| 2007 | Natural language processing and visualization in the molecular imaging domain
P. Karina Tulipano, Ying Tao, William S. Millar, Pat Zanzonico, Katherine Kolbert, Hua Xu 0001, Hong Yu 0001, Lifeng Chen, Yves A. Lussier, Carol Friedman |
J. Biomed. Informatics | 7 |
| 2007 | Using MEDLINE as a knowledge source for disambiguating abbreviations and acronyms in full-text biomedical journal articles
Hong Yu 0001, Won Kim 0003, Vasileios Hatzivassiloglou, W. John Wilbur |
J. Biomed. Informatics | 1 |
| 2007 | Development, implementation, and a cognitive evaluation of a definitional question answering system for physicians
Hong Yu 0001, Minsuk Lee, David R. Kaufman, John W. Ely, Jerome A. Osheroff, George Hripcsak, James J. Cimino |
J. Biomed. Informatics | 1 |
| 2006 | Beyond Information Retrieval - Medical Question Answering
Minsuk Lee, James J. Cimino, Hai Ran Zhu, Carl L. Sable, Vijay Shanker, John W. Ely, Hong Yu 0001 |
AMIA | 7 |
| 2006 | BioEx: A Novel User-Interface that Accesses Images from Abstract Sentences
Hong Yu 0001, Minsuk Lee |
HLT-NAACL | 1 |
| 2006 | Exploring supervised and unsupervised methods to detect topics in biomedical textabstractBACKGROUND: Topic detection is a task that automatically identifies topics (e.g., "biochemistry" and "protein structure") in scientific articles based on information content. Topic detection will benefit many other natural language processing tasks including information retrieval, text summarization and question answering; and is a necessary step towards the building of an information system that provides an efficient way for biologists to seek information from an ocean of literature. RESULTS: We have explored the methods of Topic Spotting, a task of text categorization that applies the supervised machine-learning technique naïve Bayes to assign automatically a document into one or more predefined topics; and Topic Clustering, which apply unsupervised hierarchical clustering algorithms to aggregate documents into clusters such that each cluster represents a topic. We have applied our methods to detect topics of more than fifteen thousand of articles that represent over sixteen thousand entries in the Online Mendelian Inheritance in Man (OMIM) database. We have explored bag of words as the features. Additionally, we have explored semantic features; namely, the Medical Subject Headings (MeSH) that are assigned to the MEDLINE records, and the Unified Medical Language System (UMLS) semantic types that correspond to the MeSH terms, in addition to bag of words, to facilitate the tasks of topic detection. Our results indicate that incorporating the MeSH terms and the UMLS semantic types as additional features enhances the performance of topic detection and the naïve Bayes has the highest accuracy, 66.4%, for predicting the topic of an OMIM article as one of the total twenty-five topics. CONCLUSION: Our results indicate that the supervised topic spotting methods outperformed the unsupervised topic clustering; on the other hand, the unsupervised topic clustering methods have the advantages of being robust and applicable in real world settings. Minsuk Lee, Hong Yu 0001 |
BMC Bioinform. | 3 |
| 2006 | A large scale, corpus-based approach for automatically disambiguating biomedical abbreviationsabstractAbbreviations and acronyms are widely used in the biomedical literature and many of them represent important biomedical concepts. Because many abbreviations are ambiguous (e.g., CAT denotes both chloramphenicol acetyl transferase and computed axial tomography , depending on the context), recognizing the full form associated with each abbreviation is in most cases equivalent to identifying the meaning of the abbreviation. This, in turn, allows us to perform more accurate natural language processing, information extraction, and retrieval. In this study, we have developed supervised approaches to identifying the full forms of ambiguous abbreviations within the context they appear. We first automatically assigned multiple possible full forms for each abbreviation; we then treated the in-context full-form prediction for each specific abbreviation occurrence as a case of word-sense disambiguation. We generated automatically a dictionary of all possible full forms for each abbreviation. We applied supervised machine-learning algorithms for disambiguation. Because some of the links between abbreviations and their corresponding full forms are explicitly given in the text and can be recovered automatically, we can use these explicit links to automatically provide training data for disambiguating the abbreviations that are not linked to a full form within a text. We evaluated our methods on over 150 thousand abstracts and obtain for coverage and precision results of 82% and 92%, respectively, when performed as tenfold cross-validation, and 79% and 80%, respectively, when evaluated against an external set of abstracts in which the abbreviations are not defined. Hong Yu 0001, Won Kim 0003, Vasileios Hatzivassiloglou, W. John Wilbur |
ACM Trans. Inf. Syst. | 1 |
| 2005 | Question Analysis for Biomedical Question Answering
Carl L. Sable, Minsuk Lee, Hai Ran Zhu, Hong Yu 0001 |
AMIA | 4 |
| 2004 | Using MEDLINE as a Knowledge Source for Disambiguating Abbreviations in Full-Text Biomedical Journal ArticlesabstractBiomedical abbreviations and acronyms are widely used in biomedical literature. Since many abbreviations represent important content in biomedical literature, information retrieval and extraction benefits from identifying the meanings of biomedical abbreviations. Since many abbreviations are ambiguous, it would be important to map abbreviations to their full forms, which ultimately represent the meanings of the abbreviations. In this study, we present a novel unsupervised method that applies MEDLINE records as a knowledge source for disambiguating abbreviations in full-text biomedical journal articles. We first automatically generated from MEDLINE records a knowledge source or dictionary of abbreviation-full pairs. We then trained on MEDLINE records and predicted the full forms of abbreviations in full-text journal articles by applying supervised machine-learning algorithms in an unsupervised fashion. We report up to 92% prediction precision and up to 91% coverage. Hong Yu 0001, Won Kim 0003, Vasileios Hatzivassiloglou, W. John Wilbur |
CBMS | 1 |
| 2004 | GeneWays: a system for extracting, analyzing, visualizing, and integrating molecular pathway data
Andrey Rzhetsky, Ivan Iossifov, Tomohiro Koike, Michael Krauthammer, Pauline Kra, Mitzi Morris, Hong Yu 0001, Pablo Ariel Duboue, Wubin Weng, W. John Wilbur |
J. Biomed. Informatics | 7 |
| 2002 | Automatic extraction of gene and protein synonyms from MEDLINE and journal articles
Hong Yu 0001, Vasileios Hatzivassiloglou, Carol Friedman, Andrey Rzhetsky, W. John Wilbur |
AMIA | 1 |
| 2002 | Research Paper: Mapping Abbreviations to Full Forms in Biomedical ArticlesabstractOBJECTIVE: To develop methods that automatically map abbreviations to their full forms in biomedical articles. METHODS: The authors developed two methods of mapping defined and undefined abbreviations (defined abbreviations are paired with their full forms in the articles, whereas undefined ones are not). For defined abbreviations, they developed a set of pattern-matching rules to map an abbreviation to its full form and implemented the rules into a software program, AbbRE (for "abbreviation recognition and extraction"). Using the opinions of domain experts as a reference standard, they evaluated the recall and precision of AbbRE for defined abbreviations in ten biomedical articles randomly selected from the ten most frequently cited medical and biological journals. They also measured the percentage of undefined abbreviations in the same set of articles, and they investigated whether they could map undefined abbreviations to any of four public abbreviation databases (GenBank LocusLink, SWISSPROT, LRABR of the UMLS Specialist Lexicon, and BioABACUS). RESULTS: AbbRE had an average 0.70 recall and 0.95 precision for the defined abbreviations. The authors found that an average of 25 percent of abbreviations were defined in biomedical articles and that of a randomly selected subset of undefined abbreviations, 68 percent could be mapped to any of four abbreviation databases. They also found that many abbreviations are ambiguous (i.e., they map to more than one full form in abbreviation databases). CONCLUSION: AbbRE is efficient for mapping defined abbreviations. To couple AbbRE with abbreviation databases for the mapping of undefined abbreviations, not only exhaustive abbreviation databases but also a method to resolve the ambiguity of abbreviations in the databases are needed. Hong Yu 0001, George Hripcsak, Carol Friedman |
J. Am. Medical Informatics Assoc. | 1 |
| 2002 | Automatically identifying gene/protein terms in MEDLINE abstracts
Hong Yu 0001, Vasileios Hatzivassiloglou, Andrey Rzhetsky, W. John Wilbur |
J. Biomed. Informatics | 1 |
| 2000 | Hereditary Disease Discovery from a Clinical Data Warehouse
Hong Yu 0001, George Hripcsak |
AMIA | 1 |
| 2000 | A Large Scale, Cross-disease Family Health History Data Set
Hong Yu 0001, George Hripcsak |
AMIA | 1 |
| 1999 | Representing genomic knowledge in the UMLS semantic network
Hong Yu 0001, Carol Friedman, Andrey Rzhetsky, Pauline Kra |
AMIA | 1 |