VLDB 2026 Research / reviewers in the wild / expert
Cui Tao
dblp:03/4748
· DBLP profile ↗
112ranked-venue papers
11as first author
37since 2021 · last 2026
0000-0002-4267-1924ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 94 · 5 first-author · 30 since 2021Databases, data management, data science and information retrieval · 13 · 6 first-author · 2 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 5 since 2021Computer networks · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AcuKG: a comprehensive knowledge graph for medical acupunctureabstractBACKGROUND: Acupuncture, a key modality in traditional Chinese medicine, is gaining global recognition as a complementary therapy and a subject of increasing scientific interest. However, fragmented and unstructured acupuncture knowledge spread across diverse sources poses challenges for semantic retrieval, reasoning, and in-depth analysis. To address this gap, we developed AcuKG, a comprehensive knowledge graph that systematically organizes acupuncture-related knowledge to support sharing, discovery, and artificial intelligence-driven innovation in the field. METHODS: AcuKG integrates data from multiple sources, including online resources, guidelines, PubMed literature, ClinicalTrials.gov, and multiple ontologies (SNOMED CT, UBERON, and MeSH). We employed entity recognition, relation extraction, and ontology mapping to establish AcuKG, with human-in-the-loop to ensure data quality. Two cases evaluated AcuKG's usability: (1) how AcuKG advances acupuncture research for obesity and (2) how AcuKG enhances large language model (LLM) application on acupuncture question-answering. RESULTS: AcuKG comprises 1839 entities and 11 527 relations, mapped to 1836 standard concepts in 3 ontologies. Two use cases demonstrated AcuKG's effectiveness and potential in advancing acupuncture research and supporting LLM applications. In the obesity use case, AcuKG identified highly relevant acupoints (eg, ST25, ST36) and uncovered novel research insights based on evidence from clinical trials and literature. When applied to LLMs in answering acupuncture-related questions, integrating AcuKG with GPT-4o and LLaMA 3 significantly improved accuracy (GPT-4o: 46% → 54%, P = .03; LLaMA 3: 17% → 28%, P = .01). CONCLUSION: AcuKG is an open dataset that provides a structured and computational framework for acupuncture applications, bridging traditional practices with acupuncture research and cutting-edge LLM technologies. Xueqing Peng, Su-Yuan Peng, Jianfu Li, Donghong Pei, Fang Li 0011, Yongqun He, Cui Tao, Hua Xu 0001, Na Hong |
J. Am. Medical Informatics Assoc. | 12 |
| 2026 | Optimizing Retrieval-Augmented Generation (RAG) in clinical medicine: methods and performance evaluationabstractOBJECTIVE: Evaluate how RAG architecture, including corpus structure, retrieval strategy, and pipeline complexity, affects LLM-based medical problem solving and knowledge retrieval in sleep medicine. MATERIALS AND METHODS: We benchmarked four open-source LLMs (Llama-3-8B, Llama -3 -70B, Qwen 2.5-14B, and Qwen 2.5-235B) using a knowledge base of five sleep medicine textbooks. We compared performance across three dimensions: corpus structure (raw text vs table-of-contents aligned.), retrieval strategy (dense embedding vs hybrid sparse-dense), and pipeline complexity (baseline vs augmented). Evaluation metrics included board-style multiple choice question (MCQ) accuracy and clinical case vignette diagnostic ranking. RESULTS: RAG improved MCQ accuracy for all models. Llama-8B saw the largest gain of 10.6% (61.8% to 72.4%), while Qwen-235B reached 87.3%. In clinical cases, Llama-8B accuracy dropped by 7.1% when using raw text and dense retrieval due to context noise. This was corrected by using structured hybrid configurations. Hybrid retrieval consistently outperformed dense-only methods. Overall, structured corpora improved primary diagnosis accuracy by 6.1% on average, with Qwen-235B reaching a peak 10.2% increase. DISCUSSION: RAG effectiveness depends on the balance between model size and data structure. Large models handle uncurated text well, but smaller models are easily distracted by irrelevant data. Hybrid retrieval is necessary to maintain precision with specialized medical terms. A structured corpus paired with a baseline hybrid pipeline offers the best stability and speed for clinical use. CONCLUSION: Rigorous data curation and hybrid retrieval are as essential as model scale for deploying safe, guideline-compliant AI in sleep medicine. Pengze Li, Anshum Patel, Sai Krishna Vallamchetla, Hayden Heninger, Het Contractor, Cui Tao, Joseph Cheung |
J. Am. Medical Informatics Assoc. | 6 |
| 2026 | Validation of 13 102 International Classification of Diseases, Tenth Revision, Clinical Modification codes using a large language model-based systemabstractOBJECTIVES: To comprehensively evaluate the validity of International Classification of Diseases, Tenth Revision, Clinical Modification (ICD-10-CM) codes for both prevalent diagnoses and less common diseases, and to assess the performance of a large language model (LLM)-based system in validating these codes. MATERIALS AND METHODS: This retrospective study analyzed hospital admissions from the Medical Information Mart for Intensive Care IV (MIMIC-IV) database. We developed a validated LLM-based system using GPT-4o, refined through iterative prompt engineering, to assess ICD-10-CM code validity. We measured the positive predictive value (PPV) of ICD-10-CM codes, PPV of principal and secondary diagnoses, and the performance of an LLM-based system in code validation. RESULTS: Among 865 079 assigned codes, the PPV was 84.6% (95% CI, 84.5%-84.6%). Principal diagnoses had a PPV of 93.9% (95% CI, 93.7%-94.1%), while secondary diagnoses had a PPV of 83.8% (95% CI, 83.7%-83.9%). The LLM system demonstrated high performance in validating ICD codes, achieving 93.6% accuracy, 95.4% sensitivity, and 85.2% specificity. Among correctly assigned secondary diagnoses, the majority (67.9%) represented historical or baseline conditions, while 32.1% reflected active conditions that deviated from baseline status; 22.3% of these emerged after hospital admission. PPV decreases with later diagnosis positions, with the largest decline occurring between principal and secondary diagnoses. DISCUSSION AND CONCLUSION: In this large-scale evaluation, ICD-10-CM codes exhibited generally high accuracy, though variability existed by position and condition type. A validated LLM system performed comparably to physician review and offers a scalable means to improve coding accuracy. These findings support the potential for integrating LLM-based auditing into routine workflows to strengthen the quality of administrative and research data. Yilin Song, Rex Siu, Induja R. Nimma, Thomas R. Savage, Zhichen Li, Daryl Ramai, Dilhana Badurdeen, Cui Tao, Vivek Kumbhari |
J. Am. Medical Informatics Assoc. | 12 |
| 2026 | Exploring the role of reinforcement learning in vision-language models for cardiovascular disease decision support
Pengze Li, Jianfu Li, Shuteng Niu, Farris K. Timimi, Joseph Cheung, Clark Otley, Sonya Makhni, Fang Li 0011, Jingna Feng, Xinyue Hu 0002, Yue Yu 0012, Cui Tao |
J. Biomed. Informatics | 13 |
| 2026 | UCVA ontology: Standardizing local context factors to support the analysis of unwarranted clinical variation
Apollo McOwiti, Xubing Hao, Rebecca Lin, Laila Rasmy-Bekhet, Cui Tao, Susan H. Fenton |
J. Biomed. Informatics | 5 |
| 2025 | Leveraging Vulnerabilities in Temporal Graph Neural Networks via Strategic High-Impact AssaultsabstractTemporal Graph Neural Networks (TGNNs) have become indispensable for analyzing dynamic graphs in critical applications such as social networks, communication systems, and financial networks. However, the robustness of TGNNs against adversarial attacks, particularly sophisticated attacks that exploit the temporal dimension, remains a significant challenge. Existing attack methods for Spatio-Temporal Dynamic Graphs (STDGs) often rely on simplistic, easily detectable perturbations (e.g., random edge additions/deletions) and fail to strategically target the most influential nodes and edges for maximum impact. We introduce the High Impact Attack (HIA), a novel restricted black-box attack framework specifically designed to overcome these limitations and expose critical vulnerabilities in TGNNs. HIA leverages a data-driven surrogate model to identify structurally important nodes (central to network connectivity) and dynamically important nodes (critical for the graph's temporal evolution). It then employs a hybrid perturbation strategy, combining strategic edge injection (to create misleading connections) and targeted edge deletion (to disrupt essential pathways), maximizing TGNN performance degradation. Importantly, HIA minimizes the number of perturbations to enhance stealth, making it more challenging to detect. Comprehensive experiments on five real-world datasets and four representative TGNN architectures (TGN, JODIE, DySAT, and TGAT) demonstrate that HIA significantly reduces TGNN accuracy on the link prediction task, achieving up to a 35.55% decrease in Mean Reciprocal Rank (MRR) - a substantial improvement over state-of-the-art baselines. These results highlight fundamental vulnerabilities in current STDG models and underscore the urgent need for robust defenses that account for both structural and temporal dynamics. Code and Data are available at https://github.com/ryandhjeon/hia. Donghyun Jeon, Lijing Zhu, Haifang Li 0003, Pengze Li, Jingna Feng, Tiehang Duan, Houbing Song, Cui Tao, Shuteng Niu |
CIKM | 8 |
| 2025 | ETT-CKGE: Efficient Task-Driven Tokens for Continual Knowledge Graph Embedding
Lijing Zhu, Qizhen Lan, Qing Tian 0003, Xi Xiao 0003, Tiehang Duan, Cui Tao, Shuteng Niu |
ECML/PKDD (6) | 10 |
| 2025 | Temporal Ensemble Logic for Integrative Representation of the Entirety of Clinical Trials
Yan Huang 0034, Rashmie Abeysinghe, Zenan Sun, Pengze Li, Xing He 0003, Shiqiang Tao, Cui Tao, Jiang Bian 0001, Licong Cui, Guo-Qiang Zhang 0001 |
TIME | 9 |
| 2025 | Quantitatively assessing the impact of the quality of SNOMED CT subtype hierarchy on cohort queriesabstractOBJECTIVE: SNOMED CT provides a standardized terminology for clinical concepts, allowing cohort queries over heterogeneous clinical data including Electronic Health Records (EHRs). While it is intuitive that missing and inaccurate subtype (or is-a) relations in SNOMED CT reduce the recall and precision of cohort queries, the extent of these impacts has not been formally assessed. This study fills this gap by developing quantitative metrics to measure these impacts and performing statistical analysis on their significance. MATERIAL AND METHODS: We used the Optum de-identified COVID-19 Electronic Health Record dataset. We defined micro-averaged and macro-averaged recall and precision metrics to assess the impact of missing and inaccurate is-a relations on cohort queries. Both practical and simulated analyses were performed. Practical analyses involved 407 missing and 48 inaccurate is-a relations confirmed by domain experts, with statistical testing using Wilcoxon signed-rank tests. Simulated analyses used two random sets of 400 is-a relations to simulate missing and inaccurate is-a relations. RESULTS: Wilcoxon signed-rank tests from both practical and simulated analyses (P-values < .001) showed that missing is-a relations significantly reduced the micro- and macro-averaged recall, and inaccurate is-a relations significantly reduced the micro- and macro-averaged precision. DISCUSSION: The introduced impact metrics can assist SNOMED CT maintainers in prioritizing critical hierarchical defects for quality enhancement. These metrics are generally applicable for assessing the quality impact of a terminology's subtype hierarchy on its cohort query applications. CONCLUSION: Our results indicate a significant impact of missing and inaccurate is-a relations in SNOMED CT on the recall and precision of cohort queries. Our work highlights the importance of high-quality terminology hierarchy for cohort queries over EHR data and provides valuable insights for prioritizing quality improvements of SNOMED CT's hierarchy. Xubing Hao, Yan Huang 0034, Jay Shi, Rashmie Abeysinghe, Cui Tao, Kirk Roberts, Guo-Qiang Zhang 0001, Licong Cui |
J. Am. Medical Informatics Assoc. | 6 |
| 2025 | A comparative study of recent large language models on generating hospital discharge summaries for lung cancer patients
Fang Li 0011, Na Hong, Manqi Li, Kirk Roberts, Licong Cui, Cui Tao, Hua Xu 0001 |
J. Biomed. Informatics | 7 |
| 2025 | Exploring multimodal large language models on transthoracic Echocardiogram (TTE) tasks for cardiovascular decision support
Jianfu Li, Zenan Sun, Evan Yu, Ahmed M. Abdelhameed, Weiguo Cao, Jianping He 0002, Pengze Li, Jingna Feng, Yue Yu 0012, Xinyue Hu 0002, Manqi Li, Yifang Dang, Fang Li 0011, Shahyar M. Gharacholou, Cui Tao |
J. Biomed. Informatics | 18 |
| 2025 | Improving entity recognition using ensembles of deep learning and fine-tuned large language models: A case study on adverse event extraction from VAERS and social media
Deepthi Viswaroopan, William He, Jianfu Li, Xu Zuo, Hua Xu 0001, Cui Tao |
J. Biomed. Informatics | 7 |
| 2024 | Relation extraction using large language models: a case study on acupuncture point locationsabstractOBJECTIVE: In acupuncture therapy, the accurate location of acupoints is essential for its effectiveness. The advanced language understanding capabilities of large language models (LLMs) like Generative Pre-trained Transformers (GPTs) and Llama present a significant opportunity for extracting relations related to acupoint locations from textual knowledge sources. This study aims to explore the performance of LLMs in extracting acupoint-related location relations and assess the impact of fine-tuning on GPT's performance. MATERIALS AND METHODS: We utilized the World Health Organization Standard Acupuncture Point Locations in the Western Pacific Region (WHO Standard) as our corpus, which consists of descriptions of 361 acupoints. Five types of relations ("direction_of", "distance_of", "part_of", "near_acupoint", and "located_near") (n = 3174) between acupoints were annotated. Four models were compared: pre-trained GPT-3.5, fine-tuned GPT-3.5, pre-trained GPT-4, as well as pretrained Llama 3. Performance metrics included micro-average exact match precision, recall, and F1 scores. RESULTS: Our results demonstrate that fine-tuned GPT-3.5 consistently outperformed other models in F1 scores across all relation types. Overall, it achieved the highest micro-average F1 score of 0.92. DISCUSSION: The superior performance of the fine-tuned GPT-3.5 model, as shown by its F1 scores, underscores the importance of domain-specific fine-tuning in enhancing relation extraction capabilities for acupuncture-related tasks. In light of the findings from this study, it offers valuable insights into leveraging LLMs for developing clinical decision support and creating educational modules in acupuncture. CONCLUSION: This study underscores the effectiveness of LLMs like GPT and Llama in extracting relations related to acupoint locations, with implications for accurately modeling acupuncture knowledge and promoting standard implementation in acupuncture training and practice. The findings also contribute to advancing informatics applications in traditional and complementary medicine, showcasing the potential of LLMs in natural language processing. Xueqing Peng, Jianfu Li, Xu Zuo, Su-Yuan Peng, Donghong Pei, Cui Tao, Hua Xu 0001, Na Hong |
J. Am. Medical Informatics Assoc. | 7 |
| 2024 | Ensemble pretrained language models to extract biomedical knowledge from literatureabstractOBJECTIVES: The rapid expansion of biomedical literature necessitates automated techniques to discern relationships between biomedical concepts from extensive free text. Such techniques facilitate the development of detailed knowledge bases and highlight research deficiencies. The LitCoin Natural Language Processing (NLP) challenge, organized by the National Center for Advancing Translational Science, aims to evaluate such potential and provides a manually annotated corpus for methodology development and benchmarking. MATERIALS AND METHODS: For the named entity recognition (NER) task, we utilized ensemble learning to merge predictions from three domain-specific models, namely BioBERT, PubMedBERT, and BioM-ELECTRA, devised a rule-driven detection method for cell line and taxonomy names and annotated 70 more abstracts as additional corpus. We further finetuned the T0pp model, with 11 billion parameters, to boost the performance on relation extraction and leveraged entites' location information (eg, title, background) to enhance novelty prediction performance in relation extraction (RE). RESULTS: Our pioneering NLP system designed for this challenge secured first place in Phase I-NER and second place in Phase II-relation extraction and novelty prediction, outpacing over 200 teams. We tested OpenAI ChatGPT 3.5 and ChatGPT 4 in a Zero-Shot setting using the same test set, revealing that our finetuned model considerably surpasses these broad-spectrum large language models. DISCUSSION AND CONCLUSION: Our outcomes depict a robust NLP system excelling in NER and RE across various biomedical entities, emphasizing that task-specific models remain superior to generic large ones. Such insights are valuable for endeavors like knowledge graph development and hypothesis formulation in biomedical research. Qiang Wei 0002, Liang-Chin Huang, Jianfu Li, Yao-Shun Chuang, Jianping He 0002, Avisha Das, Vipina Kuttichi Keloth, Yuntao Yang, Chiamaka S. Diala, Kirk Roberts, Cui Tao, Xiaoqian Jiang, W. Jim Zheng, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 13 |
| 2024 | RefAI: a GPT-powered retrieval-augmented generative tool for biomedical literature recommendation and summarizationabstractOBJECTIVES: Precise literature recommendation and summarization are crucial for biomedical professionals. While the latest iteration of generative pretrained transformer (GPT) incorporates 2 distinct modes-real-time search and pretrained model utilization-it encounters challenges in dealing with these tasks. Specifically, the real-time search can pinpoint some relevant articles but occasionally provides fabricated papers, whereas the pretrained model excels in generating well-structured summaries but struggles to cite specific sources. In response, this study introduces RefAI, an innovative retrieval-augmented generative tool designed to synergize the strengths of large language models (LLMs) while overcoming their limitations. MATERIALS AND METHODS: RefAI utilized PubMed for systematic literature retrieval, employed a novel multivariable algorithm for article recommendation, and leveraged GPT-4 turbo for summarization. Ten queries under 2 prevalent topics ("cancer immunotherapy and target therapy" and "LLMs in medicine") were chosen as use cases and 3 established counterparts (ChatGPT-4, ScholarAI, and Gemini) as our baselines. The evaluation was conducted by 10 domain experts through standard statistical analyses for performance comparison. RESULTS: The overall performance of RefAI surpassed that of the baselines across 5 evaluated dimensions-relevance and quality for literature recommendation, accuracy, comprehensiveness, and reference integration for summarization, with the majority exhibiting statistically significant improvements (P-values <.05). DISCUSSION: RefAI demonstrated substantial improvements in literature recommendation and summarization over existing tools, addressing issues like fabricated papers, metadata inaccuracies, restricted recommendations, and poor reference integration. CONCLUSION: By augmenting LLM with external resources and a novel ranking algorithm, RefAI is uniquely capable of recommending high-quality literature and generating well-structured summaries, holding the potential to meet the critical needs of biomedical professionals in navigating and synthesizing vast amounts of scientific literature. Jeff Zhao, Manqi Li, Yifang Dang, Evan Yu, Jianfu Li, Zenan Sun, Usama Hussein, Jianguo Wen, Ahmed M. Abdelhameed, Junhua Mai, Shenduo Li, Yue Yu 0012, Xinyue Hu 0002, Daowei Yang, Jingna Feng, Zehan Li, Jianping He 0002, Tiehang Duan, Yanyan Lou, Fang Li 0011, Cui Tao |
J. Am. Medical Informatics Assoc. | 23 |
| 2024 | FedFSA: Hybrid and federated framework for functional status ascertainment across institutions
Sunyang Fu, Heling Jia, Maria Vassilaki, Vipina Kuttichi Keloth, Yifang Dang, Yujia Zhou 0003, Muskan Garg, Ronald C. Petersen, Jennifer L. St. Sauver, Sungrim Moon, Liwei Wang 0010, Andrew Wen, Fang Li 0011, Hua Xu 0001, Cui Tao, Jungwei Fan 0001, Sunghwan Sohn |
J. Biomed. Informatics | 15 |
| 2024 | Artificial intelligence-powered pharmacovigilance: A review of machine and deep learning in clinical text-based adverse drug event detection for benchmark datasets
Zehan Li, Zenan Sun, Fang Li 0011, Susan H. Fenton, Hua Xu 0001, Cui Tao |
J. Biomed. Informatics | 8 |
| 2024 | Improving tabular data extraction in scanned laboratory reports using deep learning models
Qiang Wei 0002, Xinghan Chen, Jianfu Li, Cui Tao, Hua Xu 0001 |
J. Biomed. Informatics | 5 |
| 2024 | Online continual decoding of streaming EEG signal with a balanced and informative memory buffer
Tiehang Duan, Zhenyi Wang 0001, Fang Li 0011, Gianfranco Doretto, Donald A. Adjeroh, Yiyi Yin, Cui Tao |
Neural Networks | 7 |
| 2023 | Replay with Stochastic Neural Transformation for Online Continual EEG ClassificationabstractBrain computer interface (BCI) systems used for clinical assistance purposes such as wheelchair control require decoding of streaming brain signals i.e. electroencephalography (EEG) signals over a long period of time with subject shift in the middle. Numerous challenges arise during this online continual brain signal decoding process: 1) the EEG decoder needs to deal with streaming EEG signals from sequentially arriving subjects, with no data available beforehand for large-scale pretraining; 2) the EEG decoder should avoid catastrophic forgetting on previous subjects after learning on a new subject; 3) the EEG decoder should perform well on noisy signals with high variance across subjects. We proposed a principled replay-based approach for this general decoding scenario, forming a bi-level optimization framework with stochastic neural transformation for dynamic memory evolution, making them representative in feature space and encouraging the model to generalize well. The evolved signal segments are stored and replayed during later decoding stages to achieve optimal model performance on all previous subjects. The stochastic neural transformation performed in inner sup of bi-level optimization significantly enhances the diversity of stored signal segments and improves model robustness during online continual decoding. We perform detailed theoretical analysis on model’s generalization ability in addition to the empirical evaluations. We construct multiple new benchmarks to mimic real-world online sequential EEG decoding scenarios with underlying subject shifts. The extensive evaluation of the proposed approach shows it outperforms related strong baselines by a large margin. Tiehang Duan, Zhenyi Wang 0001, Gianfranco Doretto, Fang Li 0011, Cui Tao, Donald A. Adjeroh |
BIBM | 5 |
| 2023 | Distributionally Robust Cross Subject EEG DecodingabstractRecently, deep learning has shown to be effective for Electroencephalography (EEG) decoding tasks. Yet, its performance can be negatively influenced by two key factors: 1) the high variance and different types of corruption that are inherent in the signal, 2) the EEG datasets are usually relatively small given the acquisition cost, annotation cost and amount of effort needed. Data augmentation approaches for alleviation of this problem have been empirically studied, with augmentation operations on spatial domain, time domain or frequency domain handcrafted based on expertise of domain knowledge. In this work, we propose a principled approach to perform dynamic evolution on the data for improvement of decoding robustness. The approach is based on distributionally robust optimization and achieves robustness by optimizing on a family of evolved data distributions instead of the single training data distribution. We derived a general data evolution framework based on Wasserstein gradient flow (WGF) and provides two different forms of evolution within the framework. Intuitively, the evolution process helps the EEG decoder to learn more robust and diverse features. It is worth mentioning that the proposed approach can be readily integrated with other data augmentation approaches for further improvements. We performed extensive experiments on the proposed approach and tested its performance on different types of corrupted EEG signals. The model significantly outperforms competitive baselines on challenging decoding scenarios. Tiehang Duan, Zhenyi Wang 0001, Gianfranco Doretto, Fang Li 0011, Cui Tao, Donald A. Adjeroh |
ECAI | 5 |
| 2023 | Honoring Heritage, Managing Health: A Mobile Diabetes Self-Management App for Native Americans with Cultural Sensitivity and Local FactorsabstractDiabetes has a disproportionate impact on Native Americans (NAs) as a chronic health condition, yet there is a dearth of mobile apps specifically designed for this population. In this paper, we present the design and development of a culturally tailored mobile app for NAs, taking into account their cultural traditions. Our app incorporates NA's traditional foods, food availability, the importance of family and community, cultural practices and beliefs, local resources, and heritage heroes into the app interface and self-management design. The app includes personalized nutrition guidance, family and community-based support, seamless connection to tribal health providers, access to local resources, and integration of cultural elements. By considering the cultural context of NAs, the developed app has the potential to provide culturally sensitive and relevant features that address the unique needs and preferences of NA users, facilitating effective self-management of diabetes. Wordh Ul Hasan, Juan Li 0004, Shadi Alian, Vikram Pandey, Kimia Tuz Zaman, Cui Tao |
ISCC | 8 |
| 2023 | Empowering Caregivers of Alzheimer's Disease and Related Dementias (ADRD) with a GPT-Powered Voice Assistant: Leveraging Peer Insights from Social MediaabstractCaring for individuals with Alzheimer's Disease and Related Dementias (ADRD) is a complex and challenging task, especially for unprofessional caregivers who often lack the necessary training and resources. While online peer support groups have been shown to be useful in providing caregivers with information and emotional support, many caregivers are unable to benefit from them due to time constraints and limited knowledge of social media platforms. To address this issue, we propose the development of a voice assistant app that can collect relevant information and discussions from online peer support groups on social media. This app will use the collected information as a knowledge base and fine-tune a Generative Pre-trained Transformers (GPT) model to facilitate caregivers in accessing shared experiences and practical tips from peers. Initial evaluation of the app has shown promising results in terms of feasibility and potential impact on caregivers. Kimia Tuz Zaman, Wordh Ul Hasan, Juan Li 0004, Cui Tao |
ISCC | 4 |
| 2023 | Machine learning-based donor permission extraction from informed consent documentsabstractBACKGROUND: With more clinical trials are offering optional participation in the collection of bio-specimens for biobanking comes the increasing complexity of requirements of informed consent forms. The aim of this study is to develop an automatic natural language processing (NLP) tool to annotate informed consent documents to promote biorepository data regulation, sharing, and decision support. We collected informed consent documents from several publicly available sources, then manually annotated them, covering sentences containing permission information about the sharing of either bio-specimens or donor data, or conducting genetic research or future research using bio-specimens or donor data. RESULTS: We evaluated a variety of machine learning algorithms including random forest (RF) and support vector machine (SVM) for the automatic identification of these sentences. 120 informed consent documents containing 29,204 sentences were annotated, of which 1250 sentences (4.28%) provide answers to a permission question. A support vector machine (SVM) model achieved a F-1 score of 0.95 on classifying the sentences when using a gold standard, which is a prefiltered corpus containing all relevant sentences. CONCLUSIONS: This study provides the feasibility of using machine learning tools to classify permission-related sentences in informed consent documents. Madhuri Sankaranarayanapillai, Jingcheng Du, Yang Xiang 0003, Frank J. Manion, Marcelline R. Harris, Cooper Stansbury, Huy Anh Pham, Cui Tao |
BMC Bioinform. | 9 |
| 2023 | Systematic design and data-driven evaluation of social determinants of health ontology (SDoHO)abstractOBJECTIVE: Social determinants of health (SDoH) play critical roles in health outcomes and well-being. Understanding the interplay of SDoH and health outcomes is critical to reducing healthcare inequalities and transforming a "sick care" system into a "health-promoting" system. To address the SDOH terminology gap and better embed relevant elements in advanced biomedical informatics, we propose an SDoH ontology (SDoHO), which represents fundamental SDoH factors and their relationships in a standardized and measurable way. MATERIAL AND METHODS: Drawing on the content of existing ontologies relevant to certain aspects of SDoH, we used a top-down approach to formally model classes, relationships, and constraints based on multiple SDoH-related resources. Expert review and coverage evaluation, using a bottom-up approach employing clinical notes data and a national survey, were performed. RESULTS: We constructed the SDoHO with 708 classes, 106 object properties, and 20 data properties, with 1,561 logical axioms and 976 declaration axioms in the current version. Three experts achieved 0.967 agreement in the semantic evaluation of the ontology. A comparison between the coverage of the ontology and SDOH concepts in 2 sets of clinical notes and a national survey instrument also showed satisfactory results. DISCUSSION: SDoHO could potentially play an essential role in providing a foundation for a comprehensive understanding of the associations between SDoH and health outcomes and paving the way for health equity across populations. CONCLUSION: SDoHO has well-designed hierarchies, practical objective properties, and versatile functionalities, and the comprehensive semantic and coverage evaluation achieved promising performance compared to the existing ontologies relevant to SDoH. Yifang Dang, Fang Li 0011, Xinyue Hu 0002, Vipina Kuttichi Keloth, Sunyang Fu, Muhammad Amith, J. Wilfred Fan, Jingcheng Du, Evan Yu, Xiaoqian Jiang, Hua Xu 0001, Cui Tao |
J. Am. Medical Informatics Assoc. | 14 |
| 2023 | An NLP approach to identify SDoH-related circumstance and suicide crisis from death investigation narrativesabstractOBJECTIVES: Suicide presents a major public health challenge worldwide, affecting people across the lifespan. While previous studies revealed strong associations between Social Determinants of Health (SDoH) and suicide deaths, existing evidence is limited by the reliance on structured data. To resolve this, we aim to adapt a suicide-specific SDoH ontology (Suicide-SDoHO) and use natural language processing (NLP) to effectively identify individual-level SDoH-related social risks from death investigation narratives. MATERIALS AND METHODS: We used the latest National Violent Death Report System (NVDRS), which contains 267 804 victim suicide data from 2003 to 2019. After adapting the Suicide-SDoHO, we developed a transformer-based model to identify SDoH-related circumstances and crises in death investigation narratives. We applied our model retrospectively to annotate narratives whose crisis variables were not coded in NVDRS. The crisis rates were calculated as the percentage of the group's total suicide population with the crisis present. RESULTS: The Suicide-SDoHO contains 57 fine-grained circumstances in a hierarchical structure. Our classifier achieves AUCs of 0.966 and 0.942 for classifying circumstances and crises, respectively. Through the crisis trend analysis, we observed that not everyone is equally affected by SDoH-related social risks. For the economic stability crisis, our result showed a significant increase in crisis rate in 2007-2009, parallel with the Great Recession. CONCLUSIONS: This is the first study curating a Suicide-SDoHO using death investigation narratives. We showcased that our model can effectively classify SDoH-related social risks through NLP approaches. We hope our study will facilitate the understanding of suicide crises and inform effective prevention strategies. Song Wang 0026, Yifang Dang, Zhaoyi Sun, Ying Ding 0001, Jyotishman Pathak, Cui Tao, Yunyu Xiao, Yifan Peng 0002 |
J. Am. Medical Informatics Assoc. | 6 |
| 2022 | Ontologies in the Behavioral Sciences: Accelerating Research and the Accessibility and Use of Knowledge
Mark A. Musen, Jiang Bian 0001, Bruce Chorpita, Vimla L. Patel, Cui Tao |
AMIA | 5 |
| 2022 | Combining Transfer Learning with Graph Attention Models for HIV Risk Prediction
Evan Yu, Cui Tao, Jingcheng Du, Degui Zhi, Yang Xiang 0003, Kayo Fujimoto, John A. Schneider |
AMIA | 2 |
| 2022 | Toward a standard formal semantic representation of the model card reportabstractBACKGROUND: Model card reports aim to provide informative and transparent description of machine learning models to stakeholders. This report document is of interest to the National Institutes of Health's Bridge2AI initiative to address the FAIR challenges with artificial intelligence-based machine learning models for biomedical research. We present our early undertaking in developing an ontology for capturing the conceptual-level information embedded in model card reports. RESULTS: Sourcing from existing ontologies and developing the core framework, we generated the Model Card Report Ontology. Our development efforts yielded an OWL2-based artifact that represents and formalizes model card report information. The current release of this ontology utilizes standard concepts and properties from OBO Foundry ontologies. Also, the software reasoner indicated no logical inconsistencies with the ontology. With sample model cards of machine learning models for bioinformatics research (HIV social networks and adverse outcome prediction for stent implantation), we showed the coverage and usefulness of our model in transforming static model card reports to a computable format for machine-based processing. CONCLUSIONS: The benefit of our work is that it utilizes expansive and standard terminologies and scientific rigor promoted by biomedical ontologists, as well as, generating an avenue to make model cards machine-readable using semantic web technology. Our future goal is to assess the veracity of our model and later expand the model to include additional concepts to address terminological gaps. We discuss tools and software that will utilize our ontology for potential application services. Muhammad Amith, Licong Cui, Degui Zhi, Kirk Roberts, Xiaoqian Jiang, Fang Li 0011, Evan Yu, Cui Tao |
BMC Bioinform. | 8 |
| 2022 | Mining on Alzheimer's diseases related knowledge graph to identity potential AD-related semantic triples for drug repurposingabstractBACKGROUND: To date, there are no effective treatments for most neurodegenerative diseases. Knowledge graphs can provide comprehensive and semantic representation for heterogeneous data, and have been successfully leveraged in many biomedical applications including drug repurposing. Our objective is to construct a knowledge graph from literature to study the relations between Alzheimer's disease (AD) and chemicals, drugs and dietary supplements in order to identify opportunities to prevent or delay neurodegenerative progression. We collected biomedical annotations and extracted their relations using SemRep via SemMedDB. We used both a BERT-based classifier and rule-based methods during data preprocessing to exclude noise while preserving most AD-related semantic triples. The 1,672,110 filtered triples were used to train with knowledge graph completion algorithms (i.e., TransE, DistMult, and ComplEx) to predict candidates that might be helpful for AD treatment or prevention. RESULTS: Among three knowledge graph completion models, TransE outperformed the other two (MR = 10.53, Hits@1 = 0.28). We leveraged the time-slicing technique to further evaluate the prediction results. We found supporting evidence for most highly ranked candidates predicted by our model which indicates that our approach can inform reliable new knowledge. CONCLUSION: This paper shows that our graph mining model can predict reliable new relationships between AD and other entities (i.e., dietary supplements, chemicals, and drugs). The knowledge graph constructed can facilitate data-driven knowledge discoveries and the generation of novel hypotheses. Yi Nian, Xinyue Hu 0002, Rui Zhang 0028, Jingna Feng, Jingcheng Du, Fang Li 0011, Larry Bu, Yuji Zhang 0001, Yong Chen 0016, Cui Tao |
BMC Bioinform. | 10 |
| 2021 | Expressing and Executing Informed Consent Permissions using SWRL: The All of Us Use Case
Muhammad Amith, Marcelline R. Harris, Cooper Stansbury, Kathleen Ford, Frank J. Manion, Cui Tao |
AMIA | 6 |
| 2021 | FHIRTime: Standardizing Temporal Patterns Identified from Clinical Narratives Using HL7 FHIR
Daniel J. Stone, Sijia Liu 0002, Yuan Luo 0001, Andrew Wen, Nansu Zong, Luke V. Rasmussen, Prakash Adekkanattu, Pascal S. Brandt, Jennifer A. Pacheco, Fei Wang 0001, Cui Tao, Jyotishman Pathak, Guoqian Jiang |
AMIA | 11 |
| 2021 | Data and Model Biases in Social Media Analyses: A Case Study of COVID-19 Tweets
Pengfei Yin, Yongqiu Li, Xing He 0003, Jingcheng Du, Cui Tao, Yi Guo 0005, Mattia Prosperi, Pierangelo Veltri, Xi Yang 0015, Yonghui Wu 0001, Jiang Bian 0001 |
AMIA | 6 |
| 2021 | Developing Ontologies to Standardize Descriptions of Visual and Dermoscopic ElementsabstractWith the growing importance of dermoscopic analysis to diagnose skin diseases, the vocabulary of dermoscopy has rapidly expanded without standardized control. Many metaphoric terms are ambiguous and create barriers for general research and education. Ontologies are computable artifacts that represent and model information from a domain space that can later be leveraged by software to understand domain knowledge. We aim to standardize dermoscopic vocabulary through the introduction of two ontologies: 1. the Elements of Visuals Ontology (EVO), a foundational ontology to decompose the visual elements of physical entities; and 2. the Dermoscopy Elements of Visuals Ontology (DEVO), a domain ontology that harnesses EVO to formalize the definitions of dermoscopic metaphoric terms. We discuss how DEVO would enhance both trainee education and patient care, with the future goal of generating responses to queries about dermoscopic features and integrating these features with diagnostic rules for skin diseases. Rebecca Lin, Muhammad Amith, Xinyuan Zhang 0003, Cynthia Wang, Jeremy Light, John Strickley, Cui Tao |
BIBM | 7 |
| 2021 | Dental EHR-infused Persona Ontologies to Enrich Dental Dialogue Interaction of AgentsabstractThe quality of patient-provider communication can predict the healthcare outcomes in patients, and therefore, training dental providers to handle the communication effort with patients is crucial. In our previous work, we developed an ontology model that can standardize and represent patient-provider communication, which can later be integrated in conversational agents as tools for dental communication training. In this study, we embark on enriching our previous model with an ontology of patient personas to portray and express types of dental patient archetypes. The Ontology of Patient Personas that we developed was rooted in terminologies from an OBO Foundry ontology and dental electronic health record data elements. We discuss how this ontology aims to enhance the aforementioned dialogue ontology and future direction in executing our model in software agents to train dental students. Patricia Ngantcha, Muhammad Amith, Kirk Roberts, John A. Valenza, Muhammad F. Walji, Cui Tao |
BIBM | 6 |
| 2021 | Extracting postmarketing adverse events from safety reports in the vaccine adverse event reporting system (VAERS) using deep learningabstractOBJECTIVE: Automated analysis of vaccine postmarketing surveillance narrative reports is important to understand the progression of rare but severe vaccine adverse events (AEs). This study implemented and evaluated state-of-the-art deep learning algorithms for named entity recognition to extract nervous system disorder-related events from vaccine safety reports. MATERIALS AND METHODS: We collected Guillain-Barré syndrome (GBS) related influenza vaccine safety reports from the Vaccine Adverse Event Reporting System (VAERS) from 1990 to 2016. VAERS reports were selected and manually annotated with major entities related to nervous system disorders, including, investigation, nervous_AE, other_AE, procedure, social_circumstance, and temporal_expression. A variety of conventional machine learning and deep learning algorithms were then evaluated for the extraction of the above entities. We further pretrained domain-specific BERT (Bidirectional Encoder Representations from Transformers) using VAERS reports (VAERS BERT) and compared its performance with existing models. RESULTS AND CONCLUSIONS: Ninety-one VAERS reports were annotated, resulting in 2512 entities. The corpus was made publicly available to promote community efforts on vaccine AEs identification. Deep learning-based methods (eg, bi-long short-term memory and BERT models) outperformed conventional machine learning-based methods (ie, conditional random fields with extensive features). The BioBERT large model achieved the highest exact match F-1 scores on nervous_AE, procedure, social_circumstance, and temporal_expression; while VAERS BERT large models achieved the highest exact match F-1 scores on investigation and other_AE. An ensemble of these 2 models achieved the highest exact match microaveraged F-1 score at 0.6802 and the second highest lenient match microaveraged F-1 score at 0.8078 among peer models. Jingcheng Du, Yang Xiang 0003, Madhuri Sankaranarayanapillai, Yuqi Si, Huy Anh Pham, Hua Xu 0001, Yong Chen 0016, Cui Tao |
J. Am. Medical Informatics Assoc. | 10 |
| 2021 | COVID-19 trial graph: a linked graph for COVID-19 clinical trialsabstractOBJECTIVE: Clinical trials are an essential part of the effort to find safe and effective prevention and treatment for COVID-19. Given the rapid growth of COVID-19 clinical trials, there is an urgent need for a better clinical trial information retrieval tool that supports searching by specifying criteria, including both eligibility criteria and structured trial information. MATERIALS AND METHODS: We built a linked graph for registered COVID-19 clinical trials: the COVID-19 Trial Graph, to facilitate retrieval of clinical trials. Natural language processing tools were leveraged to extract and normalize the clinical trial information from both their eligibility criteria free texts and structured information from ClinicalTrials.gov. We linked the extracted data using the COVID-19 Trial Graph and imported it to a graph database, which supports both querying and visualization. We evaluated trial graph using case queries and graph embedding. RESULTS: The graph currently (as of October 5, 2020) contains 3392 registered COVID-19 clinical trials, with 17 480 nodes and 65 236 relationships. Manual evaluation of case queries found high precision and recall scores on retrieving relevant clinical trials searching from both eligibility criteria and trial-structured information. We observed clustering in clinical trials via graph embedding, which also showed superiority over the baseline (0.870 vs 0.820) in evaluating whether a trial can complete its recruitment successfully. CONCLUSIONS: The COVID-19 Trial Graph is a novel representation of clinical trials that allows diverse search queries and provides a graph-based visualization of COVID-19 clinical trials. High-dimensional vectors mapped by graph embedding for clinical trials would be potentially beneficial for many downstream applications, such as trial end recruitment status prediction and trial similarity comparison. Our methodology also is generalizable to other clinical trials. Jingcheng Du, Prerana Ramesh, Yang Xiang 0003, Xiaoqian Jiang, Cui Tao |
J. Am. Medical Informatics Assoc. | 7 |
| 2020 | Leverage Real-World Longitudinal Data in Large Clinical Research Networks for Alzheimer's Disease and Related Dementia (ADRD)
Rui Duan 0004, Zhaoyi Chen, Jiayi Tong, Chongliang Luo, Tianchen Lyu, Cui Tao, Demetrius Maraganore, Jiang Bian 0001, Yong Chen 0016 |
AMIA | 6 |
| 2020 | Identifying Clinical Risk Factors for Opioid Use Disorder using a Distributed Algorithm to Combine Real-World Data from a Large Clinical Data Research Network
Jiayi Tong, Zhaoyi Chen, Rui Duan 0004, Wei-Hsuan Lo-Ciganic, Tianchen Lyu, Cui Tao, Peter A. Merkel, Henry R. Kranzler, Jiang Bian 0001, Yong Chen 0016 |
AMIA | 6 |
| 2020 | A health consumer ontology of fast food informationabstractA variety of severe health issues can be attributed to poor nutrition and poor eating behaviors. Research has explored the impact of nutritional knowledge on an individual's inclination to purchase and consume certain foods. This paper introduces the Ontology of Fast Food Facts, a knowledge base that models consumer nutritional data from major fast food establishments. This artifact serves as an aggregate knowledge base to centralize nutritional information for consumers. As a semantically-linked data source, the Ontology of Fast Food Facts could engender methods and tools to further the research and impact the health consumers' diet and behavior, which is a factor in many severe health outcomes. We describe the initial development of this ontology and future directions we plan with this knowledge base. Muhammad Amith, Grace Xiong, Kirk Roberts, Cui Tao |
BIBM | 5 |
| 2020 | Development and Evaluation of ADCareOnto - an Ontology for Personalized Home Care for Persons with Alzheimer's DiseaseabstractAlzheimer's disease (AD) poses serious challenges for both patients and their family caregivers. In this paper we present the design, development, and evaluation of an ontology model, ADCareOnto, to assist family caregivers providing personalized care for persons living with AD. ADCareOnto includes top-level categories, concepts, and relations about informal care for persons with AD. To enable personalization in care, ADCareOnto also includes a comprehensive user profile modeling that includes various characteristics of both AD patients and caregivers. AD care thus can be tailored based on the user's unique concerns, preferences, and needs. We verified and validated the design of ADCareOnto and evaluated it using a real use case. The results support the quality of its content and techniques. Juan Li 0004, Rasha Hendawi, Vikram Pandey, Rafa Alenezi, Bo Xie 0001, Cui Tao |
HealthCom | 7 |
| 2020 | Time event ontology (TEO): to support semantic representation and reasoning of complex temporal relations of clinical eventsabstractOBJECTIVE: The goal of this study is to develop a robust Time Event Ontology (TEO), which can formally represent and reason both structured and unstructured temporal information. MATERIALS AND METHODS: Using our previous Clinical Narrative Temporal Relation Ontology 1.0 and 2.0 as a starting point, we redesigned concept primitives (clinical events and temporal expressions) and enriched temporal relations. Specifically, 2 sets of temporal relations (Allen's interval algebra and a novel suite of basic time relations) were used to specify qualitative temporal order relations, and a Temporal Relation Statement was designed to formalize quantitative temporal relations. Moreover, a variety of data properties were defined to represent diversified temporal expressions in clinical narratives. RESULTS: TEO has a rich set of classes and properties (object, data, and annotation). When evaluated with real electronic health record data from the Mayo Clinic, it could faithfully represent more than 95% of the temporal expressions. Its reasoning ability was further demonstrated on a sample drug adverse event report annotated with respect to TEO. The results showed that our Java-based TEO reasoner could answer a set of frequently asked time-related queries, demonstrating that TEO has a strong capability of reasoning complex temporal relations. CONCLUSION: TEO can support flexible temporal relation representation and reasoning. Our next step will be to apply TEO to the natural language processing field to facilitate automated temporal information annotation, extraction, and timeline reasoning to better support time-based clinical decision-making. Fang Li 0011, Jingcheng Du, Yongqun He, Hsing-yi Song, Mohcine Madkour, Guozheng Rao, Yang Xiang 0003, Henry W. Chen, Sijia Liu 0002, Liwei Wang 0010, Hua Xu 0001, Cui Tao |
J. Am. Medical Informatics Assoc. | 14 |
| 2020 | Representation of EHR data for predictive modeling: a comparison between UMLS and other terminologiesabstractOBJECTIVE: Predictive disease modeling using electronic health record data is a growing field. Although clinical data in their raw form can be used directly for predictive modeling, it is a common practice to map data to standard terminologies to facilitate data aggregation and reuse. There is, however, a lack of systematic investigation of how different representations could affect the performance of predictive models, especially in the context of machine learning and deep learning. MATERIALS AND METHODS: We projected the input diagnoses data in the Cerner HealthFacts database to Unified Medical Language System (UMLS) and 5 other terminologies, including CCS, CCSR, ICD-9, ICD-10, and PheWAS, and evaluated the prediction performances of these terminologies on 2 different tasks: the risk prediction of heart failure in diabetes patients and the risk prediction of pancreatic cancer. Two popular models were evaluated: logistic regression and a recurrent neural network. RESULTS: For logistic regression, using UMLS delivered the optimal area under the receiver operating characteristics (AUROC) results in both dengue hemorrhagic fever (81.15%) and pancreatic cancer (80.53%) tasks. For recurrent neural network, UMLS worked best for pancreatic cancer prediction (AUROC 82.24%), second only (AUROC 85.55%) to PheWAS (AUROC 85.87%) for dengue hemorrhagic fever prediction. DISCUSSION/CONCLUSION: In our experiments, terminologies with larger vocabularies and finer-grained representations were associated with better prediction performances. In particular, UMLS is consistently 1 of the best-performing ones. We believe that our work may help to inform better designs of predictive models, although further investigation is warranted. Laila Rasmy, Firat Tiryaki, Yujia Zhou 0003, Yang Xiang 0003, Cui Tao, Hua Xu 0001, Degui Zhi |
J. Am. Medical Informatics Assoc. | 5 |
| 2020 | iDISK: the integrated DIetary Supplements Knowledge baseabstractOBJECTIVE: To build a knowledge base of dietary supplement (DS) information, called the integrated DIetary Supplement Knowledge base (iDISK), which integrates and standardizes DS-related information from 4 existing resources. MATERIALS AND METHODS: iDISK was built through an iterative process comprising 3 phases: 1) establishment of the content scope, 2) development of the data model, and 3) integration of existing resources. Four well-regarded DS resources were integrated into iDISK: The Natural Medicines Comprehensive Database, the "About Herbs" page on the Memorial Sloan Kettering Cancer Center website, the Dietary Supplement Label Database, and the Natural Health Products Database. We evaluated the iDISK build process by manually checking that the data elements associated with 50 randomly selected ingredients were correctly extracted and integrated from their respective sources. RESULTS: iDISK encompasses a terminology of 4208 DS ingredient concepts, which are linked via 6 relationship types to 495 drugs, 776 diseases, 985 symptoms, 605 therapeutic classes, 17 system organ classes, and 137 568 DS products. iDISK also contains 7 concept attribute types and 3 relationship attribute types. Evaluation of the data extraction and integration process showed average errors of 0.3%, 2.6%, and 0.4% for concepts, relationships and attributes, respectively. CONCLUSION: We developed iDISK, a publicly available standardized DS knowledge base that can facilitate more efficient and meaningful dissemination of DS knowledge. Rubina F. Rizvi, Jake Vasilakes, Terrence Adam, Genevieve B. Melton, Jeffrey R. Bishop, Jiang Bian 0001, Cui Tao, Rui Zhang 0028 |
J. Am. Medical Informatics Assoc. | 7 |
| 2020 | A study of deep learning approaches for medication and adverse drug event extraction from clinical textabstractOBJECTIVE: This article presents our approaches to extraction of medications and associated adverse drug events (ADEs) from clinical documents, which is the second track of the 2018 National NLP Clinical Challenges (n2c2) shared task. MATERIALS AND METHODS: The clinical corpus used in this study was from the MIMIC-III database and the organizers annotated 303 documents for training and 202 for testing. Our system consists of 2 components: a named entity recognition (NER) and a relation classification (RC) component. For each component, we implemented deep learning-based approaches (eg, BI-LSTM-CRF) and compared them with traditional machine learning approaches, namely, conditional random fields for NER and support vector machines for RC, respectively. In addition, we developed a deep learning-based joint model that recognizes ADEs and their relations to medications in 1 step using a sequence labeling approach. To further improve the performance, we also investigated different ensemble approaches to generating optimal performance by combining outputs from multiple approaches. RESULTS: Our best-performing systems achieved F1 scores of 93.45% for NER, 96.30% for RC, and 89.05% for end-to-end evaluation, which ranked #2, #1, and #1 among all participants, respectively. Additional evaluations show that the deep learning-based approaches did outperform traditional machine learning algorithms in both NER and RC. The joint model that simultaneously recognizes ADEs and their relations to medications also achieved the best performance on RC, indicating its promise for relation extraction. CONCLUSION: In this study, we developed deep learning approaches for extracting medications and their attributes such as ADEs, and demonstrated its superior performance compared with traditional machine learning algorithms, indicating its uses in broader NER and RC tasks in the medical domain. Qiang Wei 0002, Zongcheng Ji, Zhiheng Li 0004, Jingcheng Du, Jun Xu 0007, Yang Xiang 0003, Firat Tiryaki, Stephen Wu 0004, Yaoyun Zhang, Cui Tao, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 11 |
| 2020 | Mining Twitter to assess the determinants of health behavior toward human papillomavirus vaccination in the United StatesabstractOBJECTIVES: The study sought to test the feasibility of using Twitter data to assess determinants of consumers' health behavior toward human papillomavirus (HPV) vaccination informed by the Integrated Behavior Model (IBM). MATERIALS AND METHODS: We used 3 Twitter datasets spanning from 2014 to 2018. We preprocessed and geocoded the tweets, and then built a rule-based model that classified each tweet into either promotional information or consumers' discussions. We applied topic modeling to discover major themes and subsequently explored the associations between the topics learned from consumers' discussions and the responses of HPV-related questions in the Health Information National Trends Survey (HINTS). RESULTS: We collected 2 846 495 tweets and analyzed 335 681 geocoded tweets. Through topic modeling, we identified 122 high-quality topics. The most discussed consumer topic is "cervical cancer screening"; while in promotional tweets, the most popular topic is to increase awareness of "HPV causes cancer." A total of 87 of the 122 topics are correlated between promotional information and consumers' discussions. Guided by IBM, we examined the alignment between our Twitter findings and the results obtained from HINTS. Thirty-five topics can be mapped to HINTS questions by keywords, 112 topics can be mapped to IBM constructs, and 45 topics have statistically significant correlations with HINTS responses in terms of geographic distributions. CONCLUSIONS: Mining Twitter to assess consumers' health behaviors can not only obtain results comparable to surveys, but also yield additional insights via a theory-driven approach. Limitations exist; nevertheless, these encouraging results impel us to develop innovative ways of leveraging social media in the changing health communication landscape. Hansi Zhang, Christopher Wheldon, Adam G. Dunn, Cui Tao, Jinhai Huo, Rui Zhang 0028, Mattia Prosperi, Yi Guo 0005, Jiang Bian 0001 |
J. Am. Medical Informatics Assoc. | 4 |
| 2019 | A Scoping Review to Identify Biobank Classification and Permissions to Share Metadata
Marcelline R. Harris, Frank J. Manion, Madhuri Sankaranarayanapillai, Hsing-yi Song, Marisa Conte, Cui Tao |
AMIA | 6 |
| 2019 | Refactoring and Expanding the Informed Consent Ontology (ICO)
Jonathan Vajda, J. Neil Otte, Cooper Stansbury, Marcelline R. Harris, Frank J. Manion, Cui Tao |
AMIA | 6 |
| 2019 | Expanding the Representation of Permissions and Deontic Roles in the Informed Consent Ontology (ICO)
Jonathan Vajda, J. Neil Otte, Cooper Stansbury, Marcelline R. Harris, Elizabeth Umberfield, Frank J. Manion, Cui Tao |
AMIA | 7 |
| 2019 | Relation Extraction from Clinical Narratives Using Pre-trained Language Models
Qiang Wei 0002, Zongcheng Ji, Yuqi Si, Jingcheng Du, Firat Tiryaki, Stephen Wu 0004, Cui Tao, Kirk Roberts, Hua Xu 0001 |
AMIA | 8 |
| 2019 | Ontology of Consumer Health Vocabulary: providing a formal and interoperable semantic resource for linking lay language and medical terminologyabstractThe Consumer Health Vocabulary has been an important contribution to the health informatics field since its introduction in 2006. Many studies have utilized the vocabulary for various scientific research to bridge the gap between consumers and health experts. Given the flat file format of the Consumer Health Vocabulary dataset, we developed a SKOS-based ontology of the dataset. As an ontology, this dataset can be semantically linked to other resources to provide consumer-level meaning. In addition with this artifact, we plan to further expand the terminology. Muhammad Amith, Licong Cui, Kirk Roberts, Hua Xu 0001, Cui Tao |
BIBM | 5 |
| 2019 | Alexa, What Should I Eat? : A Personalized Virtual Nutrition Coach for Native American Diabetes Patients Using Amazon's Smart Speaker TechnologyabstractNative Americans are disproportionately affected by diabetes and diabetes complications. To control this disease, self-management, especially diet management is very important. There have appeared many electronic tools to help diabetic patients to manage their diet and control their blood glucose level. However, due to their lack of consideration of the special requirement of this ethical group, these tools are not well-accepted by Native American communities. In this paper, we propose a culturally appropriate tool to help this population to manage their disease. Specifically, we propose a voice-based Artificial Intelligence-powered virtual assistant to help Native American diabetic patients to manage their daily diet, and to learn food and nutrition-related knowledge. Voice is the most natural communication modality and it is easy to use without any technical background. In addition, the communication and recommendation provided by the system are personalized based on each user's physical, social, and cultural profile. Therefore, it would be easy to be accepted by the target audience. The proposed virtual assistant has been implemented on the Amazon Alexa platform. Preliminary experiments have demonstrated the usefulness of the virtual assistant. Bikesh Maharjan, Juan Li 0004, Cui Tao |
HealthCom | 4 |
| 2019 | Conceiving an application ontology to model patient human papillomavirus vaccine counseling for dialogue managementabstractBACKGROUND: In the United States and parts of the world, the human papillomavirus vaccine uptake is below the prescribed coverage rate for the population. Some research have noted that dialogue that communicates the risks and benefits, as well as patient concerns, can improve the uptake levels. In this paper, we introduce an application ontology for health information dialogue called Patient Health Information Dialogue Ontology for patient-level human papillomavirus vaccine counseling and potentially for any health-related counseling. RESULTS: The ontology's class level hierarchy is segmented into 4 basic levels - Discussion, Goal, Utterance, and Speech Task. The ontology also defines core low-level utterance interaction for communicating human papillomavirus health information. We discuss the design of the ontology and the execution of the utterance interaction. CONCLUSION: With an ontology that represents patient-centric dialogue to communicate health information, we have an application-driven model that formalizes the structure for the communication of health information, and a reusable scaffold that can be integrated for software agents. Our next step will to be develop the software engine that will utilize the ontology and automate the dialogue interaction of a software agent. Muhammad Amith, Kirk Roberts, Cui Tao |
BMC Bioinform. | 3 |
| 2019 | A 2018 workshop: vaccine and drug ontology studies (VDOS 2018)abstractThis Editorial first introduces the background of the vaccine and drug relations and how biomedical terminologies and ontologies have been used to support their studies. The history of the seven workshops, initially named VDOSME, and then named VDOS, is also summarized and introduced. Then the 7th International Workshop on Vaccine and Drug Ontology Studies (VDOS 2018), held on August 10th, 2018, Corvallis, Oregon, USA, is introduced in detail. These VDOS workshops have greatly supported the development, applications, and discussion of vaccine- and drug-related terminology and drug studies. Junguk Hur, Cui Tao, Yongqun He |
BMC Bioinform. | 2 |
| 2019 | ML-Net: multi-label classification of biomedical texts with deep neural networksabstractOBJECTIVE: In multi-label text classification, each textual document is assigned 1 or more labels. As an important task that has broad applications in biomedicine, a number of different computational methods have been proposed. Many of these methods, however, have only modest accuracy or efficiency and limited success in practical use. We propose ML-Net, a novel end-to-end deep learning framework, for multi-label classification of biomedical texts. MATERIALS AND METHODS: ML-Net combines a label prediction network with an automated label count prediction mechanism to provide an optimal set of labels. This is accomplished by leveraging both the predicted confidence score of each label and the deep contextual information (modeled by ELMo) in the target document. We evaluate ML-Net on 3 independent corpora in 2 text genres: biomedical literature and clinical notes. For evaluation, we use example-based measures, such as precision, recall, and the F measure. We also compare ML-Net with several competitive machine learning and deep learning baseline models. RESULTS: Our benchmarking results show that ML-Net compares favorably to state-of-the-art methods in multi-label classification of biomedical text. ML-Net is also shown to be robust when evaluated on different text genres in biomedicine. CONCLUSION: ML-Net is able to accuractely represent biomedical document context and dynamically estimate the label count in a more systematic and accurate manner. Unlike traditional machine learning methods, ML-Net does not require human effort for feature engineering and is a highly efficient and scalable approach to tasks with a large set of labels, so there is no need to build individual classifiers for each separate label. Jingcheng Du, Qingyu Chen 0001, Yifan Peng 0002, Yang Xiang 0003, Cui Tao, Zhiyong Lu |
J. Am. Medical Informatics Assoc. | 5 |
| 2019 | Network context matters: graph convolutional network model over social networks improves the detection of unknown HIV infections among young men who have sex with menabstractOBJECTIVE: HIV infection risk can be estimated based on not only individual features but also social network information. However, there have been insufficient studies using n machine learning methods that can maximize the utility of such information. Leveraging a state-of-the-art network topology modeling method, graph convolutional networks (GCN), our main objective was to include network information for the task of detecting previously unknown HIV infections. MATERIALS AND METHODS: We used multiple social network data (peer referral, social, sex partners, and affiliation with social and health venues) that include 378 young men who had sex with men in Houston, TX, collected between 2014 and 2016. Due to the limited sample size, an ensemble approach was engaged by integrating GCN for modeling information flow and statistical machine learning methods, including random forest and logistic regression, to efficiently model sparse features in individual nodes. RESULTS: Modeling network information using GCN effectively increased the prediction of HIV status in the social network. The ensemble approach achieved 96.6% on accuracy and 94.6% on F1 measure, which outperformed the baseline methods (GCN, logistic regression, and random forest: 79.0%, 90.5%, 94.4% on accuracy, respectively; and 57.7%, 80.2%, 90.4% on F1). In the networks with missing HIV status, the ensemble also produced promising results. CONCLUSION: Network context is a necessary component in modeling infectious disease transmissions such as HIV. GCN, when combined with traditional machine learning approaches, achieved promising performance in detecting previously unknown HIV infections, which may provide a useful tool for combatting the HIV epidemic. Yang Xiang 0003, Kayo Fujimoto, John A. Schneider, Yuxi Jia, Degui Zhi, Cui Tao |
J. Am. Medical Informatics Assoc. | 6 |
| 2018 | Development of HeartData, a Data Discovery Index Prototype for Cardiovascular data
Mandana Salimi, Anupama E. Gururaj, Cui Tao, Degui Zhi, Hua Xu 0001 |
AMIA | 4 |
| 2018 | Mining Human Papillomavirus Vaccination Health Beliefs from Twitter Using Attentive Recurrent Neural Network
Jingcheng Du, Fang Li 0011, Yuxi Jia, Yang Xiang 0003, Sahiti Myneni, Cui Tao |
AMIA | 7 |
| 2018 | Construction of Drug Repurposing-oriented Alzheimer's Disease Ontology
Fang Li 0011, Jingcheng Du, Guozheng Rao, Cui Tao |
AMIA | 4 |
| 2018 | Development and Validation of an Ontological Representation of the US Common Rule
Frank J. Manion, Marcelline R. Harris, Muhamamd F. Amith, Cui Tao |
AMIA | 4 |
| 2018 | Identification of Rare Adverse Events with Year-varying Reporting Rates for FLU4 Vaccine in VAERS
Jiayi Tong, Jing Huang 0021, Jingcheng Du, Cui Tao, Yong Chen 0016 |
AMIA | 5 |
| 2018 | Asthma Onset Prediction Using Structured EMR Data
Yang Xiang 0003, Jun Xu 0007, Jingcheng Du, Degui Zhi, Cui Tao |
AMIA | 5 |
| 2018 | Computer Aided Diagnosis of Chest X-rays using by Combining Deep Convolutional Neural Networks
Xinyuan Zhang 0006, Luca Giancardo, Cui Tao |
AMIA | 4 |
| 2018 | OntoKeeper: Semiotic-driven Ontology Evaluation Tool For Biomedical Ontologists
Muhammad Amith, Frank J. Manion, Chen Liang 0005, Marcelline R. Harris, Dennis Wang, Yongqun He, Cui Tao |
BIBM | 7 |
| 2018 | X-A-BiLSTM: a Deep Learning Approach for Depression Detection in Imbalanced Data
Qing Cong, Zhiyong Feng 0002, Fang Li 0011, Yang Xiang 0003, Guozheng Rao, Cui Tao |
BIBM | 6 |
| 2018 | Constructing Biomedical Knowledge Graph Based on SemMedDB and Linked Open Data
Qing Cong, Zhiyong Feng 0002, Fang Li 0011, Li Zhang 0059, Guozheng Rao, Cui Tao |
BIBM | 6 |
| 2018 | Comparing adverse effects of Hepatitis C drugs using FAERS data
Jing Huang 0021, Xinyuan Zhang 0003, Jiayi Tong, Jingcheng Du, Rui Duan 0004, Liu Yang 0026, Jason H. Moore, Yong Chen 0016, Cui Tao |
BIBM | 9 |
| 2018 | Toward a normalized clinical drug knowledge base in China - applying the RxNorm model to Chinese clinical drugsabstractObjective: In recent years, electronic health record systems have been widely implemented in China, making clinical data available electronically. However, little effort has been devoted to making drug information exchangeable among these systems. This study aimed to build a Normalized Chinese Clinical Drug (NCCD) knowledge base, by applying and extending the information model of RxNorm to Chinese clinical drugs. Methods: Chinese drugs were collected from 4 major resources-China Food and Drug Administration, China Health Insurance Systems, Hospital Pharmacy Systems, and China Pharmacopoeia-for integration and normalization in NCCD. Chemical drugs were normalized using the information model in RxNorm without much change. Chinese patent drugs (i.e., Chinese herbal extracts), however, were represented using an expanded RxNorm model to incorporate the unique characteristics of these drugs. A hybrid approach combining automated natural language processing technologies and manual review by domain experts was then applied to drug attribute extraction, normalization, and further generation of drug names at different specification levels. Lastly, we reported the statistics of NCCD, as well as the evaluation results using several sets of randomly selected Chinese drugs. Results: The current version of NCCD contains 16 976 chemical drugs and 2663 Chinese patent medicines, resulting in 19 639 clinical drugs, 250 267 unique concepts, and 2 602 760 relations. By manual review of 1700 chemical drugs and 250 Chinese patent drugs randomly selected from NCCD (about 10%), we showed that the hybrid approach could achieve an accuracy of 98.60% for drug name extraction and normalization. Using a collection of 500 chemical drugs and 500 Chinese patent drugs from other resources, we showed that NCCD achieved coverages of 97.0% and 90.0% for chemical drugs and Chinese patent drugs, respectively. Conclusion: Evaluation results demonstrated the potential to improve interoperability across various electronic drug systems in China. Li Wang 0077, Yaoyun Zhang, Min Jiang 0007, Jiancheng Dong, Yun Liu 0020, Cui Tao, Guoqian Jiang, Yi Zhou 0005, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 7 |
| 2018 | Assessing the practice of biomedical ontology evaluation: Gaps and opportunities
Muhammad Amith, Zhe He 0001, Jiang Bian 0001, Juan Antonio Lossio-Ventura, Cui Tao |
J. Biomed. Informatics | 5 |
| 2017 | From Scoping Review to Metadata: An Evidence-Based Approach to Identifying Informed Consent Metadata for Biorepositories
Marcelline R. Harris, Frank J. Manion, Hsing-yi Song, Yongqun He, Muhamamd F. Amith, Cui Tao |
AMIA | 6 |
| 2017 | Big Data to Knowledge (BD2K) and the Application of Metadata
Guoqian Jiang, Walter S. Campbell, Timothy Clark, Cui Tao, Mark A. Musen |
AMIA | 4 |
| 2017 | A Twitter Study of the Relationship between the Geographic Variations of HPV Vaccination Rates and Online Information in the United States
Hansi Zhang, Christopher Wheldon, Xinsong Du, Terrell Brown, Yi Guo 0005, Stephanie Staras, Jingcheng Du, Cui Tao, Jiang Bian 0001 |
AMIA | 9 |
| 2017 | A pilot study of mining association between psychiatric stressors and symptoms in tweetsabstractSuicide is a significant public health issue, causing huge impacts on individuals as well as their families. Psychiatric stressors are major suicide risk factors and can profoundly impact a person in many aspects. In order to facilitate the understanding of psychiatric stressors and the associated symptoms, we extracted stressors and symptoms terms from major online knowledge repositories and psychiatric clinical notes. The current vocabulary collection contains 1,292 psychiatric symptoms and 715 psychiatric stressors, which was leveraged to study the associations between stressors and symptoms in a corpus of suicide related tweets. Using Chi-Square test with Bonferroni correction, 3,500 symptom-stressor pairs were identified with significant association (p-value<;0.01). Jingcheng Du, Yaoyun Zhang, Cui Tao, Hua Xu 0001 |
BIBM | 3 |
| 2017 | Towards practical temporal relation extraction from clinical notes: An analysis of direct temporal relationsabstractFollowing the conventions developed in general domain, most of the current work on clinical temporal relation identification aims to identify a comprehensive set of temporal relations from source documents. This includes both explicit relations that is described in the documents and implicit relations that are identifiable only through inference. Although such an approach may provide a complete view of temporal information provided in a document, some temporal relations may not be practically essential, depending on the clinical application at hand. In addition, the performances of current systems that identify both explicit and implicit relations are still low and how to enhance the performances to be enough for practical use is not clear yet. In this paper, we propose focus on a subset of temporal relations, in order to provide insights into how to develop practically useful temporal information extraction methods for clinical text. We focus on “direct” temporal relations, which are intra-sentential temporal relations between a time expression and an event mention with limited syntactic distance. A corpus of 120 discharge summaries is constructed, leveraging an existing corpus, the 2012 i2b2 corpus. We show that the direct temporal relations constitute a major category of temporal relations. In addition, we show that the performance of the state-of-the art temporal relation extraction system, which is developed for both implicit and explicit relations, on direct temporal relations is still low. This indicates the need for development of methods tailored to direct temporal relations. Hee-Jin Lee, Yaoyun Zhang, Jun Xu 0007, Cui Tao, Hua Xu 0001, Min Jiang 0007 |
BIBM | 4 |
| 2017 | Designing an ontology for emotion-driven visual representationsabstractEmotions influence our perceptions and decisions and are often felt more strongly in situations related to healthcare. Therefore, it is important to understand how both providers and patients express their emotions in face-to-face scenarios. An ontology is a way to represent domain concepts and the relationships between them in a polyarchical manner. We have created an ontological model called the Visualized Emotion Ontology (VEO) that expresses the semantic definitions and visualizations of 25 emotions based on published research. With VEO, we can augment patient-facing software tools, like embodied conversational agents, to improve patient-provider interaction in clinical environments. Rebecca Lin, Muhammad Amith, Chen Liang 0005, Cui Tao |
BIBM | 4 |
| 2017 | Computer-aided diagnosis of four common cutaneous diseases using deep learning algorithmabstractWith the emergence of deep-learning algorithms, the accuracy of computer-aided supporting systems advanced., However, their adoption in the field of medicine has been limited, partially due to the challenges of generating reliable and timely results. In this research, we focused on classifying four common cutaneous diseases based on dermoscopic images using deep learning algorithms. Xinyuan Zhang 0006, Shiqi Wang 0008, Jie Liu 0085, Cui Tao |
BIBM | 4 |
| 2017 | Knowledge-Based Approach for Named Entity Recognition in Biomedical Literature: A Use Case in Biomedical Software Identification
Muhammad Amith, Yaoyun Zhang, Hua Xu 0001, Cui Tao |
IEA/AIE (2) | 4 |
| 2017 | Using Pathfinder networks to discover alignment between expert and consumer conceptual knowledge from online vaccine content
Muhammad Amith, Rachel Cunningham, Lara S. Savas, Julie Boom, Roger W. Schvaneveldt, Cui Tao, Trevor Cohen |
J. Biomed. Informatics | 6 |
| 2016 | A Scalable Dataset Indexing Infrastructure for the bioCADDIE Data Discovery System
Jeffrey S. Grethe, Ibrahim Burak Özyurt, Hua Xu 0001, Ruiling Liu, Ergin Soysal, Anupama E. Gururaj, Hyeon-Eui Kim, Trevor Cohen, Todd R. Johnson, Mandana Salimi, Saeid Pournejati, Min Jiang 0007, Claudiu Farcas, Alejandra N. González-Beltrán, Philippe Rocca-Serra, Muhamamd F. Amith, Cui Tao, Ian Fore, Ronald Margolis, George Alter, Susanna-Assunta Sansone, Lucila Ohno-Machado |
AMIA | 18 |
| 2016 | Development of DataMed, a Data Discovery Index Prototype by bioCADDIE: Laying the Groundwork for Biomedical Data Discovery
Hua Xu 0001, Jeffrey S. Grethe, Ruiling Liu, Ergin Soysal, Anupama E. Gururaj, Yueling Li, Ibrahim Burak Özyurt, Hyeon-Eui Kim, Trevor Cohen, Todd R. Johnson, Mandana Salimi, Saeid Pournejati, Min Jiang 0007, Claudiu Farcas, Alejandra N. González-Beltrán, Philippe Rocca-Serra, Muhamamd F. Amith, Cui Tao, Ian Fore, Ronald Margolis, George Alter, Susanna-Assunta Sansone, Lucila Ohno-Machado |
AMIA | 19 |
| 2016 | A representational analysis of a temporal indeterminancy display in clinical eventsabstractThis paper describes a proposition for representing temporal indeterminacy in events from clinical narratives using fuzzy sets membership functions. This approach leverages both temporal and semantic information of events and has been proved by representational analysis evaluation method. We demonstrate that membership functions' graphs can be used for representing temporal approximation and granularity of events. We also show that this approach is helpful for the construction of fine timeline of clinical events, and can be used for calculating accurate metrics for ordering events. Mohcine Madkour, Hsing-yi Song, Jingcheng Du, Cui Tao |
BIBM | 4 |
| 2015 | Model Checking for Verification of Interactive Health IT Systems
Keith A. Butler, Eric Mercer, Ali Bahrami, Cui Tao |
AMIA | 4 |
| 2015 | Quality Assurance of Cancer Study Common Data Elements Using A Post-Coordination Approach
Guoqian Jiang, Harold R. Solbrig, Eric Prud'hommeaux, Cui Tao, Chunhua Weng, Christopher G. Chute |
AMIA | 4 |
| 2015 | Analysis of empty responses from electronic resources in infobutton managers
Nathan C. Hulse, Cui Tao |
AMIA | 3 |
| 2015 | Transformation of standardized clinical models based on OWL technologies: from CEM to OpenEHR archetypesabstractINTRODUCTION: The semantic interoperability of electronic healthcare records (EHRs) systems is a major challenge in the medical informatics area. International initiatives pursue the use of semantically interoperable clinical models, and ontologies have frequently been used in semantic interoperability efforts. The objective of this paper is to propose a generic, ontology-based, flexible approach for supporting the automatic transformation of clinical models, which is illustrated for the transformation of Clinical Element Models (CEMs) into openEHR archetypes. METHODS: Our transformation method exploits the fact that the information models of the most relevant EHR specifications are available in the Web Ontology Language (OWL). The transformation approach is based on defining mappings between those ontological structures. We propose a way in which CEM entities can be transformed into openEHR by using transformation templates and OWL as common representation formalism. The transformation architecture exploits the reasoning and inferencing capabilities of OWL technologies. RESULTS: We have devised a generic, flexible approach for the transformation of clinical models, implemented for the unidirectional transformation from CEM to openEHR, a series of reusable transformation templates, a proof-of-concept implementation, and a set of openEHR archetypes that validate the methodological approach. CONCLUSIONS: We have been able to transform CEM into archetypes in an automatic, flexible, reusable transformation approach that could be extended to other clinical model specifications. We exploit the potential of OWL technologies for supporting the transformation process. We believe that our approach could be useful for international efforts in the area of semantic interoperability of EHR systems. María Del Carmen Legaz-García, Marcos Menárguez Tortosa, Jesualdo Tomás Fernández-Breis, Christopher G. Chute, Cui Tao |
J. Am. Medical Informatics Assoc. | 5 |
| 2014 | Enhancing the TURF Framework with a Workflow Ontology
Craig Harrington, Cui Tao, Keith A. Butler |
AMIA | 2 |
| 2014 | Enabling Locally-Developed Content For Access Through the Infobutton By Means of Automated Concept Annotation
Nathan C. Hulse, Cui Tao |
AMIA | 4 |
| 2014 | An Integrative Framework for Drug Target Prediction and Repurposing
Jingchun Sun, Cui Tao, Kevin W. Zhu, W. Jim Zheng, Hua Xu 0001 |
AMIA | 2 |
| 2014 | Pharmacological class data representation in the Web Ontology Language (OWL)abstractDozens of drug terminologies and resources capture drug and/or drug class information; they range greatly in their coverage and their adequacy of representation. However, there are no transformative ways to link these resources together in a standard and formal fashion, which hinders data integration and data representation for supporting drug-related clinical and translational studies. In this study, we generated a standardized Pharmacological Class Profile Ontology, named PCPO, which integrates multiple drug resources in the Web Ontology Language (OWL). More specifically, we mapped two well-known drug class resources, Anatomical Therapeutic Chemical classification system (ATC) and National Drug File Reference Terminology (NDF-RT), as the pharmacological class backbone. Furthermore we extended the PCPO with individual clinical drug information extracted from RxNorm and Structured Product Labeling. In parallel, we calculated and compared chemical structure similarity for each drug pair from ATC and NDF-RT, and re-grouped drugs with a similar structure into a same drug class. PCPO will not only present big drug data into well-organized formal form, OWL with possible inference capability, but also potentially support computational drug repurposing application development. Qian Zhu 0003, Cui Tao |
IEEE BigData | 2 |
| 2013 | Combining Infobuttons and Semantic Web Rules for Identifying Patterns and Delivering Highly-Personalized Education Materials
Nathan C. Hulse, Cui Tao |
AMIA | 3 |
| 2013 | Coreference Resolution from Medical Corpus with Topic Modeling
Dingcheng Li, Liwei Wang 0010, Cui Tao, Christopher G. Chute |
AMIA | 3 |
| 2013 | Adding Search Engine Functionality to Infobutton Manager Platform
Nathan C. Hulse, Cui Tao, Guilherme Del Fiol |
AMIA | 3 |
| 2013 | Introduction and Implementation of Common Terminology Services 2 (CTS2)
Cui Tao, Harold R. Solbrig, Craig Stancle, Kevin J. Peterson, Cory M. Endle, Scott Bauer, Deepak K. Sharma, Christopher G. Chute |
AMIA | 1 |
| 2013 | An integrative computational approach to identify disease-specific networks from PubMed literature informationabstractA huge amount of association relationships among biological entities (e.g., diseases, drugs, and genes) are scattered in biomedical literature. How to extract and analyze such heterogeneous data still remains a challenging task for most researchers in the biomedical field. Natural language processing (NLP) has the potential in extracting associations among biological entities from literature. However, association information extracted through NLP can be large, noisy, and redundant which poses significant challenges to biomedical researchers to use such information. To address this challenge, we propose a computational framework to facilitate the use of NLP results. We apply Latent Dirichlet Allocation (LDA) to discover topics based on associations. The networks extracted from each topic provide a disease-specific network for downstream bioinformatics analysis of associations for each topic. We illustrated the framework through the construction of disease-specific networks from Semantic MEDLINE, an NLP-generated association database, followed by the analysis of network properties, such as hub nodes and degree distribution. The results demonstrate that (1) LDA-based approach can group related diseases into the same disease topic; (2) the disease-specific association network follows the scale-free network property, in which hub nodes are enriched in related diseases, genes and drugs. Yuji Zhang 0001, Dingcheng Li, Cui Tao, Feichen Shen |
BIBM | 3 |
| 2013 | Comprehensive temporal information detection from clinical text: medical events, time, and TLINK identificationabstractBACKGROUND: Temporal information detection systems have been developed by the Mayo Clinic for the 2012 i2b2 Natural Language Processing Challenge. OBJECTIVE: To construct automated systems for EVENT/TIMEX3 extraction and temporal link (TLINK) identification from clinical text. MATERIALS AND METHODS: The i2b2 organizers provided 190 annotated discharge summaries as the training set and 120 discharge summaries as the test set. Our Event system used a conditional random field classifier with a variety of features including lexical information, natural language elements, and medical ontology. The TIMEX3 system employed a rule-based method using regular expression pattern match and systematic reasoning to determine normalized values. The TLINK system employed both rule-based reasoning and machine learning. All three systems were built in an Apache Unstructured Information Management Architecture framework. RESULTS: Our TIMEX3 system performed the best (F-measure of 0.900, value accuracy 0.731) among the challenge teams. The Event system produced an F-measure of 0.870, and the TLINK system an F-measure of 0.537. CONCLUSIONS: Our TIMEX3 system demonstrated good capability of regular expression rules to extract and normalize time information. Event and TLINK machine learning systems required well-defined feature sets to perform well. We could also leverage expert knowledge as part of the machine learning features to further improve TLINK identification performance. Sunghwan Sohn, Kavishwar B. Wagholikar, Dingcheng Li, Siddhartha Jonnalagadda, Cui Tao, K. E. Ravikumar |
J. Am. Medical Informatics Assoc. | 5 |
| 2013 | A semantic-web oriented representation of the clinical element model for secondary use of electronic health records dataabstractThe clinical element model (CEM) is an information model designed for representing clinical information in electronic health records (EHR) systems across organizations. The current representation of CEMs does not support formal semantic definitions and therefore it is not possible to perform reasoning and consistency checking on derived models. This paper introduces our efforts to represent the CEM specification using the Web Ontology Language (OWL). The CEM-OWL representation connects the CEM content with the Semantic Web environment, which provides authoring, reasoning, and querying tools. This work may also facilitate the harmonization of the CEMs with domain knowledge represented in terminology models as well as other clinical information models such as the openEHR archetype model. We have created the CEM-OWL meta ontology based on the CEM specification. A convertor has been implemented in Java to automatically translate detailed CEMs from XML to OWL. A panel evaluation has been conducted, and the results show that the OWL modeling can faithfully represent the CEM specification and represent patient data. Cui Tao, Guoqian Jiang, Thomas A. Oniki, Robert R. Freimuth, Qian Zhu 0003, Deepak K. Sharma, Jyotishman Pathak, Stanley M. Huff, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 1 |
| 2013 | Terminology representation guidelines for biomedical ontologies in the semantic web notations
Cui Tao, Jyotishman Pathak, Harold R. Solbrig, Wei-Qi Wei, Christopher G. Chute |
J. Biomed. Informatics | 1 |
| 2013 | Semantator: Semantic annotator for converting biomedical text to linked data
Cui Tao, Dezhao Song, Deepak K. Sharma, Christopher G. Chute |
J. Biomed. Informatics | 1 |
| 2013 | Harmonization and semantic annotation of data dictionaries from the Pharmacogenomics Research Network: A case study
Qian Zhu 0003, Robert R. Freimuth, Zonghui Lian, Scott Bauer, Jyotishman Pathak, Cui Tao, Matthew J. Durski, Christopher G. Chute |
J. Biomed. Informatics | 6 |
| 2012 | Common Terminology Services 2 (CTS2) for Biomedical Community
Cui Tao, Harold R. Solbrig, Pradip Kanjamala, Kevin J. Peterson, Craig Stancle, Christopher G. Chute |
AMIA | 1 |
| 2012 | Managing interoperability and compleXity in health systems - MIXHS'12abstractData management and knowledge engineering have long been important research fields in computer science, and rapid progress in recent years have increasingly seen these technologies successfully applied to solve complex biomedical challenges and support health services professionals in the course of their intellectually-demanding clinical duties, such as through the use of decision-support or expert systems. Yet, as the biomedical knowledge available in the modern digital world grows exponentially, there is a pressing need for a focused forum to promote technology and knowledge transfer from basic research to biomedical applications as well as allowing for implementers of healthcare systems to share their experiences with the research community. The Managing Interoperability and Complexity in Health Systems, MIXHS workshops are designed for such a purpose with the view that multi-disciplinary approaches within a holistic forum is essential to rise to the ever new challenges of biomedical knowledge complexity and interoperability of health systems and services. Cui Tao, Matt-Mouley Bouamrane |
CIKM | 1 |
| 2012 | Unified Medical Language System term occurrences in clinical notes: a large-scale corpus analysisabstractOBJECTIVE: To characterise empirical instances of Unified Medical Language System (UMLS) Metathesaurus term strings in a large clinical corpus, and to illustrate what types of term characteristics are generalisable across data sources. DESIGN: Based on the occurrences of UMLS terms in a 51 million document corpus of Mayo Clinic clinical notes, this study computes statistics about the terms' string attributes, source terminologies, semantic types and syntactic categories. Term occurrences in 2010 i2b2/VA text were also mapped; eight example filters were designed from the Mayo-based statistics and applied to i2b2/VA data. RESULTS: For the corpus analysis, negligible numbers of mapped terms in the Mayo corpus had over six words or 55 characters. Of source terminologies in the UMLS, the Consumer Health Vocabulary and Systematized Nomenclature of Medicine-Clinical Terms (SNOMED-CT) had the best coverage in Mayo clinical notes at 106426 and 94788 unique terms, respectively. Of 15 semantic groups in the UMLS, seven groups accounted for 92.08% of term occurrences in Mayo data. Syntactically, over 90% of matched terms were in noun phrases. For the cross-institutional analysis, using five example filters on i2b2/VA data reduces the actual lexicon to 19.13% of the size of the UMLS and only sees a 2% reduction in matched terms. CONCLUSION: The corpus statistics presented here are instructive for building lexicons from the UMLS. Features intrinsic to Metathesaurus terms (well formedness, length and language) generalise easily across clinical institutions, but term frequencies should be adapted with caution. The semantic groups of mapped terms may differ slightly from institution to institution, but they differ greatly when moving to the biomedical literature domain. Stephen T. Wu, Dingcheng Li, Cui Tao, Mark A. Musen, Christopher G. Chute, Nigam H. Shah |
J. Am. Medical Informatics Assoc. | 4 |
| 2012 | Building a robust, scalable and standards-driven infrastructure for secondary use of EHR data: The SHARPn project
Susan Rea, Jyotishman Pathak, Guergana K. Savova, Thomas A. Oniki, Les Westberg, Calvin E. Beebe, Cui Tao, Craig G. Parker, Peter J. Haug, Stanley M. Huff, Christopher G. Chute |
J. Biomed. Informatics | 7 |
| 2011 | Managing interoperability and complexity inhealth systems: MIXHS'11 workshop summaryabstractManaging Interoperability and Complexity in Health Systems, MIXHS'11, aims to be a forum focussing on recent research and technical results in knowledge management and information systems in bio-medical and electronic health systems. The workshop will provide an opportunity for sharing practical experiences and best practices in e-Health information infrastructure development and management. Of particular interest to the workshop themes are technical solutions to recurring practical systems deployment issues, including harnessing the complexity of bio-medical domain knowledge and the interoperability of heterogeneous health systems. The workshop will gather experts, researchers, system developers, practitioners and policymakers designing and implementing solutions for managing clinical data and integrating existing and future electronic health systems infrastructures. Matt-Mouley Bouamrane, Cui Tao |
CIKM | 2 |
| 2010 | Time-Oriented Question Answering from Clinical Narratives Using Semantic-Web Techniques
Cui Tao, Harold R. Solbrig, Deepak K. Sharma, Wei-Qi Wei, Guergana K. Savova, Christopher G. Chute |
ISWC (2) | 1 |
| 2009 | FOCIH: Form-Based Ontology Creation and Information Harvesting
Cui Tao, David W. Embley, Stephen W. Liddle |
ER | 1 |
| 2009 | Automatic hidden-web table interpretation, conceptualization, and semantic annotation
Cui Tao, David W. Embley |
Data Knowl. Eng. | 1 |
| 2008 | A Conceptual-Model-Based Computational Alembic for a Web of Knowledge
David W. Embley, Stephen W. Liddle, Deryle W. Lonsdale, George Nagy, Yuri A. Tijerino, Robert Clawson, Jordan Crabtree, Yihong Ding, Piyushee Jha, Zonghui Lian, Stephen Lynn, Raghav K. Padmanabhan, Jeff Peters, Cui Tao, Robby Watts, Charla Woodbury, Andrew Zitzelberger |
ER | 14 |
| 2007 | Automatic Hidden-Web Table Interpretation by Sibling Page Comparison
Cui Tao, David W. Embley |
ER | 1 |
| 2006 | Toward Making Online Biological Data Machine Understandable
Cui Tao |
ISWC | 1 |
| 2005 | Automating the extraction of data from HTML tables with unknown structure
David W. Embley, Cui Tao, Stephen W. Liddle |
Data Knowl. Eng. | 2 |
| 2002 | Automatically Extracting Ontologically Specified Data from HTML Tables of Unknown Structure
David W. Embley, Cui Tao, Stephen W. Liddle |
ER | 2 |