EDBT 2026 Demo / reviewers in the wild / expert
Jingcheng Du
dblp:156/1121
· DBLP profile ↗
28ranked-venue papers
5as first author
13since 2021 · last 2024
0000-0002-0322-4566ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 28 · 5 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | AutoCriteria: a generalizable clinical trial eligibility criteria extraction system powered by large language modelsabstractOBJECTIVES: We aim to build a generalizable information extraction system leveraging large language models to extract granular eligibility criteria information for diverse diseases from free text clinical trial protocol documents. We investigate the model's capability to extract criteria entities along with contextual attributes including values, temporality, and modifiers and present the strengths and limitations of this system. MATERIALS AND METHODS: The clinical trial data were acquired from https://ClinicalTrials.gov/. We developed a system, AutoCriteria, which comprises the following modules: preprocessing, knowledge ingestion, prompt modeling based on GPT, postprocessing, and interim evaluation. The final system evaluation was performed, both quantitatively and qualitatively, on 180 manually annotated trials encompassing 9 diseases. RESULTS: AutoCriteria achieves an overall F1 score of 89.42 across all 9 diseases in extracting the criteria entities, with the highest being 95.44 for nonalcoholic steatohepatitis and the lowest of 84.10 for breast cancer. Its overall accuracy is 78.95% in identifying all contextual information across all diseases. Our thematic analysis indicated accurate logic interpretation of criteria as one of the strengths and overlooking/neglecting the main criteria as one of the weaknesses of AutoCriteria. DISCUSSION: AutoCriteria demonstrates strong potential to extract granular eligibility criteria information from trial documents without requiring manual annotations. The prompts developed for AutoCriteria generalize well across different disease areas. Our evaluation suggests that the system handles complex scenarios including multiple arm conditions and logics. CONCLUSION: AutoCriteria currently encompasses a diverse range of diseases and holds potential to extend to more in the future. This signifies a generalizable and scalable solution, poised to address the complexities of clinical trial application in real-world settings. Surabhi Datta, Kyeryoung Lee, Hunki Paek, Frank J. Manion, Nneka Ofoegbu, Jingcheng Du, Liang-Chin Huang, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 6 |
| 2024 | Improving large language models for clinical named entity recognition via prompt engineeringabstractIMPORTANCE: The study highlights the potential of large language models, specifically GPT-3.5 and GPT-4, in processing complex clinical data and extracting meaningful information with minimal training data. By developing and refining prompt-based strategies, we can significantly enhance the models' performance, making them viable tools for clinical NER tasks and possibly reducing the reliance on extensive annotated datasets. OBJECTIVES: This study quantifies the capabilities of GPT-3.5 and GPT-4 for clinical named entity recognition (NER) tasks and proposes task-specific prompts to improve their performance. MATERIALS AND METHODS: We evaluated these models on 2 clinical NER tasks: (1) to extract medical problems, treatments, and tests from clinical notes in the MTSamples corpus, following the 2010 i2b2 concept extraction shared task, and (2) to identify nervous system disorder-related adverse events from safety reports in the vaccine adverse event reporting system (VAERS). To improve the GPT models' performance, we developed a clinical task-specific prompt framework that includes (1) baseline prompts with task description and format specification, (2) annotation guideline-based prompts, (3) error analysis-based instructions, and (4) annotated samples for few-shot learning. We assessed each prompt's effectiveness and compared the models to BioClinicalBERT. RESULTS: Using baseline prompts, GPT-3.5 and GPT-4 achieved relaxed F1 scores of 0.634, 0.804 for MTSamples and 0.301, 0.593 for VAERS. Additional prompt components consistently improved model performance. When all 4 components were used, GPT-3.5 and GPT-4 achieved relaxed F1 socres of 0.794, 0.861 for MTSamples and 0.676, 0.736 for VAERS, demonstrating the effectiveness of our prompt framework. Although these results trail BioClinicalBERT (F1 of 0.901 for the MTSamples dataset and 0.802 for the VAERS), it is very promising considering few training samples are needed. DISCUSSION: The study's findings suggest a promising direction in leveraging LLMs for clinical NER tasks. However, while the performance of GPT models improved with task-specific prompts, there's a need for further development and refinement. LLMs like GPT-4 show potential in achieving close performance to state-of-the-art models like BioClinicalBERT, but they still require careful prompt engineering and understanding of task-specific knowledge. The study also underscores the importance of evaluation schemas that accurately reflect the capabilities and performance of LLMs in clinical settings. CONCLUSION: While direct application of GPT models to clinical NER tasks falls short of optimal performance, our task-specific prompt framework, incorporating medical knowledge and training samples, significantly enhances GPT models' feasibility for potential clinical applications. Qingyu Chen 0001, Jingcheng Du, Xueqing Peng, Vipina Kuttichi Keloth, Xu Zuo, Yujia Zhou 0003, Zehan Li, Xiaoqian Jiang, Zhiyong Lu, Kirk Roberts, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 3 |
| 2023 | Machine learning-based donor permission extraction from informed consent documentsabstractBACKGROUND: With more clinical trials are offering optional participation in the collection of bio-specimens for biobanking comes the increasing complexity of requirements of informed consent forms. The aim of this study is to develop an automatic natural language processing (NLP) tool to annotate informed consent documents to promote biorepository data regulation, sharing, and decision support. We collected informed consent documents from several publicly available sources, then manually annotated them, covering sentences containing permission information about the sharing of either bio-specimens or donor data, or conducting genetic research or future research using bio-specimens or donor data. RESULTS: We evaluated a variety of machine learning algorithms including random forest (RF) and support vector machine (SVM) for the automatic identification of these sentences. 120 informed consent documents containing 29,204 sentences were annotated, of which 1250 sentences (4.28%) provide answers to a permission question. A support vector machine (SVM) model achieved a F-1 score of 0.95 on classifying the sentences when using a gold standard, which is a prefiltered corpus containing all relevant sentences. CONCLUSIONS: This study provides the feasibility of using machine learning tools to classify permission-related sentences in informed consent documents. Madhuri Sankaranarayanapillai, Jingcheng Du, Yang Xiang 0003, Frank J. Manion, Marcelline R. Harris, Cooper Stansbury, Huy Anh Pham, Cui Tao |
BMC Bioinform. | 3 |
| 2023 | Systematic design and data-driven evaluation of social determinants of health ontology (SDoHO)abstractOBJECTIVE: Social determinants of health (SDoH) play critical roles in health outcomes and well-being. Understanding the interplay of SDoH and health outcomes is critical to reducing healthcare inequalities and transforming a "sick care" system into a "health-promoting" system. To address the SDOH terminology gap and better embed relevant elements in advanced biomedical informatics, we propose an SDoH ontology (SDoHO), which represents fundamental SDoH factors and their relationships in a standardized and measurable way. MATERIAL AND METHODS: Drawing on the content of existing ontologies relevant to certain aspects of SDoH, we used a top-down approach to formally model classes, relationships, and constraints based on multiple SDoH-related resources. Expert review and coverage evaluation, using a bottom-up approach employing clinical notes data and a national survey, were performed. RESULTS: We constructed the SDoHO with 708 classes, 106 object properties, and 20 data properties, with 1,561 logical axioms and 976 declaration axioms in the current version. Three experts achieved 0.967 agreement in the semantic evaluation of the ontology. A comparison between the coverage of the ontology and SDOH concepts in 2 sets of clinical notes and a national survey instrument also showed satisfactory results. DISCUSSION: SDoHO could potentially play an essential role in providing a foundation for a comprehensive understanding of the associations between SDoH and health outcomes and paving the way for health equity across populations. CONCLUSION: SDoHO has well-designed hierarchies, practical objective properties, and versatile functionalities, and the comprehensive semantic and coverage evaluation achieved promising performance compared to the existing ontologies relevant to SDoH. Yifang Dang, Fang Li 0011, Xinyue Hu 0002, Vipina Kuttichi Keloth, Sunyang Fu, Muhammad Amith, J. Wilfred Fan, Jingcheng Du, Evan Yu, Xiaoqian Jiang, Hua Xu 0001, Cui Tao |
J. Am. Medical Informatics Assoc. | 9 |
| 2023 | SHAPE: A Sample-Adaptive Hierarchical Prediction Network for Medication RecommendationabstractEffectively medication recommendation with complex multimorbidity conditions is a critical yet challenging task in healthcare. Most existing works predicted medications based on longitudinal records, which assumed the encoding format of intra-visit medical events are serialized and information transmitted patterns of learning longitudinal sequence data are stable. However, the following conditions may have been ignored: 1) A more compact encoder for intra-relationship in the intra-visit medical event is urgent; 2) Strategies for learning accurate representations of the variable longitudinal sequences of patients are different. In this article, we proposed a novel Sample-adaptive Hierarchical medicAtion Prediction nEtwork, termed SHAPE, to tackle the above challenges in the medication recommendation task. Specifically, we design a compact intra-visit set encoder to encode the relationship in the medical event for obtaining visit-level representation and then develop an inter-visit longitudinal encoder to learn the patient-level longitudinal representation efficiently. To endow the model with the capability of modeling the variable visit length, we introduce a soft curriculum learning method to assign the difficulty of each sample automatically by the visit length. Extensive experiments on a benchmark dataset verify the superiority of our model compared with several state-of-the-art baselines. Sicen Liu, Xiaolong Wang 0001, Jingcheng Du, Yongshuai Hou, Xianbing Zhao, Hui Wang 0030, Yang Xiang 0003, Buzhou Tang |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | Combining Transfer Learning with Graph Attention Models for HIV Risk Prediction
Evan Yu, Cui Tao, Jingcheng Du, Degui Zhi, Yang Xiang 0003, Kayo Fujimoto, John A. Schneider |
AMIA | 3 |
| 2022 | Mining on Alzheimer's diseases related knowledge graph to identity potential AD-related semantic triples for drug repurposingabstractBACKGROUND: To date, there are no effective treatments for most neurodegenerative diseases. Knowledge graphs can provide comprehensive and semantic representation for heterogeneous data, and have been successfully leveraged in many biomedical applications including drug repurposing. Our objective is to construct a knowledge graph from literature to study the relations between Alzheimer's disease (AD) and chemicals, drugs and dietary supplements in order to identify opportunities to prevent or delay neurodegenerative progression. We collected biomedical annotations and extracted their relations using SemRep via SemMedDB. We used both a BERT-based classifier and rule-based methods during data preprocessing to exclude noise while preserving most AD-related semantic triples. The 1,672,110 filtered triples were used to train with knowledge graph completion algorithms (i.e., TransE, DistMult, and ComplEx) to predict candidates that might be helpful for AD treatment or prevention. RESULTS: Among three knowledge graph completion models, TransE outperformed the other two (MR = 10.53, Hits@1 = 0.28). We leveraged the time-slicing technique to further evaluate the prediction results. We found supporting evidence for most highly ranked candidates predicted by our model which indicates that our approach can inform reliable new knowledge. CONCLUSION: This paper shows that our graph mining model can predict reliable new relationships between AD and other entities (i.e., dietary supplements, chemicals, and drugs). The knowledge graph constructed can facilitate data-driven knowledge discoveries and the generation of novel hypotheses. Yi Nian, Xinyue Hu 0002, Rui Zhang 0028, Jingna Feng, Jingcheng Du, Fang Li 0011, Larry Bu, Yuji Zhang 0001, Yong Chen 0016, Cui Tao |
BMC Bioinform. | 5 |
| 2022 | LitMC-BERT: Transformer-Based Multi-Label Classification of Biomedical Literature With An Application on COVID-19 Literature CurationabstractThe rapid growth of biomedical literature poses a significant challenge for curation and interpretation. This has become more evident during the COVID-19 pandemic. LitCovid, a literature database of COVID-19 related papers in PubMed, has accumulated over 200,000 articles with millions of accesses. Approximately 10,000 new articles are added to LitCovid every month. A main curation task in LitCovid is topic annotation where an article is assigned with up to eight topics, e.g., Treatment and Diagnosis. The annotated topics have been widely used both in LitCovid (e.g., accounting for ∼18% of total uses) and downstream studies such as network generation. However, it has been a primary curation bottleneck due to the nature of the task and the rapid literature growth. This study proposes LITMC-BERT, a transformer-based multi-label classification method in biomedical literature. It uses a shared transformer backbone for all the labels while also captures label-specific features and the correlations between label pairs. We compare LITMC-BERT with three baseline models on two datasets. Its micro-F1 and instance-based F1 are 5% and 4% higher than the current best results, respectively, and only requires ∼18% of the inference time than the Binary BERT baseline. The related datasets and models are available via https://github.com/ncbi/ml-transformer. Qingyu Chen 0001, Jingcheng Du, Alexis Allot, Zhiyong Lu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | AM2BERT: attention guided and regularized transformer-based multi-label classification model for COVID-19 literature curation
Qingyu Chen 0001, Jingcheng Du, Alexis Allot, Zhiyong Lu |
AMIA | 2 |
| 2021 | Data and Model Biases in Social Media Analyses: A Case Study of COVID-19 Tweets
Pengfei Yin, Yongqiu Li, Xing He 0003, Jingcheng Du, Cui Tao, Yi Guo 0005, Mattia Prosperi, Pierangelo Veltri, Xi Yang 0015, Yonghui Wu 0001, Jiang Bian 0001 |
AMIA | 5 |
| 2021 | Extracting postmarketing adverse events from safety reports in the vaccine adverse event reporting system (VAERS) using deep learningabstractOBJECTIVE: Automated analysis of vaccine postmarketing surveillance narrative reports is important to understand the progression of rare but severe vaccine adverse events (AEs). This study implemented and evaluated state-of-the-art deep learning algorithms for named entity recognition to extract nervous system disorder-related events from vaccine safety reports. MATERIALS AND METHODS: We collected Guillain-Barré syndrome (GBS) related influenza vaccine safety reports from the Vaccine Adverse Event Reporting System (VAERS) from 1990 to 2016. VAERS reports were selected and manually annotated with major entities related to nervous system disorders, including, investigation, nervous_AE, other_AE, procedure, social_circumstance, and temporal_expression. A variety of conventional machine learning and deep learning algorithms were then evaluated for the extraction of the above entities. We further pretrained domain-specific BERT (Bidirectional Encoder Representations from Transformers) using VAERS reports (VAERS BERT) and compared its performance with existing models. RESULTS AND CONCLUSIONS: Ninety-one VAERS reports were annotated, resulting in 2512 entities. The corpus was made publicly available to promote community efforts on vaccine AEs identification. Deep learning-based methods (eg, bi-long short-term memory and BERT models) outperformed conventional machine learning-based methods (ie, conditional random fields with extensive features). The BioBERT large model achieved the highest exact match F-1 scores on nervous_AE, procedure, social_circumstance, and temporal_expression; while VAERS BERT large models achieved the highest exact match F-1 scores on investigation and other_AE. An ensemble of these 2 models achieved the highest exact match microaveraged F-1 score at 0.6802 and the second highest lenient match microaveraged F-1 score at 0.8078 among peer models. Jingcheng Du, Yang Xiang 0003, Madhuri Sankaranarayanapillai, Yuqi Si, Huy Anh Pham, Hua Xu 0001, Yong Chen 0016, Cui Tao |
J. Am. Medical Informatics Assoc. | 1 |
| 2021 | COVID-19 trial graph: a linked graph for COVID-19 clinical trialsabstractOBJECTIVE: Clinical trials are an essential part of the effort to find safe and effective prevention and treatment for COVID-19. Given the rapid growth of COVID-19 clinical trials, there is an urgent need for a better clinical trial information retrieval tool that supports searching by specifying criteria, including both eligibility criteria and structured trial information. MATERIALS AND METHODS: We built a linked graph for registered COVID-19 clinical trials: the COVID-19 Trial Graph, to facilitate retrieval of clinical trials. Natural language processing tools were leveraged to extract and normalize the clinical trial information from both their eligibility criteria free texts and structured information from ClinicalTrials.gov. We linked the extracted data using the COVID-19 Trial Graph and imported it to a graph database, which supports both querying and visualization. We evaluated trial graph using case queries and graph embedding. RESULTS: The graph currently (as of October 5, 2020) contains 3392 registered COVID-19 clinical trials, with 17 480 nodes and 65 236 relationships. Manual evaluation of case queries found high precision and recall scores on retrieving relevant clinical trials searching from both eligibility criteria and trial-structured information. We observed clustering in clinical trials via graph embedding, which also showed superiority over the baseline (0.870 vs 0.820) in evaluating whether a trial can complete its recruitment successfully. CONCLUSIONS: The COVID-19 Trial Graph is a novel representation of clinical trials that allows diverse search queries and provides a graph-based visualization of COVID-19 clinical trials. High-dimensional vectors mapped by graph embedding for clinical trials would be potentially beneficial for many downstream applications, such as trial end recruitment status prediction and trial similarity comparison. Our methodology also is generalizable to other clinical trials. Jingcheng Du, Prerana Ramesh, Yang Xiang 0003, Xiaoqian Jiang, Cui Tao |
J. Am. Medical Informatics Assoc. | 1 |
| 2021 | Deep representation learning of patient data from Electronic Health Records (EHR): A systematic review
Yuqi Si, Jingcheng Du, Xiaoqian Jiang, Timothy A. Miller, Fei Wang 0001, W. Jim Zheng, Kirk Roberts |
J. Biomed. Informatics | 2 |
| 2020 | Deep Representation Learning of Patient Data from Electronic Health Records: A Systematic Review
Yuqi Si, Jingcheng Du, Xiaoqian Jiang, Timothy A. Miller, Fei Wang 0001, W. Jim Zheng, Kirk Roberts |
AMIA | 2 |
| 2020 | Time event ontology (TEO): to support semantic representation and reasoning of complex temporal relations of clinical eventsabstractOBJECTIVE: The goal of this study is to develop a robust Time Event Ontology (TEO), which can formally represent and reason both structured and unstructured temporal information. MATERIALS AND METHODS: Using our previous Clinical Narrative Temporal Relation Ontology 1.0 and 2.0 as a starting point, we redesigned concept primitives (clinical events and temporal expressions) and enriched temporal relations. Specifically, 2 sets of temporal relations (Allen's interval algebra and a novel suite of basic time relations) were used to specify qualitative temporal order relations, and a Temporal Relation Statement was designed to formalize quantitative temporal relations. Moreover, a variety of data properties were defined to represent diversified temporal expressions in clinical narratives. RESULTS: TEO has a rich set of classes and properties (object, data, and annotation). When evaluated with real electronic health record data from the Mayo Clinic, it could faithfully represent more than 95% of the temporal expressions. Its reasoning ability was further demonstrated on a sample drug adverse event report annotated with respect to TEO. The results showed that our Java-based TEO reasoner could answer a set of frequently asked time-related queries, demonstrating that TEO has a strong capability of reasoning complex temporal relations. CONCLUSION: TEO can support flexible temporal relation representation and reasoning. Our next step will be to apply TEO to the natural language processing field to facilitate automated temporal information annotation, extraction, and timeline reasoning to better support time-based clinical decision-making. Fang Li 0011, Jingcheng Du, Yongqun He, Hsing-yi Song, Mohcine Madkour, Guozheng Rao, Yang Xiang 0003, Henry W. Chen, Sijia Liu 0002, Liwei Wang 0010, Hua Xu 0001, Cui Tao |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | A study of deep learning approaches for medication and adverse drug event extraction from clinical textabstractOBJECTIVE: This article presents our approaches to extraction of medications and associated adverse drug events (ADEs) from clinical documents, which is the second track of the 2018 National NLP Clinical Challenges (n2c2) shared task. MATERIALS AND METHODS: The clinical corpus used in this study was from the MIMIC-III database and the organizers annotated 303 documents for training and 202 for testing. Our system consists of 2 components: a named entity recognition (NER) and a relation classification (RC) component. For each component, we implemented deep learning-based approaches (eg, BI-LSTM-CRF) and compared them with traditional machine learning approaches, namely, conditional random fields for NER and support vector machines for RC, respectively. In addition, we developed a deep learning-based joint model that recognizes ADEs and their relations to medications in 1 step using a sequence labeling approach. To further improve the performance, we also investigated different ensemble approaches to generating optimal performance by combining outputs from multiple approaches. RESULTS: Our best-performing systems achieved F1 scores of 93.45% for NER, 96.30% for RC, and 89.05% for end-to-end evaluation, which ranked #2, #1, and #1 among all participants, respectively. Additional evaluations show that the deep learning-based approaches did outperform traditional machine learning algorithms in both NER and RC. The joint model that simultaneously recognizes ADEs and their relations to medications also achieved the best performance on RC, indicating its promise for relation extraction. CONCLUSION: In this study, we developed deep learning approaches for extracting medications and their attributes such as ADEs, and demonstrated its superior performance compared with traditional machine learning algorithms, indicating its uses in broader NER and RC tasks in the medical domain. Qiang Wei 0002, Zongcheng Ji, Zhiheng Li 0004, Jingcheng Du, Jun Xu 0007, Yang Xiang 0003, Firat Tiryaki, Stephen Wu 0004, Yaoyun Zhang, Cui Tao, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 4 |
| 2020 | Deep learning in clinical natural language processing: a methodical reviewabstractOBJECTIVE: This article methodically reviews the literature on deep learning (DL) for natural language processing (NLP) in the clinical domain, providing quantitative analysis to answer 3 research questions concerning methods, scope, and context of current research. MATERIALS AND METHODS: We searched MEDLINE, EMBASE, Scopus, the Association for Computing Machinery Digital Library, and the Association for Computational Linguistics Anthology for articles using DL-based approaches to NLP problems in electronic health records. After screening 1,737 articles, we collected data on 25 variables across 212 papers. RESULTS: DL in clinical NLP publications more than doubled each year, through 2018. Recurrent neural networks (60.8%) and word2vec embeddings (74.1%) were the most popular methods; the information extraction tasks of text classification, named entity recognition, and relation extraction were dominant (89.2%). However, there was a "long tail" of other methods and specific tasks. Most contributions were methodological variants or applications, but 20.8% were new methods of some kind. The earliest adopters were in the NLP community, but the medical informatics community was the most prolific. DISCUSSION: Our analysis shows growing acceptance of deep learning as a baseline for NLP research, and of DL-based NLP in the medical community. A number of common associations were substantiated (eg, the preference of recurrent neural networks for sequence-labeling named entity recognition), while others were surprisingly nuanced (eg, the scarcity of French language clinical NLP with deep learning). CONCLUSION: Deep learning has not yet fully penetrated clinical NLP and is growing rapidly. This review highlighted both the popular and unique trends in this active field. Stephen Wu 0004, Kirk Roberts, Surabhi Datta, Jingcheng Du, Zongcheng Ji, Yuqi Si, Sarvesh Soni, Qiang Wei 0002, Yang Xiang 0003, Bo Zhao 0001, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 4 |
| 2019 | Relation Extraction from Clinical Narratives Using Pre-trained Language Models
Qiang Wei 0002, Zongcheng Ji, Yuqi Si, Jingcheng Du, Firat Tiryaki, Stephen Wu 0004, Cui Tao, Kirk Roberts, Hua Xu 0001 |
AMIA | 4 |
| 2019 | ML-Net: multi-label classification of biomedical texts with deep neural networksabstractOBJECTIVE: In multi-label text classification, each textual document is assigned 1 or more labels. As an important task that has broad applications in biomedicine, a number of different computational methods have been proposed. Many of these methods, however, have only modest accuracy or efficiency and limited success in practical use. We propose ML-Net, a novel end-to-end deep learning framework, for multi-label classification of biomedical texts. MATERIALS AND METHODS: ML-Net combines a label prediction network with an automated label count prediction mechanism to provide an optimal set of labels. This is accomplished by leveraging both the predicted confidence score of each label and the deep contextual information (modeled by ELMo) in the target document. We evaluate ML-Net on 3 independent corpora in 2 text genres: biomedical literature and clinical notes. For evaluation, we use example-based measures, such as precision, recall, and the F measure. We also compare ML-Net with several competitive machine learning and deep learning baseline models. RESULTS: Our benchmarking results show that ML-Net compares favorably to state-of-the-art methods in multi-label classification of biomedical text. ML-Net is also shown to be robust when evaluated on different text genres in biomedicine. CONCLUSION: ML-Net is able to accuractely represent biomedical document context and dynamically estimate the label count in a more systematic and accurate manner. Unlike traditional machine learning methods, ML-Net does not require human effort for feature engineering and is a highly efficient and scalable approach to tasks with a large set of labels, so there is no need to build individual classifiers for each separate label. Jingcheng Du, Qingyu Chen 0001, Yifan Peng 0002, Yang Xiang 0003, Cui Tao, Zhiyong Lu |
J. Am. Medical Informatics Assoc. | 1 |
| 2018 | Mining Human Papillomavirus Vaccination Health Beliefs from Twitter Using Attentive Recurrent Neural Network
Jingcheng Du, Fang Li 0011, Yuxi Jia, Yang Xiang 0003, Sahiti Myneni, Cui Tao |
AMIA | 1 |
| 2018 | Construction of Drug Repurposing-oriented Alzheimer's Disease Ontology
Fang Li 0011, Jingcheng Du, Guozheng Rao, Cui Tao |
AMIA | 2 |
| 2018 | Identification of Rare Adverse Events with Year-varying Reporting Rates for FLU4 Vaccine in VAERS
Jiayi Tong, Jing Huang 0021, Jingcheng Du, Cui Tao, Yong Chen 0016 |
AMIA | 3 |
| 2018 | Asthma Onset Prediction Using Structured EMR Data
Yang Xiang 0003, Jun Xu 0007, Jingcheng Du, Degui Zhi, Cui Tao |
AMIA | 3 |
| 2018 | Detecting Anatomical Entities in Clinical Text
Jun Xu 0007, Yaoyun Zhang, Jingcheng Du, Firat Tiryaki, Hua Xu 0001 |
AMIA | 3 |
| 2018 | Comparing adverse effects of Hepatitis C drugs using FAERS data
Jing Huang 0021, Xinyuan Zhang 0003, Jiayi Tong, Jingcheng Du, Rui Duan 0004, Liu Yang 0026, Jason H. Moore, Yong Chen 0016, Cui Tao |
BIBM | 4 |
| 2017 | A Twitter Study of the Relationship between the Geographic Variations of HPV Vaccination Rates and Online Information in the United States
Hansi Zhang, Christopher Wheldon, Xinsong Du, Terrell Brown, Yi Guo 0005, Stephanie Staras, Jingcheng Du, Cui Tao, Jiang Bian 0001 |
AMIA | 8 |
| 2017 | A pilot study of mining association between psychiatric stressors and symptoms in tweetsabstractSuicide is a significant public health issue, causing huge impacts on individuals as well as their families. Psychiatric stressors are major suicide risk factors and can profoundly impact a person in many aspects. In order to facilitate the understanding of psychiatric stressors and the associated symptoms, we extracted stressors and symptoms terms from major online knowledge repositories and psychiatric clinical notes. The current vocabulary collection contains 1,292 psychiatric symptoms and 715 psychiatric stressors, which was leveraged to study the associations between stressors and symptoms in a corpus of suicide related tweets. Using Chi-Square test with Bonferroni correction, 3,500 symptom-stressor pairs were identified with significant association (p-value<;0.01). Jingcheng Du, Yaoyun Zhang, Cui Tao, Hua Xu 0001 |
BIBM | 1 |
| 2016 | A representational analysis of a temporal indeterminancy display in clinical eventsabstractThis paper describes a proposition for representing temporal indeterminacy in events from clinical narratives using fuzzy sets membership functions. This approach leverages both temporal and semantic information of events and has been proved by representational analysis evaluation method. We demonstrate that membership functions' graphs can be used for representing temporal approximation and granularity of events. We also show that this approach is helpful for the construction of fine timeline of clinical events, and can be used for calculating accurate metrics for ordering events. Mohcine Madkour, Hsing-yi Song, Jingcheng Du, Cui Tao |
BIBM | 3 |