Qiang Wei 0002

dblp:50/5190-2 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
4since 2021 · last 2024
0000-0002-8665-0201ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Ensemble pretrained language models to extract biomedical knowledge from literature
abstract
OBJECTIVES: The rapid expansion of biomedical literature necessitates automated techniques to discern relationships between biomedical concepts from extensive free text. Such techniques facilitate the development of detailed knowledge bases and highlight research deficiencies. The LitCoin Natural Language Processing (NLP) challenge, organized by the National Center for Advancing Translational Science, aims to evaluate such potential and provides a manually annotated corpus for methodology development and benchmarking. MATERIALS AND METHODS: For the named entity recognition (NER) task, we utilized ensemble learning to merge predictions from three domain-specific models, namely BioBERT, PubMedBERT, and BioM-ELECTRA, devised a rule-driven detection method for cell line and taxonomy names and annotated 70 more abstracts as additional corpus. We further finetuned the T0pp model, with 11 billion parameters, to boost the performance on relation extraction and leveraged entites' location information (eg, title, background) to enhance novelty prediction performance in relation extraction (RE). RESULTS: Our pioneering NLP system designed for this challenge secured first place in Phase I-NER and second place in Phase II-relation extraction and novelty prediction, outpacing over 200 teams. We tested OpenAI ChatGPT 3.5 and ChatGPT 4 in a Zero-Shot setting using the same test set, revealing that our finetuned model considerably surpasses these broad-spectrum large language models. DISCUSSION AND CONCLUSION: Our outcomes depict a robust NLP system excelling in NER and RE across various biomedical entities, emphasizing that task-specific models remain superior to generic large ones. Such insights are valuable for endeavors like knowledge graph development and hypothesis formulation in biomedical research.
Qiang Wei 0002, Liang-Chin Huang, Jianfu Li, Yao-Shun Chuang, Jianping He 0002, Avisha Das, Vipina Kuttichi Keloth, Yuntao Yang, Chiamaka S. Diala, Kirk Roberts, Cui Tao, Xiaoqian Jiang, W. Jim Zheng, Hua Xu 0001
J. Am. Medical Informatics Assoc.2
2024 Improving tabular data extraction in scanned laboratory reports using deep learning models
Qiang Wei 0002, Xinghan Chen, Jianfu Li, Cui Tao, Hua Xu 0001
J. Biomed. Informatics2
2022 Fast Textual Corpus Retrieval from Electronic Health Records using PDF Parser and TxtSplit Tools
Qiang Wei 0002, Yujia Zhou 0003, Hua Xu 0001
AMIA2
2022 ClinicalLayoutLM: A Pre-trained Multi-modal Model for Understanding Scanned Document in Electronic Health Records
abstract
Scanned documents (e.g., faxes) are still widely used in clinical practice and are prevalent in Electronic Health Records (EHR). Unlocking information in scanned documents in EHRs is critical for clinical operation and research. However, it is challenging as it requires converting images to texts before applying information extraction technologies. Here we propose a multi-modal approach (ClinicalLayoutLM) that jointly models text extracted from Optical Character Recognition (OCR) and layout/image information to classify scanned clinical documents into different categories (e.g., lab reports and CT scans). Using a clinical corpus of 348, 311 scanned documents, we continually pretrained ClinicalLayoutLM based on LayoutLMv3, a multi-modal model from the open domain. For the task to classify the scanned clinical documents into 16 categories, ClinicalLayoutLM achieved an F1 score of 0.9051, which outperformed the baseline model (0.8840) that was based on text from OCR only. ClinicalLayoutLM is the first of its kind of multi-modal models for clinical documents and we believe it could benefit other clinical natural language processing (NLP) tasks such as layout analysis, information extraction and so on. The code is available at https://github.com/UTHealth-CCB/ClinicalLayoutLM and the pre-trained model is available upon request.
Qiang Wei 0002, Xu Zuo, Omer Anjum, Ryan Denlinger, Elmer V. Bernstam, Martin J. Citardi, Hua Xu 0001
IEEE Big Data1
2020 A study of deep learning approaches for medication and adverse drug event extraction from clinical text
abstract
OBJECTIVE: This article presents our approaches to extraction of medications and associated adverse drug events (ADEs) from clinical documents, which is the second track of the 2018 National NLP Clinical Challenges (n2c2) shared task. MATERIALS AND METHODS: The clinical corpus used in this study was from the MIMIC-III database and the organizers annotated 303 documents for training and 202 for testing. Our system consists of 2 components: a named entity recognition (NER) and a relation classification (RC) component. For each component, we implemented deep learning-based approaches (eg, BI-LSTM-CRF) and compared them with traditional machine learning approaches, namely, conditional random fields for NER and support vector machines for RC, respectively. In addition, we developed a deep learning-based joint model that recognizes ADEs and their relations to medications in 1 step using a sequence labeling approach. To further improve the performance, we also investigated different ensemble approaches to generating optimal performance by combining outputs from multiple approaches. RESULTS: Our best-performing systems achieved F1 scores of 93.45% for NER, 96.30% for RC, and 89.05% for end-to-end evaluation, which ranked #2, #1, and #1 among all participants, respectively. Additional evaluations show that the deep learning-based approaches did outperform traditional machine learning algorithms in both NER and RC. The joint model that simultaneously recognizes ADEs and their relations to medications also achieved the best performance on RC, indicating its promise for relation extraction. CONCLUSION: In this study, we developed deep learning approaches for extracting medications and their attributes such as ADEs, and demonstrated its superior performance compared with traditional machine learning algorithms, indicating its uses in broader NER and RC tasks in the medical domain.
Qiang Wei 0002, Zongcheng Ji, Zhiheng Li 0004, Jingcheng Du, Jun Xu 0007, Yang Xiang 0003, Firat Tiryaki, Stephen Wu 0004, Yaoyun Zhang, Cui Tao, Hua Xu 0001
J. Am. Medical Informatics Assoc.1
2020 Deep learning in clinical natural language processing: a methodical review
abstract
OBJECTIVE: This article methodically reviews the literature on deep learning (DL) for natural language processing (NLP) in the clinical domain, providing quantitative analysis to answer 3 research questions concerning methods, scope, and context of current research. MATERIALS AND METHODS: We searched MEDLINE, EMBASE, Scopus, the Association for Computing Machinery Digital Library, and the Association for Computational Linguistics Anthology for articles using DL-based approaches to NLP problems in electronic health records. After screening 1,737 articles, we collected data on 25 variables across 212 papers. RESULTS: DL in clinical NLP publications more than doubled each year, through 2018. Recurrent neural networks (60.8%) and word2vec embeddings (74.1%) were the most popular methods; the information extraction tasks of text classification, named entity recognition, and relation extraction were dominant (89.2%). However, there was a "long tail" of other methods and specific tasks. Most contributions were methodological variants or applications, but 20.8% were new methods of some kind. The earliest adopters were in the NLP community, but the medical informatics community was the most prolific. DISCUSSION: Our analysis shows growing acceptance of deep learning as a baseline for NLP research, and of DL-based NLP in the medical community. A number of common associations were substantiated (eg, the preference of recurrent neural networks for sequence-labeling named entity recognition), while others were surprisingly nuanced (eg, the scarcity of French language clinical NLP with deep learning). CONCLUSION: Deep learning has not yet fully penetrated clinical NLP and is growing rapidly. This review highlighted both the popular and unique trends in this active field.
Stephen Wu 0004, Kirk Roberts, Surabhi Datta, Jingcheng Du, Zongcheng Ji, Yuqi Si, Sarvesh Soni, Qiang Wei 0002, Yang Xiang 0003, Bo Zhao 0001, Hua Xu 0001
J. Am. Medical Informatics Assoc.9
2019 Relation Extraction from Clinical Narratives Using Pre-trained Language Models
Qiang Wei 0002, Zongcheng Ji, Yuqi Si, Jingcheng Du, Firat Tiryaki, Stephen Wu 0004, Cui Tao, Kirk Roberts, Hua Xu 0001
AMIA1
2019 Cost-aware active learning for named entity recognition in clinical text
abstract
OBJECTIVE: Active Learning (AL) attempts to reduce annotation cost (ie, time) by selecting the most informative examples for annotation. Most approaches tacitly (and unrealistically) assume that the cost for annotating each sample is identical. This study introduces a cost-aware AL method, which simultaneously models both the annotation cost and the informativeness of the samples and evaluates both via simulation and user studies. MATERIALS AND METHODS: We designed a novel, cost-aware AL algorithm (Cost-CAUSE) for annotating clinical named entities; we first utilized lexical and syntactic features to estimate annotation cost, then we incorporated this cost measure into an existing AL algorithm. Using the 2010 i2b2/VA data set, we then conducted a simulation study comparing Cost-CAUSE with noncost-aware AL methods, and a user study comparing Cost-CAUSE with passive learning. RESULTS: Our cost model fit empirical annotation data well, and Cost-CAUSE increased the simulation area under the learning curve (ALC) scores by up to 5.6% and 4.9%, compared with random sampling and alternate AL methods. Moreover, in a user annotation task, Cost-CAUSE outperformed passive learning on the ALC score and reduced annotation time by 20.5%-30.2%. DISCUSSION: Although AL has proven effective in simulations, our user study shows that a real-world environment is far more complex. Other factors have a noticeable effect on the AL method, such as the annotation accuracy of users, the tiredness of users, and even the physical and mental condition of users. CONCLUSION: Cost-CAUSE saves significant annotation cost compared to random sampling.
Qiang Wei 0002, Yukun Chen 0001, Mandana Salimi, Joshua C. Denny, Qiaozhu Mei, Thomas A. Lasko, Qingxia Chen, Stephen Wu 0004, Amy Franklin, Trevor Cohen, Hua Xu 0001
J. Am. Medical Informatics Assoc.1
2018 Clinical text annotation - what factors are associated with the cost of time?
Qiang Wei 0002, Amy Franklin, Trevor Cohen, Hua Xu 0001
AMIA1
2018 CLAMP-PA: A machine learning based pre-annotation pipeline for corpus construction of clinical concepts
Yaoyun Zhang, Qiang Wei 0002, Hua Xu 0001
AMIA2
2017 Detecting Contradictory and Consistent Citations in Biomedical Literature
Jun Xu 0007, Yonghui Wu 0001, Yaoyun Zhang, Qiang Wei 0002, Hua Xu 0001
AMIA4
2016 A Study of Active Learning for Document Selection in Clinical Named Entity Recognition
Qiang Wei 0002, Yukun Chen 0001, Sungrim Moon, Trevor Cohen, Hua Xu 0001
AMIA1