Yikuan Li

dblp:230/3520 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Deep Reinforcement Learning for Efficient and Fair Allocation of Healthcare Resources
abstract
The scarcity of health care resources, such as ventilators, often leads to the unavoidable consequence of rationing, particularly during public health emergencies or in resource-constrained settings like pandemics. The absence of a universally accepted standard for resource allocation protocols results in governments relying on varying criteria and heuristic-based approaches, often yielding suboptimal and inequitable outcomes. This study addresses the societal challenge of fair and effective critical care resource allocation by leveraging deep reinforcement learning to optimize policy decisions. We propose a transformer-based deep Q-network that integrates individual patient disease progression and interaction effects among patients to enhance allocation decisions. Our method aims to improve both fairness and overall patient outcomes. Experiments using metrics such as normalized survival rates and interracial allocation rate differences demonstrate that our approach significantly reduces excess deaths and achieves more equitable resource allocation compared to severity- and comorbidity-based protocols currently in use. Our findings highlight the potential of deep reinforcement learning to address critical health care challenges.
Yikuan Li, Chengsheng Mao, Kaixuan Huang, Hanyin Wang, Mengdi Wang 0001, Yuan Luo 0001
IJCAI1
2024 Targeted-BEHRT: Deep Learning for Observational Causal Inference on Longitudinal Electronic Health Records
abstract
Observational causal inference is useful for decision-making in medicine when randomized clinical trials (RCTs) are infeasible or nongeneralizable. However, traditional approaches do not always deliver unconfounded causal conclusions in practice. The rise of "doubly robust" nonparametric tools coupled with the growth of deep learning for capturing rich representations of multimodal data offers a unique opportunity to develop and test such models for causal inference on comprehensive electronic health records (EHRs). In this article, we investigate causal modeling of an RCT-established causal association: the effect of classes of antihypertensive on incident cancer risk. We develop a transformer-based model, targeted bidirectional EHR transformer (T-BEHRT) coupled with doubly robust estimation to estimate average risk ratio (RR). We compare our model to benchmark statistical and deep learning models for causal inference in multiple experiments on semi-synthetic derivations of our dataset with various types and intensities of confounding. In order to further test the reliability of our approach, we test our model on situations of limited data. We find that our model provides more accurate estimates of relative risk [least sum absolute error (SAE) from ground truth] compared with benchmark estimations. Finally, our model provides an estimate of class-wise antihypertensive effect on cancer risk that is consistent with results derived from RCTs.
Shishir Rao, Mohammad Mamouei, Yikuan Li, Rema Ramakrishnan, Abdelaali Hassaïne, Dexter Canoy, Kazem Rahimi
IEEE Trans. Neural Networks Learn. Syst.4
2023 Deep Reinforcement Learning for Cost-Effective Medical Diagnosis
Yikuan Li, Joseph C. Kim, Kaixuan Huang, Yuan Luo 0001, Mengdi Wang 0001
ICLR2
2023 A comparative study of pretrained language models for long clinical text
abstract
OBJECTIVE: Clinical knowledge-enriched transformer models (eg, ClinicalBERT) have state-of-the-art results on clinical natural language processing (NLP) tasks. One of the core limitations of these transformer models is the substantial memory consumption due to their full self-attention mechanism, which leads to the performance degradation in long clinical texts. To overcome this, we propose to leverage long-sequence transformer models (eg, Longformer and BigBird), which extend the maximum input sequence length from 512 to 4096, to enhance the ability to model long-term dependencies in long clinical texts. MATERIALS AND METHODS: Inspired by the success of long-sequence transformer models and the fact that clinical notes are mostly long, we introduce 2 domain-enriched language models, Clinical-Longformer and Clinical-BigBird, which are pretrained on a large-scale clinical corpus. We evaluate both language models using 10 baseline tasks including named entity recognition, question answering, natural language inference, and document classification tasks. RESULTS: The results demonstrate that Clinical-Longformer and Clinical-BigBird consistently and significantly outperform ClinicalBERT and other short-sequence transformers in all 10 downstream tasks and achieve new state-of-the-art results. DISCUSSION: Our pretrained language models provide the bedrock for clinical NLP using long texts. We have made our source code available at https://github.com/luoyuanlab/Clinical-Longformer, and the pretrained models available for public download at: https://huggingface.co/yikuan8/Clinical-Longformer. CONCLUSION: This study demonstrates that clinical knowledge-enriched long-sequence transformers are able to learn long-term dependencies in long clinical text. Our methods can also inspire the development of other domain-enriched long-sequence transformers.
Yikuan Li, Ramsey M. Wehbe, Faraz S. Ahmad, Hanyin Wang, Yuan Luo 0001
J. Am. Medical Informatics Assoc.1
2023 Patterns of diverse and changing sentiments towards COVID-19 vaccines: a sentiment analysis study integrating 11 million tweets and surveillance data across over 180 countries
abstract
OBJECTIVES: Vaccines are crucial components of pandemic responses. Over 12 billion coronavirus disease 2019 (COVID-19) vaccines were administered at the time of writing. However, public perceptions of vaccines have been complex. We integrated social media and surveillance data to unravel the evolving perceptions of COVID-19 vaccines. MATERIALS AND METHODS: Applying human-in-the-loop deep learning models, we analyzed sentiments towards COVID-19 vaccines in 11 211 672 tweets of 2 203 681 users from 2020 to 2022. The diverse sentiment patterns were juxtaposed against user demographics, public health surveillance data of over 180 countries, and worldwide event timelines. A subanalysis was performed targeting the subpopulation of pregnant people. Additional feature analyses based on user-generated content suggested possible sources of vaccine hesitancy. RESULTS: Our trained deep learning model demonstrated performances comparable to educated humans, yielding an accuracy of 0.92 in sentiment analysis against our manually curated dataset. Albeit fluctuations, sentiments were found more positive over time, followed by a subsequence upswing in population-level vaccine uptake. Distinguishable patterns were revealed among subgroups stratified by demographic variables. Encouraging news or events were detected surrounding positive sentiments crests. Sentiments in pregnancy-related tweets demonstrated a lagged pattern compared with the general population, with delayed vaccine uptake trends. Feature analysis detected hesitancies stemmed from clinical trial logics, risks and complications, and urgency of scientific evidence. DISCUSSION: Integrating social media and public health surveillance data, we associated the sentiments at individual level with observed populational-level vaccination patterns. By unraveling the distinctive patterns across subpopulations, the findings provided evidence-based strategies for improving vaccine promotion during pandemics.
Hanyin Wang, Yikuan Li, Meghan Hutch, Adrienne S. Kline, Sebastian Otero, Leena B. Mithal, Emily S. Miller, Andrew Naidech, Yuan Luo 0001
J. Am. Medical Informatics Assoc.2
2023 AD-BERT: Using pre-trained language model to predict the progression from mild cognitive impairment to Alzheimer's disease
Chengsheng Mao, Jie Xu 0012, Luke V. Rasmussen, Yikuan Li, Prakash Adekkanattu, Jennifer A. Pacheco, Borna Bonakdarpour, Robert Vassar, Li Shen 0001, Guoqian Jiang, Fei Wang 0001, Jyotishman Pathak, Yuan Luo 0001
J. Biomed. Informatics4
2023 Hi-BEHRT: Hierarchical Transformer-Based Model for Accurate Prediction of Clinical Events Using Multimodal Longitudinal Electronic Health Records
abstract
Electronic health records (EHR) represent a holistic overview of patients' trajectories. Their increasing availability has fueled new hopes to leverage them and develop accurate risk prediction models for a wide range of diseases. Given the complex interrelationships of medical records and patient outcomes, deep learning models have shown clear merits in achieving this goal. However, a key limitation of current study remains their capacity in processing long sequences, and long sequence modelling and its application in the context of healthcare and EHR remains unexplored. Capturing the whole history of medical encounters is expected to lead to more accurate predictions, but the inclusion of records collected for decades and from multiple resources can inevitably exceed the receptive field of the most existing deep learning architectures. This can result in missing crucial, long-term dependencies. To address this gap, we present Hi-BEHRT, a hierarchical Transformer-based model that can significantly expand the receptive field of Transformers and extract associations from much longer sequences. Using a multimodal large-scale linked longitudinal EHR, the Hi-BEHRT exceeds the state-of-the-art deep learning models 1% to 5% for area under the receiver operating characteristic (AUROC) curve and 1% to 8% for area under the precision recall (AUPRC) curve on average, and 2% to 8% (AUROC) and 2% to 11% (AUPRC) for patients with long medical history for 5-year heart failure, diabetes, chronic kidney disease, and stroke risk prediction. Additionally, because pretraining for hierarchical Transformer is not well-established, we provide an effective end-to-end contrastive pre-training strategy for Hi-BEHRT using EHR, improving its transferability on predicting clinical events with relatively small training dataset.
Yikuan Li, Mohammad Mamouei, Shishir Rao, Abdelaali Hassaïne, Dexter Canoy, Thomas Lukasiewicz, Kazem Rahimi
IEEE J. Biomed. Health Informatics1
2022 An Explainable Transformer-Based Deep Learning Model for the Prediction of Incident Heart Failure
abstract
Predicting the incidence of complex chronic conditions such as heart failure is challenging. Deep learning models applied to rich electronic health records may improve prediction but remain unexplainable hampering their wider use in medical practice. We aimed to develop a deep-learning framework for accurate and yet explainable prediction of 6-month incident heart failure (HF). Using 100,071 patients from longitudinal linked electronic health records across the U.K., we applied a novel Transformer-based risk model using all community and hospital diagnoses and medications contextualized within the age and calendar year for each patient's clinical encounter. Feature importance was investigated with an ablation analysis to compare model performance when alternatively removing features and by comparing the variability of temporal representations. A post-hoc perturbation technique was conducted to propagate the changes in the input to the outcome for feature contribution analyses. Our model achieved 0.93 area under the receiver operator curve and 0.69 area under the precision-recall curve on internal 5-fold cross validation and outperformed existing deep learning models. Ablation analysis indicated medication is important for predicting HF risk, calendar year is more important than chronological age, which was further reinforced by temporal variability analysis. Contribution analyses identified risk factors that are closely related to HF. Many of them were consistent with existing knowledge from clinical and epidemiological research but several new associations were revealed which had not been considered in expert-driven risk prediction models. In conclusion, the results highlight that our deep learning model, in addition high predictive performance, can inform data-driven risk factor identification.
Shishir Rao, Yikuan Li, Rema Ramakrishnan, Abdelaali Hassaïne, Dexter Canoy, John G. F. Cleland, Thomas Lukasiewicz, Kazem Rahimi
IEEE J. Biomed. Health Informatics2
2021 Early Prediction of Mortality in Critical Care Setting in Sepsis Patients Using Structured Features and Unstructured Clinical Notes
abstract
Sepsis is an important cause of mortality, especially in intensive care unit(ICU) patients. Developing novel methods to identify early mortality is critical for improving survival outcomes in sepsis patients. Using the MIMIC-III database, we integrated demographic data, physiological measurements and clinical notes. We built and applied several machine learning models to predict the risk of hospital mortality and 30-day mortality in sepsis patients. From the clinical notes, we generated clinically meaningful word representations and embeddings. Supervised learning classifiers and a deep learning architecture were used to construct prediction models. The configurations that utilized both structured and unstructured clinical features yielded competitive F-measure of 0.512. Our results showed that the approaches integrating both structured and unstructured clinical features can be effectively applied to assist clinicians in identifying the risk of mortality in sepsis patients upon admission to the ICU.
Jiyoung Shin, Yikuan Li, Yuan Luo 0001
BIBM2
2020 A Comparison of Pre-trained Vision-and-Language Models for Multimodal Representation Learning across Medical Images and Reports
abstract
Joint image-text embedding extracted from medical images and associated contextual reports is the bedrock for most biomedical vision-and-language (V+L) tasks, including medical visual question answering, clinical image-text retrieval, clinical report auto-generation. In this study, we adopt four pre-trained V+L models: LXMERT, VisualBERT, UNIER and PixelBERT to learn multimodal representation from MIMIC-CXR images and associated reports. External evaluation using the OpenI dataset shows that the joint embedding learned by pre-trained V+L models demonstrates performance improvement of 1.4% in thoracic finding classification tasks compared to a pioneering CNN+RNN model. Ablation studies are conducted to further analyze the contribution of certain model components and validate the advantage of joint embedding over text-only embedding. Attention maps are also visualized to illustrate the attention mechanism of V+L models.
Yikuan Li, Hanyin Wang, Yuan Luo 0001
BIBM1
2020 Prediction of breast cancer distant recurrence using natural language processing and knowledge-guided convolutional neural network
Hanyin Wang, Yikuan Li, Seema A. Khan, Yuan Luo 0001
Artif. Intell. Medicine2
2020 Learning multimorbidity patterns from electronic health records using Non-negative Matrix Factorisation
Abdelaali Hassaïne, Dexter Canoy, José Roberto Ayala Solares, Yajie Zhu, Shishir Rao, Yikuan Li, Mariagrazia Zottoli, Kazem Rahimi
J. Biomed. Informatics6
2018 Early Prediction of Acute Kidney Injury in Critical Care Setting Using Clinical Notes
Yikuan Li, Chengsheng Mao, Anand Srivastava, Xiaoqian Jiang, Yuan Luo 0001
BIBM1