EDBT 2026 Demo / reviewers in the wild / expert
Wenqi Shi 0002
dblp:16/4475-2
· DBLP profile ↗
25ranked-venue papers
6as first author
23since 2021 · last 2025
0000-0001-8972-7342ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 4 first-author · 18 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-PlayabstractSearch-augmented LLMs often struggle with complex reasoning tasks due to ineffective multi-hop retrieval and limited reasoning ability. We propose AceSearcher, a cooperative self-play framework that trains a single large language model (LLM) to alternate between two roles: a decomposer that breaks down complex queries and a solver that integrates retrieved contexts for answer generation. AceSearcher couples supervised fine-tuning on a diverse mixture of search, reasoning, and decomposition tasks with reinforcement fine-tuning optimized for final answer accuracy, eliminating the need for intermediate annotations. Extensive experiments on three reasoning-intensive tasks across 10 datasets show that AceSearcher outperforms state-of-the-art baselines, achieving an average exact match improvement of 7.6%. Remarkably, on document-level finance reasoning tasks, AceSearcher-32B matches the performance of the giant DeepSeek-V3 model using less than 5% of iits parameters. Even at smaller scales (1.5B and 8B), AceSearcher often surpasses existing search-augmented LLMs with up to 9× more parameters, highlighting its exceptional efficiency and effectiveness in tackling complex reasoning tasks. Ran Xu 0002, Yuchen Zhuang, Zihan Dong, Yue Yu 0001, Joyce C. Ho, Linjun Zhang, Haoyu Wang 0003, Wenqi Shi 0002, Carl Yang 0001 |
NeurIPS | 9 |
| 2025 | Novel extraction of discriminative fine-grained feature to improve retinal vessel segmentation
Shuang Zeng, Chee Hong Lee, Micky C. Nnamdi, Wenqi Shi 0002, J. Ben Tamo, Hangzhou He, May D. Wang, Lei Zhu 0012, Yanye Lu, Qiushi Ren |
Image Vis. Comput. | 4 |
| 2025 | SBDH-Reader: a large language model-powered method for extracting social and behavioral determinants of health from clinical notesabstractOBJECTIVE: Social and behavioral determinants of health (SBDH) are increasingly recognized as essential for prognostication and informing targeted interventions. Clinical notes often contain details about SBDH in unstructured format. Conventional extraction methods for these data tend to be labor intensive, inaccurate, and/or unscalable. In this study, we aim to develop and validate a large language model (LLM)-powered method to extract structured SBDH data from clinical notes through prompt engineering. MATERIALS AND METHODS: We developed SBDH-Reader to extract 6 categories of granular SBDH data by prompting GPT-4o, including employment, housing, marital status, and substance use including alcohol, tobacco, and drug use. SBDH-Reader was developed using 7225 notes from 6382 patients in the MIMIC-III database (2001-2012) and externally validated using 971 notes from 437 patients at The University of Texas Southwestern Medical Center (UTSW; 2022-2023). We evaluated SBDH-Reader's performance against human-annotated ground truths based on precision, recall, F1, and confusion matrix. RESULTS: When tested on the UTSW validation set, SBDH-Reader achieved a macro-average F1 ranging from 0.94 to 0.98 across 6 SBDH categories. For clinically relevant adverse attributes, F1 ranged from 0.96 (employment; housing) to 0.99 (tobacco use). When extracting any adverse attributes across all SBDH categories, SBDH-Reader achieved an F1 of 0.97, recall of 0.97, and precision of 0.98 in the independent validation set. DISCUSSION: SBDH-Reader demonstrated strong performance in extracting structured SBDH data through effective prompt engineering of a general-purpose LLM, without the need for task-specific fine-tuning. Its modular design and adaptability to diverse datasets and documentation patterns support its applicability in real-world clinical settings. CONCLUSION: SBDH-Reader has the potential to serve as a scalable and effective method for collecting real-time, patient-level SBDH data to support clinical research and care. Zifan Gu, Lesi He, Awais Naeem, Pui Man Chan, Asim Mohamed, Hafsa Khalil, Yujia Guo, Jingwei Huang 0002, Ismael Villanueva-Miranda, Ying Ding 0001, Wenqi Shi 0002, Matthew E. Dupre, Guanghua Xiao, Eric D. Peterson, Ann Marie Navar, Donghan M. Yang |
J. Am. Medical Informatics Assoc. | 11 |
| 2025 | Advancing Sleep Disorder Diagnostics: A Transformer-Based EEG Model for Sleep Stage Classification and OSA PredictionabstractSleep disorders, particularly Obstructive Sleep Apnea (OSA), have a considerable effect on an individual's health and quality of life. Accurate sleep stage classification and prediction of OSA are crucial for timely diagnosis and effective management of sleep disorders. In this study, we develop a sequential network that enhances sleep stage classification by incorporating self-attention mechanisms and Conditional Random Fields (CRF) into a deep learning model comprising multi-kernel Convolutional Neural Networks (CNNs) and Transformer-based encoders. The self-attention mechanism enables the model to focus on the most discriminative features extracted from single-channel electroencephalography (EEG) recordings, while the CRF module captures the temporal dependencies between sleep stages, improving the model's ability to learn more plausible sleep stage sequences. Moreover, we explore the relationship between sleep stages and OSA severity by utilizing the predicted sleep stage features to train various regression models for Apnea-Hypopnea Index (AHI) prediction. Our experiments demonstrate an improved sleep stage classification performance of 78.7%, particularly on datasets with diverse AHI values, and highlight the potential of leveraging sleep stage information for monitoring OSA. By employing advanced deep learning techniques, we thoroughly explore the intricate relationship between sleep stages and sleep apnea, laying the foundation for more precise and automated diagnostics of sleep disorders. Micky C. Nnamdi, Wenqi Shi 0002, Benjamin M. Smith, Chad Purnell, May D. Wang |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Fairness Artificial Intelligence in Clinical Decision Support: Mitigating Effect of Health DisparityabstractThe United States, as well as the global community, experiences health disparities among socially disadvantaged populations. These disparities often manifest in the data utilized for AI model training. Without appropriate de-biasing strategies, models trained to optimize predictive performance may inadvertently capture and perpetuate these inherent biases. The utilization of biased models in clinical decision-making can inflict harm upon patients from disadvantaged groups and exacerbate disparities when these decisions are documented and employed to train subsequent AI models. Unlike conventionalcorrelation-basedmethods, we aim to mitigate the negative impacts of health disparity by answering acausal inferencequestion for fairness:would the clinical decision support system make a different decision if the patient had a different sensitive attribute (e.g., race)?Recognizing the high computational complexity of developing causal models, we propose a flexible and efficient causal-model-free algorithm,CFReg, which provides causal fairness for supervised machine learning models. In addition,CFRegalso develops a novel evaluation metric to quantify fairness within clinical settings. We first validateCFRegusing a healthcare dataset of 48,784 patients focused on care management, then generalize to another four benchmark datasets with racial and ethnic disparity, including law school admission, adult income, criminal recidivism, and violent crime prediction. Experimental results demonstrate thatCFRegoutperforms baseline approaches in both fairness and accuracy, achieving a good trade-off between model fairness and supervised classification performance. Yuanda Zhu, Wenqi Shi 0002, Li Tong 0001, May D. Wang |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Quantitative Explainability Study of Deformable Convolutional Neural Networks using Chest X-raysabstractTraditional convolutional neural networks (tCNNs) often struggle with medical image analysis due to complex transformations and irregular structures with weak boundaries. Deformable convolutional neural networks (dCNNs) address these challenges through specialized modules, showing improvements in classification, segmentation, and explainability on general image tasks. However, dCNNs remain relatively unexplored in the medical domain. This study provides a unique quantitative comparison of explainability between tCNNs and dCNNs on medical image classification tasks, focusing on lung disease classification using chest X-rays (CXRs). We tested four models with varying degrees of deformability and generated saliency maps using Guided GradCAM. To quantify interpretability, we introduced a novel metric: continuous intersection over union (cIoU). While tCNNs demonstrated slightly better classification results, dCNNs showed significant improvements in explainability. We observed a positive correlation between the number of deformable layers in a model and the quality of its saliency maps. These findings highlight the potential of dCNNs in enhancing the interpretability of medical image analysis, which is crucial for real-world clinical research and practice. Vivek K. Chundru, M. Sait Kilinc, Anthony Lim, Micky C. Nnamdi, Yishan Zhong, Wenqi Shi 0002, May D. Wang |
BIBM | 6 |
| 2024 | BMRetriever: Tuning Large Language Models as Better Biomedical Text RetrieversabstractDeveloping effective biomedical retrieval models is important for excelling at knowledge-intensive biomedical tasks but still challenging due to the lack of sufficient publicly annotated biomedical data and computational resources. We present BMRetriever, a series of dense retrievers for enhancing biomedical retrieval via unsupervised pre-training on large biomedical corpora, followed by instruction fine-tuning on a combination of labeled datasets and synthetic pairs. Experiments on 5 biomedical tasks across 11 datasets verify BMRetriever's efficacy on various biomedical applications. BMRetriever also exhibits strong parameter efficiency, with the 410M variant outperforming baselines up to 11.7 times larger, and the 2B variant matching the performance of models with over 5B parameters. The training data and model checkpoints are released at https://huggingface.co/BMRetriever to ensure transparency, reproducibility, and application to new domains. Ran Xu 0002, Wenqi Shi 0002, Yue Yu 0001, Yuchen Zhuang, Yanqiao Zhu 0001, May D. Wang, Joyce C. Ho, Chao Zhang 0014, Carl Yang 0001 |
EMNLP | 2 |
| 2024 | MedAdapter: Efficient Test-Time Adaptation of Large Language Models Towards Medical ReasoningabstractDespite their improved capabilities in generation and reasoning, adapting large language models (LLMs) to the biomedical domain remains challenging due to their immense size and privacy concerns. In this study, we propose MedAdapter, a unified post-hoc adapter for test-time adaptation of LLMs towards biomedical applications. Instead of fine-tuning the entire LLM, MedAdapter effectively adapts the original model by fine-tuning only a small BERT-sized adapter to rank candidate solutions generated by LLMs. Experiments on four biomedical tasks across eight datasets demonstrate that MedAdapter effectively adapts both white-box and black-box LLMs in biomedical reasoning, achieving average performance improvements of 18.24% and 10.96%, respectively, without requiring extensive computational resources or sharing data with third parties. MedAdapter also yields enhanced performance when combined with train-time adaptation, highlighting a flexible and complementary solution to existing adaptation methods. Faced with the challenges of balancing model performance, computational resources, and data privacy, MedAdapter provides an efficient, privacy-preserving, cost-effective, and transparent solution for adapting LLMs to the biomedical domain. Wenqi Shi 0002, Ran Xu 0002, Yuchen Zhuang, Yue Yu 0001, Carl Yang 0001, May D. Wang |
EMNLP | 1 |
| 2024 | EHRAgent: Code Empowers Large Language Models for Few-shot Complex Tabular Reasoning on Electronic Health RecordsabstractClinicians often rely on data engineers to retrieve complex patient information from electronic health record (EHR) systems, a process that is both inefficient and time-consuming. We propose EHRAgent, a large language model (LLM) agent empowered with accumulative domain knowledge and robust coding capability. EHRAgent enables autonomous code generation and execution to facilitate clinicians in directly interacting with EHRs using natural language. Specifically, we formulate a multi-tabular reasoning task based on EHRs as a tool-use planning process, efficiently decomposing a complex task into a sequence of manageable actions with external toolsets. We first inject relevant medical information to enable EHRAgent to effectively reason about the given query, identifying and extracting the required records from the appropriate tables. By integrating interactive coding and execution feedback, EHRAgent then effectively learns from error messages and iteratively improves its originally generated code. Experiments on three real-world EHR datasets show that EHRAgent outperforms the strongest baseline by up to 29.6% in success rate, verifying its strong capacity to tackle complex clinical tasks with minimal demonstrations. Wenqi Shi 0002, Ran Xu 0002, Yuchen Zhuang, Yue Yu 0001, Jieyu Zhang 0001, Yuanda Zhu, Joyce C. Ho, Carl Yang 0001, May D. Wang |
EMNLP | 1 |
| 2023 | Personalized COVID-19 Early Detection Using Wearable Data Based on Self-Reported RecordsabstractThe use of wearable technology for early disease detection has gained traction as a promising avenue for improving public health outcomes. This study investigates the potential of wearable devices in early COVID-19 detection through a comprehensive methodology that integrates Long Short-Term Memory (LSTM) networks, personalized fine-tuning, and the clean-lab framework for label noise correction. Leveraging the resting heart rate (RHR) data collected from wearable devices, the study demonstrates how personalized fine-tuning improves model performance by adapting to unique physiological baselines of individuals. The incorporation of the clean-lab framework rectifies label inconsistencies arising from self-reported data inaccuracies, further enhancing model accuracy. Results indicate that the combination of personalized fine-tuning and label noise correction yields the most significant performance improvements, with heightened sensitivity, specificity, and F1 scores. The findings highlight the importance of individualized approaches and data quality control mechanisms in harnessing the potential of wearable technology for early disease detection. By addressing challenges related to variability and data quality, wearable technology may emerge as a powerful tool for timely disease identification and intervention. Jayden Myers, Wenqi Shi 0002, May D. Wang, Zewei Lei, Benoit Marteau |
BIBM | 4 |
| 2023 | Latent Topic Extraction as a Source of Labeling in Natural Language ProcessingabstractSupervised machine learning algorithms depend on accurate labeling of target data to develop models that can derive relationships between input data and the target data. One major hindrance for developing supervised machine learning models capable of predicting the correct target label of unseen data rests on the quality of the data used to train the models, which often depends on having a subject matter expert (SME) create a labeled dataset to train the model on. Given the scarcity of such experts in many fields, the time needed to analyze data for labeling, and subjective differences among experts, ways to reduce the complexity associated with creating meaningful datasets are needed. In this work, we explore the use of two unsupervised topic modeling algorithms, Latent Dirichlet Allocation (LDA) and Non-negative Matrix Factorization (NMF) as potential methods for reducing the complexities in the labeling process. Specifically, we obtained COVID patient message data labeled by a SME and compared the overlap in topics designated as COVID versus not by the two algorithms to those of the SME. For each of the topic modeling algorithms, we found a strong degree of overlap in the COVID vs. non-COVID patient message labels with that of the SME, suggesting that the methodology could be used to provide synergies for developing labeled data sets used for clinically meaningful models. Andrew Hornback, Yuanda Zhu, Monica Isgut, Wenqi Shi 0002, Arjita Nema, Blake J. Anderson, May D. Wang |
BIBM | 4 |
| 2023 | Identification of Single-Cell RNA Sequencing Molecular Signatures for COVID-19 Infection Severity ClassificationabstractIn this study, we propose a graph-based framework to identify scRNA-seq molecular signatures for COVID-19 infection severity identification. We conduct extensive experiments on scRNA-seq data from bronchoalveolar lavage fluid (BALF) with four machine learning models: Support Vector Machine, Random Forest, Graph Convolutional Network (GCN), and Graph Attention Network (GAT). In addition, we employ an explainable artificial intelligence approach, GNNExplainer, to interpret model predictions by identifying the top 15 features that contribute to the severity classification. Our finding suggests that graphical models could accurately distinguish healthy people from COVID-19 patients (F1-score > 0.9) based on patient scRNA-seq data, and traditional machine learning approaches could accurately distinguish COVID-19 patients with different severity (F1-score > 0.99), along with meaningful molecular signatures identification and discovery. Our implementation is available on a github repository: https://github.com/Da2f1eW/COVID19_Infection_Severity_Classification. Xinling Li 0003, Wei-An Chen, Wenqi Shi 0002, May D. Wang |
BIBM | 4 |
| 2023 | Development of Interpretable Machine Learning Models for COVID-19 Drug Target Docking Scores PredictionabstractWith the extensive time and financial requirements incumbent on drug discovery, computational approaches, such as protein-ligand docking predictions, are increasingly crucial for accelerating the process of drug repurposing. However, the proliferation of identified protein targets has exposed a critical knowledge gap in developing robust models that offer both generalizability and interpretability for docking score prediction. Addressing this, our study presents a machine learning-based surrogate model, employing interpretable artificial intelligence techniques for accurate docking score prediction for SARS-CoV-2 protein targets. We demonstrate the model generalization on its expansion to accommodate unseen protein targets by integrating protein target information through feature concatenation. Moreover, we leverage the SHapley Additive exPlanations (SHAP) method to identify the data-driven feature importance of molecular substructures for knowledge-based validation. Our experiments reveal that the combination of data-driven prediction and knowledge-driven validation could provide biomedical insights into the interactions between drugs and SARS-CoV-2 proteins, elucidating their consequent effects on docking scores. Wenqi Shi 0002, Mio Murakoso, May D. Wang |
BIBM | 1 |
| 2023 | Effective Surrogate Models for Docking Scores Prediction of Candidate Drug Molecules on SARS-CoV-2 Protein TargetsabstractEmerging infectious diseases, such as coronavirus disease 2019 (COVID-19), pose a major threat to public health and present a critical challenge for drug discovery. Due to the cost- and time-consuming process of new drug development, virtual pre-screening methods such as protein-ligand docking prediction have become essential tools in enhancing drug refurbishment and repurposing. In this study, we propose a machine learning-based surrogate model for docking score prediction of drug candidates on SARS-CoV-2 protein targets via deep feature concatenation. We investigate 14 different combinations of rule-based and data-driven fingerprinting methods to identify the optimal representation of candidate drug molecules. Extensive experiments on docking scores of 270,000 molecules across 18 different SARS-CoV-2 protein targets demonstrate the effectiveness of the proposed surrogate models. In addition to unseen drugs, we further investigate the generalization of the proposed framework for unseen protein targets. This study may provide an instrumental and generalizable framework for exploring ligand-protein interaction, serving as a useful tool to facilitate rapid drug pre-screening during emerging public health crises. Wenqi Shi 0002, Mio Murakoso, Linxi Xiong, Matthew Chen, May D. Wang |
BIBM | 1 |
| 2023 | Multi-Modal Deep Feature Integration for Alzheimer's Disease StagingabstractAlzheimer's disease (AD) is one of the leading causes of dementia and 7th leading cause of death in the United States. The provisional diagnosis of AD relies on comprehensive examinations, including medical history, neurological and psychiatric examinations, cognitive assessments, and neuroimaging studies. Integrating diverse sets of clinical data, including electronic health records (EHRs), medical imaging, and genomic data, enables a holistic view of AD staging analysis. In this study, we propose an end-to-end deep learning architecture to jointly learn from magnetic resonance imaging (MRI), positron emission tomography (PET), EHRs, and genomics data to classify patients into AD, mild cognitive disorders, and controls. We conduct extensive experiments to explore different feature-level and intermediate-level fusion methods. Our findings suggest intermediate multiplicative fusion achieves the best stage prediction performance on the external validation dataset. Compared with unimodal baselines, we can observe that integrative approaches that leverage all four modalities demonstrate superior performance to baselines reliant solely on one or two modalities. In an age-wise comparison, we observe a unique pattern that all fusion methods exhibited superior performance in the earlier age brackets (50–70 years), with performance diminishing as the age group advanced (70–90 years). The proposed integration framework has the potential to augment our understanding of disease diagnosis and progression by leveraging complementary information from multimodal patient data. Wenqi Shi 0002, May D. Wang |
BIBM | 2 |
| 2023 | Uncertainty-Aware Ensemble Learning Models for Out-of-Distribution Medical Imaging AnalysisabstractAdvanced deep-learning techniques have been employed to develop clinical decision support systems for diagnosis and prognosis using medical images. However, the presence of out-of-distribution (OOD) samples, which deviate from the training data distribution, poses a significant challenge. Accurate quantification of the predictive uncertainty is crucial for ensuring reliable and dependable implementation in medical settings as a clinical decision support system. In this work, we propose an ensemble model to derive predictive uncertainty estimates for uncertainty quantification on OOD medical imaging. Specifically, the models are initialized with ImageNet pre-trained weights and fine-tuned on chest Computed Tomography (CT). Moreover, we utilize Grad-CAM to visualize and interpret the areas of the image that contribute most to the model’s predictions and uncertainty estimates. This visualization technique enhances the in-terpretability of our ensemble model and supports more informed clinical decision-making. Through extensive experiments on three Chest CT datasets, we have demonstrated the effectiveness of our approach in estimating uncertainty under domain shifting. Our results provide valuable insights into the reliability and specificity of deep ensemble uncertainty predictions in medical image analysis. Our Uncertainty-Aware Ensemble (UAE) approach can enable reliable and transparent predictions for safety-critical medical applications. J. Ben Tamo, Micky C. Nnamdi, Lea Lesbats, Wenqi Shi 0002, Yishan Zhong, May D. Wang |
BIBM | 4 |
| 2023 | Explainable synthetic image generation to improve risk assessment of rare pediatric heart transplant rejectionabstractExpert microscopic analysis of cells obtained from frequent heart biopsies is vital for early detection of pediatric heart transplant rejection to prevent heart failure. Detection of this rare condition is prone to low levels of expert agreement due to the difficulty of identifying subtle rejection signs within biopsy samples. The rarity of pediatric heart transplant rejection also means that very few gold-standard images are available for developing machine learning models. To solve this urgent clinical challenge, we developed a deep learning model to automatically quantify rejection risk within digital images of biopsied tissue using an explainable synthetic data augmentation approach. We developed this explainable AI framework to illustrate how our progressive and inspirational generative adversarial network models distinguish between normal tissue images and those containing cellular rejection signs. To quantify biopsy-level rejection risk, we first detect local rejection features using a binary image classifier trained with expert-annotated and synthetic examples. We converted these local predictions into a biopsy-wide rejection score via an interpretable histogram-based approach. Our model significantly improves upon prior works with the same dataset with an area under the receiver operating curve (AUROC) of 98.84% for the local rejection detection task and 95.56% for the biopsy-rejection prediction task. A biopsy-level sensitivity of 83.33% makes our approach suitable for early screening of biopsies to prioritize expert analysis. Our framework provides a solution to rare medical imaging challenges currently limited by small datasets. Felipe O. Giuste, Ryan Sequeira, Vikranth Keerthipati, Peter Lais, Ali Mirzazadeh, Arshawn Mohseni, Yuanda Zhu, Wenqi Shi 0002, Benoit Marteau, Yishan Zhong, Li Tong 0001, Bibhuti Das 0002, Bahig M. Shehata, Shriprasad R. Deshpande, May D. Wang |
J. Biomed. Informatics | 8 |
| 2022 | Attention-based Automated Chest CT Image Segmentation Method of COVID-19 Lung InfectionabstractAccording to the World Health Organization, Artificial Intelligence (AI) technology may assist in COVID-19 management. However, existing image segmentation using AI suffers from a lack of accuracy and explainability, which prevents its adoption in actual clinical practice. In this paper, we investigated an attention-based image segmentation method for COVID-19 CT imaging with enhanced interpretation capabilities. Specifically, we developed U-Net architecture-based for segmentation with attention coefficients to produce a salient feature map. We use the DICE score and accuracy to perform a comprehensive model evaluation. We compared to other well-known methods such as Light U-Net, COPLE-Net, and Res U-Net and demonstrated that attention U-Net is superior for COVID-19 segmentation tasks in terms of performance and explainability. We also developed the tool as a web-application with a graphic user interface with the goal to translate this AI-driven clinical decision-support system for real-world clinical use. Beom J. Lee, Sarkis T. Martirosyan, Zaid Khan 0003, Han Y. Chiu, Wenqi Shi 0002, Felipe O. Giuste, Yishan Zhong, Jimin Sun, May D. Wang |
BIBE | 6 |
| 2022 | Interpretable Evaluation of Diabetic Retinopathy Grade Regarding Eye Color Fundus ImagesabstractThis paper reports an interpretable automated grading system for diabetic retinopathy using color fundus images. First, we develop shallow learners as baselines. Second, we pre-train deep neural networks to extract high-dimensional features and complex patterns from fundus images and utilize ensemble models to do automatic grading. Then we develop several explainable artificial intelligence models to visualize the extracted deep features and to interpret the predicted outcomes. We investigate the robustness of our system over two publicly available diabetic retinopathy fundus imaging datasets. In addition, we displayed both local and global explainable results to further illustrate the clinical decision-making process with deep models. The innovations of our work include (1) using ensemble models to boost the performance of diabetic retinopathy grading system, and (2) providing transparency of ensemble models using explainable artificial intelligence. The result has shown the potential to improve the effectiveness and accessibility of diabetic retinopathy screening in clinical practice and research settings. Jieh Sheng Hsu, Noaima Bari, Xu Qiu, Malvika Viswanathan, Wenqi Shi 0002, Felipe O. Giuste, Yishan Zhong, Jimin Sun, May D. Wang |
BIBE | 6 |
| 2022 | Multi-Modal Deep Learning Models for Alzheimer's Disease Prediction Using MRI and EHRabstractAlzheimer's Disease (AD) is an irreversible and progressive neurodegenerative disorder with three stages: cognitively normal (CN), mild cognitive impairment (MCI), and clinical dementia. Progression and stage prediction of dementia plays an important role in prognosis and treatment. In this work, we developed a multi-modal AD progress prediction model that integrates magnetic resonance imaging (MRI) and electronic health record (EHR) to classify patients into three stages: CN, MCI, and AD. We trained deep auto-encoder to extract features from EHR data, and ResNet and 3D U-Net for MRI imaging data. We developed an entropy-based weighted sum classification method to integrate the classification results from each individual modality to generate final prediction. We experimented on Alzheimer's Disease Neuroimaging Initiative (ADNI) data to demonstrate that the multi-modality integration model outperforms single modality models in accuracy, precision, recall, and F1 scores. In addition, our model achieves competitive performance in comparison with other state-of-the-art multi-modality integration methods on AD progression prediction. Sathvik S. Prabhu, John A. Berkebile, Neha Rajagopalan, Renjie Yao, Wenqi Shi 0002, Felipe O. Giuste, Yishan Zhong, Jimin Sun, May D. Wang |
BIBE | 5 |
| 2022 | Development of Machine Learning Regression Model for COVID-19 Drug Target PredictionabstractThere is a perennial need to identify novel, effective therapeutic agents to combat rising infections. Recently, prediction of therapeutic targets to decrease the impact of COVID-19 has posed an urgent challenge requiring innovative solutions. Successful identification of novel drug-target combinations may greatly facilitate drug development. To meet this need, we developed a COVID-19 drug target prediction model using machine learning approaches to quickly identify drug candidates for 18 COVID-19 protein targets. Specifically, we analyzed the performance of three prediction models to predict drug-target docking scores, which represents the strength of interactions between ligands and proteins. Docking scores were predicted for 300,457 molecules on 18 different COVID-19 related protein docking targets. Our proposed approach achieved a competitive performance with $\mathrm{R}^{2}$=0.69,MAE=0.285, MSE=0.627. In addition, we identify chemical structures associated with stronger binding affinities across target binding sites. We believe our work could potentially save pharmaceutical companies significant resources, especially during the early stages of drug development. Alexandra Zamitalo, Qingtong Xie, Mayar Allam, Phinu Philip, Wenqi Shi 0002, Felipe O. Giuste, Benoit Marteau, Mio Murakoso, May D. Wang |
BIBM | 5 |
| 2021 | A FHIR-compliant Application for Multi-Site and Multi-Modality Pediatric Scoliosis Patient RehabilitationabstractScoliosis is a spinal curvature that most frequently affects adolescents. Posterior spinal fusion surgery is required to correct the deformity in patients with severe scoliosis. Surgeons frequently use radiographic measurements and patient reported outcomes to aid in surgical treatment and monitor patient rehabilitation. Shriners Hospitals for Children is a large healthcare system caring for a significant percentage of pediatric patients with scoliosis. Surgeons from SHC-Greenville and SHC-Lexington have recorded data from more than 1,000 individual scoliosis patients. However, these collected data are usually dispersed across individual healthcare sites, necessitating the development of an integrated clinical data repository for data sharing and management. In this paper, we established a standardized research data repository with FHIR resources to harmonize multi-modal patient data from multiple clinical sites. Additionally, a FHIR-compliant application with a web-based user interface was prototyped to enable clinicians and researchers to access scoliosis patient data within our integrated and standardized research repository. Patient cohort definitions can be used to search these records using the same FHIR application. This standardized data-sharing framework and healthcare information system can be applied to multi-site and multimodality studies for clinical and research purposes, with the ultimate goal of improving the quality of patient care. Wenqi Shi 0002, Felipe O. Giuste, Yuanda Zhu, Ashley M. Carpenter, Henry J. Iwinski, Coleman Hilton, J. Michael Wattenbarger, May D. Wang |
BIBM | 1 |
| 2021 | COVID-19 Automatic Diagnosis With Radiographic Imaging: Explainable Attention Transfer Deep Neural NetworksabstractResearchers seek help from deep learning methods to alleviate the enormous burden of reading radiological images by clinicians during the COVID-19 pandemic. However, clinicians are often reluctant to trust deep models due to their black-box characteristics. To automatically differentiate COVID-19 and community-acquired pneumonia from healthy lungs in radiographic imaging, we propose an explainable attention-transfer classification model based on the knowledge distillation network structure. The attention transfer direction always goes from the teacher network to the student network. Firstly, the teacher network extracts global features and concentrates on the infection regions to generate attention maps. It uses a deformable attention module to strengthen the response of infection regions and to suppress noise in irrelevant regions with an expanded reception field. Secondly, an image fusion module combines attention knowledge transferred from teacher network to student network with the essential information in original input. While the teacher network focuses on global features, the student branch focuses on irregularly shaped lesion regions to learn discriminative features. Lastly, we conduct extensive experiments on public chest X-ray and CT datasets to demonstrate the explainability of the proposed architecture in diagnosing COVID-19. Wenqi Shi 0002, Li Tong 0001, Yuanda Zhu, May D. Wang |
IEEE J. Biomed. Health Informatics | 1 |
| 2018 | Integrating Sparse Reconstruction Saliency and Target-Aware Active Contour Model for Airport ExtractionabstractThis paper deals with automatic airport extraction in remote sensing images (RSIs). We present an innovative framework using sparse reconstruction saliency (SRS) and target-aware active contour model (TAACM). We begin with segmenting an image into superpixels and extracting the feature vectors. In feature space, we learn an airport target dictionary and a background dictionary for sparse representation of all image sub-regions. The saliency confidence can be determined by sparse reconstruction error. Based on the saliency map, we apply a novel target-aware active contour model (TAACM) for target contour tracking and provide accurate descriptions about the airport details. Extensive experiments demonstrate that the SRS algorithm outperforms nine competing saliency models in remote sensing scenes. In addition, the proposed airport extraction framework achieves higher detection accuracy compared with three competing methods. Qijian Zhang, Wenqi Shi 0002, Libao Zhang |
ICIP | 2 |
| 2018 | Airport Extraction via Complementary Saliency Analysis and Saliency-Oriented Active Contour ModelabstractAutomatic airport extraction in remote sensing images (RSIs) has been widely applied in military and civil applications. An efficient airport extraction framework for RSIs is constructed in this letter. In the first step, we put forward a two-way complementary saliency analysis (CSA) scheme that combines vision-oriented saliency and knowledge-oriented saliency for the airport position estimation. In the second step, we construct a saliency-oriented active contour model (SOACM) for airport contour tracking, where a saliency orientation term is incorporated into the level-set-based energy functions. Under the guidance of saliency feature representations obtained by CSA, the SOACM can acquire well-defined and highly precise object contours. Experimental results demonstrate that the proposed extraction framework shows good adaptability in remote sensing scenes, and uniformly achieves high detection rate and low false alarm rate. Compared with three state-of-the-art algorithms, our proposal can not only estimate the location of airport targets, but also extract detailed information of the airport contours. Qijian Zhang, Libao Zhang, Wenqi Shi 0002 |
IEEE Geosci. Remote. Sens. Lett. | 3 |