VLDB 2026 Research / reviewers in the wild / expert
May D. Wang
dblp:93/2701 · also May Dongmei Wang
· DBLP profile ↗
83ranked-venue papers
0as first author
26since 2021 · last 2026
0000-0003-3961-3608ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 71 · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7Artificial intelligence and machine learning · 6 · 5 since 2021Software engineering, systems software and programming languages · 6Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MetaBench: A Multi-task Benchmark for Assessing LLMs in MetabolomicsabstractYuxing Lu, Xukai Zhao, J. Ben Tamo, Micky C. Nnamdi, Rui Peng, Shuang Zeng, Xingyu Hu, Jinzhuo Wang, May Dongmei Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuxing Lu, Xukai Zhao, J. Ben Tamo, Micky C. Nnamdi, Rui Peng 0006, Shuang Zeng, Jinzhuo Wang, May D. Wang |
ACL (1) | 9 |
| 2025 | Novel extraction of discriminative fine-grained feature to improve retinal vessel segmentation
Shuang Zeng, Chee Hong Lee, Micky C. Nnamdi, Wenqi Shi 0002, J. Ben Tamo, Hangzhou He, May D. Wang, Lei Zhu 0012, Yanye Lu, Qiushi Ren |
Image Vis. Comput. | 9 |
| 2025 | Guest Editorial: Precision Health: AI Tailored to Individuals
Edward Sazonov, Bobak Mortazavi, Tayo Obafemi-Ajayi, Hassan Ghasemzadeh 0001, María Fernanda Cabrera-Umpiérrez, May D. Wang |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Advancing Sleep Disorder Diagnostics: A Transformer-Based EEG Model for Sleep Stage Classification and OSA PredictionabstractSleep disorders, particularly Obstructive Sleep Apnea (OSA), have a considerable effect on an individual's health and quality of life. Accurate sleep stage classification and prediction of OSA are crucial for timely diagnosis and effective management of sleep disorders. In this study, we develop a sequential network that enhances sleep stage classification by incorporating self-attention mechanisms and Conditional Random Fields (CRF) into a deep learning model comprising multi-kernel Convolutional Neural Networks (CNNs) and Transformer-based encoders. The self-attention mechanism enables the model to focus on the most discriminative features extracted from single-channel electroencephalography (EEG) recordings, while the CRF module captures the temporal dependencies between sleep stages, improving the model's ability to learn more plausible sleep stage sequences. Moreover, we explore the relationship between sleep stages and OSA severity by utilizing the predicted sleep stage features to train various regression models for Apnea-Hypopnea Index (AHI) prediction. Our experiments demonstrate an improved sleep stage classification performance of 78.7%, particularly on datasets with diverse AHI values, and highlight the potential of leveraging sleep stage information for monitoring OSA. By employing advanced deep learning techniques, we thoroughly explore the intricate relationship between sleep stages and sleep apnea, laying the foundation for more precise and automated diagnostics of sleep disorders. Micky C. Nnamdi, Wenqi Shi 0002, Benjamin M. Smith, Chad Purnell, May D. Wang |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Fairness Artificial Intelligence in Clinical Decision Support: Mitigating Effect of Health DisparityabstractThe United States, as well as the global community, experiences health disparities among socially disadvantaged populations. These disparities often manifest in the data utilized for AI model training. Without appropriate de-biasing strategies, models trained to optimize predictive performance may inadvertently capture and perpetuate these inherent biases. The utilization of biased models in clinical decision-making can inflict harm upon patients from disadvantaged groups and exacerbate disparities when these decisions are documented and employed to train subsequent AI models. Unlike conventionalcorrelation-basedmethods, we aim to mitigate the negative impacts of health disparity by answering acausal inferencequestion for fairness:would the clinical decision support system make a different decision if the patient had a different sensitive attribute (e.g., race)?Recognizing the high computational complexity of developing causal models, we propose a flexible and efficient causal-model-free algorithm,CFReg, which provides causal fairness for supervised machine learning models. In addition,CFRegalso develops a novel evaluation metric to quantify fairness within clinical settings. We first validateCFRegusing a healthcare dataset of 48,784 patients focused on care management, then generalize to another four benchmark datasets with racial and ethnic disparity, including law school admission, adult income, criminal recidivism, and violent crime prediction. Experimental results demonstrate thatCFRegoutperforms baseline approaches in both fairness and accuracy, achieving a good trade-off between model fairness and supervised classification performance. Yuanda Zhu, Wenqi Shi 0002, Li Tong 0001, May D. Wang |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Quantitative Explainability Study of Deformable Convolutional Neural Networks using Chest X-raysabstractTraditional convolutional neural networks (tCNNs) often struggle with medical image analysis due to complex transformations and irregular structures with weak boundaries. Deformable convolutional neural networks (dCNNs) address these challenges through specialized modules, showing improvements in classification, segmentation, and explainability on general image tasks. However, dCNNs remain relatively unexplored in the medical domain. This study provides a unique quantitative comparison of explainability between tCNNs and dCNNs on medical image classification tasks, focusing on lung disease classification using chest X-rays (CXRs). We tested four models with varying degrees of deformability and generated saliency maps using Guided GradCAM. To quantify interpretability, we introduced a novel metric: continuous intersection over union (cIoU). While tCNNs demonstrated slightly better classification results, dCNNs showed significant improvements in explainability. We observed a positive correlation between the number of deformable layers in a model and the quality of its saliency maps. These findings highlight the potential of dCNNs in enhancing the interpretability of medical image analysis, which is crucial for real-world clinical research and practice. Vivek K. Chundru, M. Sait Kilinc, Anthony Lim, Micky C. Nnamdi, Yishan Zhong, Wenqi Shi 0002, May D. Wang |
BIBM | 7 |
| 2024 | BMRetriever: Tuning Large Language Models as Better Biomedical Text RetrieversabstractDeveloping effective biomedical retrieval models is important for excelling at knowledge-intensive biomedical tasks but still challenging due to the lack of sufficient publicly annotated biomedical data and computational resources. We present BMRetriever, a series of dense retrievers for enhancing biomedical retrieval via unsupervised pre-training on large biomedical corpora, followed by instruction fine-tuning on a combination of labeled datasets and synthetic pairs. Experiments on 5 biomedical tasks across 11 datasets verify BMRetriever's efficacy on various biomedical applications. BMRetriever also exhibits strong parameter efficiency, with the 410M variant outperforming baselines up to 11.7 times larger, and the 2B variant matching the performance of models with over 5B parameters. The training data and model checkpoints are released at https://huggingface.co/BMRetriever to ensure transparency, reproducibility, and application to new domains. Ran Xu 0002, Wenqi Shi 0002, Yue Yu 0001, Yuchen Zhuang, Yanqiao Zhu 0001, May D. Wang, Joyce C. Ho, Chao Zhang 0014, Carl Yang 0001 |
EMNLP | 6 |
| 2024 | MedAdapter: Efficient Test-Time Adaptation of Large Language Models Towards Medical ReasoningabstractDespite their improved capabilities in generation and reasoning, adapting large language models (LLMs) to the biomedical domain remains challenging due to their immense size and privacy concerns. In this study, we propose MedAdapter, a unified post-hoc adapter for test-time adaptation of LLMs towards biomedical applications. Instead of fine-tuning the entire LLM, MedAdapter effectively adapts the original model by fine-tuning only a small BERT-sized adapter to rank candidate solutions generated by LLMs. Experiments on four biomedical tasks across eight datasets demonstrate that MedAdapter effectively adapts both white-box and black-box LLMs in biomedical reasoning, achieving average performance improvements of 18.24% and 10.96%, respectively, without requiring extensive computational resources or sharing data with third parties. MedAdapter also yields enhanced performance when combined with train-time adaptation, highlighting a flexible and complementary solution to existing adaptation methods. Faced with the challenges of balancing model performance, computational resources, and data privacy, MedAdapter provides an efficient, privacy-preserving, cost-effective, and transparent solution for adapting LLMs to the biomedical domain. Wenqi Shi 0002, Ran Xu 0002, Yuchen Zhuang, Yue Yu 0001, Carl Yang 0001, May D. Wang |
EMNLP | 8 |
| 2024 | EHRAgent: Code Empowers Large Language Models for Few-shot Complex Tabular Reasoning on Electronic Health RecordsabstractClinicians often rely on data engineers to retrieve complex patient information from electronic health record (EHR) systems, a process that is both inefficient and time-consuming. We propose EHRAgent, a large language model (LLM) agent empowered with accumulative domain knowledge and robust coding capability. EHRAgent enables autonomous code generation and execution to facilitate clinicians in directly interacting with EHRs using natural language. Specifically, we formulate a multi-tabular reasoning task based on EHRs as a tool-use planning process, efficiently decomposing a complex task into a sequence of manageable actions with external toolsets. We first inject relevant medical information to enable EHRAgent to effectively reason about the given query, identifying and extracting the required records from the appropriate tables. By integrating interactive coding and execution feedback, EHRAgent then effectively learns from error messages and iteratively improves its originally generated code. Experiments on three real-world EHR datasets show that EHRAgent outperforms the strongest baseline by up to 29.6% in success rate, verifying its strong capacity to tackle complex clinical tasks with minimal demonstrations. Wenqi Shi 0002, Ran Xu 0002, Yuchen Zhuang, Yue Yu 0001, Jieyu Zhang 0001, Yuanda Zhu, Joyce C. Ho, Carl Yang 0001, May D. Wang |
EMNLP | 10 |
| 2023 | Personalized COVID-19 Early Detection Using Wearable Data Based on Self-Reported RecordsabstractThe use of wearable technology for early disease detection has gained traction as a promising avenue for improving public health outcomes. This study investigates the potential of wearable devices in early COVID-19 detection through a comprehensive methodology that integrates Long Short-Term Memory (LSTM) networks, personalized fine-tuning, and the clean-lab framework for label noise correction. Leveraging the resting heart rate (RHR) data collected from wearable devices, the study demonstrates how personalized fine-tuning improves model performance by adapting to unique physiological baselines of individuals. The incorporation of the clean-lab framework rectifies label inconsistencies arising from self-reported data inaccuracies, further enhancing model accuracy. Results indicate that the combination of personalized fine-tuning and label noise correction yields the most significant performance improvements, with heightened sensitivity, specificity, and F1 scores. The findings highlight the importance of individualized approaches and data quality control mechanisms in harnessing the potential of wearable technology for early disease detection. By addressing challenges related to variability and data quality, wearable technology may emerge as a powerful tool for timely disease identification and intervention. Jayden Myers, Wenqi Shi 0002, May D. Wang, Zewei Lei, Benoit Marteau |
BIBM | 5 |
| 2023 | Latent Topic Extraction as a Source of Labeling in Natural Language ProcessingabstractSupervised machine learning algorithms depend on accurate labeling of target data to develop models that can derive relationships between input data and the target data. One major hindrance for developing supervised machine learning models capable of predicting the correct target label of unseen data rests on the quality of the data used to train the models, which often depends on having a subject matter expert (SME) create a labeled dataset to train the model on. Given the scarcity of such experts in many fields, the time needed to analyze data for labeling, and subjective differences among experts, ways to reduce the complexity associated with creating meaningful datasets are needed. In this work, we explore the use of two unsupervised topic modeling algorithms, Latent Dirichlet Allocation (LDA) and Non-negative Matrix Factorization (NMF) as potential methods for reducing the complexities in the labeling process. Specifically, we obtained COVID patient message data labeled by a SME and compared the overlap in topics designated as COVID versus not by the two algorithms to those of the SME. For each of the topic modeling algorithms, we found a strong degree of overlap in the COVID vs. non-COVID patient message labels with that of the SME, suggesting that the methodology could be used to provide synergies for developing labeled data sets used for clinically meaningful models. Andrew Hornback, Yuanda Zhu, Monica Isgut, Wenqi Shi 0002, Arjita Nema, Blake J. Anderson, May D. Wang |
BIBM | 7 |
| 2023 | Identification of Single-Cell RNA Sequencing Molecular Signatures for COVID-19 Infection Severity ClassificationabstractIn this study, we propose a graph-based framework to identify scRNA-seq molecular signatures for COVID-19 infection severity identification. We conduct extensive experiments on scRNA-seq data from bronchoalveolar lavage fluid (BALF) with four machine learning models: Support Vector Machine, Random Forest, Graph Convolutional Network (GCN), and Graph Attention Network (GAT). In addition, we employ an explainable artificial intelligence approach, GNNExplainer, to interpret model predictions by identifying the top 15 features that contribute to the severity classification. Our finding suggests that graphical models could accurately distinguish healthy people from COVID-19 patients (F1-score > 0.9) based on patient scRNA-seq data, and traditional machine learning approaches could accurately distinguish COVID-19 patients with different severity (F1-score > 0.99), along with meaningful molecular signatures identification and discovery. Our implementation is available on a github repository: https://github.com/Da2f1eW/COVID19_Infection_Severity_Classification. Xinling Li 0003, Wei-An Chen, Wenqi Shi 0002, May D. Wang |
BIBM | 5 |
| 2023 | Development of Interpretable Machine Learning Models for COVID-19 Drug Target Docking Scores PredictionabstractWith the extensive time and financial requirements incumbent on drug discovery, computational approaches, such as protein-ligand docking predictions, are increasingly crucial for accelerating the process of drug repurposing. However, the proliferation of identified protein targets has exposed a critical knowledge gap in developing robust models that offer both generalizability and interpretability for docking score prediction. Addressing this, our study presents a machine learning-based surrogate model, employing interpretable artificial intelligence techniques for accurate docking score prediction for SARS-CoV-2 protein targets. We demonstrate the model generalization on its expansion to accommodate unseen protein targets by integrating protein target information through feature concatenation. Moreover, we leverage the SHapley Additive exPlanations (SHAP) method to identify the data-driven feature importance of molecular substructures for knowledge-based validation. Our experiments reveal that the combination of data-driven prediction and knowledge-driven validation could provide biomedical insights into the interactions between drugs and SARS-CoV-2 proteins, elucidating their consequent effects on docking scores. Wenqi Shi 0002, Mio Murakoso, May D. Wang |
BIBM | 4 |
| 2023 | Effective Surrogate Models for Docking Scores Prediction of Candidate Drug Molecules on SARS-CoV-2 Protein TargetsabstractEmerging infectious diseases, such as coronavirus disease 2019 (COVID-19), pose a major threat to public health and present a critical challenge for drug discovery. Due to the cost- and time-consuming process of new drug development, virtual pre-screening methods such as protein-ligand docking prediction have become essential tools in enhancing drug refurbishment and repurposing. In this study, we propose a machine learning-based surrogate model for docking score prediction of drug candidates on SARS-CoV-2 protein targets via deep feature concatenation. We investigate 14 different combinations of rule-based and data-driven fingerprinting methods to identify the optimal representation of candidate drug molecules. Extensive experiments on docking scores of 270,000 molecules across 18 different SARS-CoV-2 protein targets demonstrate the effectiveness of the proposed surrogate models. In addition to unseen drugs, we further investigate the generalization of the proposed framework for unseen protein targets. This study may provide an instrumental and generalizable framework for exploring ligand-protein interaction, serving as a useful tool to facilitate rapid drug pre-screening during emerging public health crises. Wenqi Shi 0002, Mio Murakoso, Linxi Xiong, Matthew Chen, May D. Wang |
BIBM | 6 |
| 2023 | Multi-Modal Deep Feature Integration for Alzheimer's Disease StagingabstractAlzheimer's disease (AD) is one of the leading causes of dementia and 7th leading cause of death in the United States. The provisional diagnosis of AD relies on comprehensive examinations, including medical history, neurological and psychiatric examinations, cognitive assessments, and neuroimaging studies. Integrating diverse sets of clinical data, including electronic health records (EHRs), medical imaging, and genomic data, enables a holistic view of AD staging analysis. In this study, we propose an end-to-end deep learning architecture to jointly learn from magnetic resonance imaging (MRI), positron emission tomography (PET), EHRs, and genomics data to classify patients into AD, mild cognitive disorders, and controls. We conduct extensive experiments to explore different feature-level and intermediate-level fusion methods. Our findings suggest intermediate multiplicative fusion achieves the best stage prediction performance on the external validation dataset. Compared with unimodal baselines, we can observe that integrative approaches that leverage all four modalities demonstrate superior performance to baselines reliant solely on one or two modalities. In an age-wise comparison, we observe a unique pattern that all fusion methods exhibited superior performance in the earlier age brackets (50–70 years), with performance diminishing as the age group advanced (70–90 years). The proposed integration framework has the potential to augment our understanding of disease diagnosis and progression by leveraging complementary information from multimodal patient data. Wenqi Shi 0002, May D. Wang |
BIBM | 3 |
| 2023 | Uncertainty-Aware Ensemble Learning Models for Out-of-Distribution Medical Imaging AnalysisabstractAdvanced deep-learning techniques have been employed to develop clinical decision support systems for diagnosis and prognosis using medical images. However, the presence of out-of-distribution (OOD) samples, which deviate from the training data distribution, poses a significant challenge. Accurate quantification of the predictive uncertainty is crucial for ensuring reliable and dependable implementation in medical settings as a clinical decision support system. In this work, we propose an ensemble model to derive predictive uncertainty estimates for uncertainty quantification on OOD medical imaging. Specifically, the models are initialized with ImageNet pre-trained weights and fine-tuned on chest Computed Tomography (CT). Moreover, we utilize Grad-CAM to visualize and interpret the areas of the image that contribute most to the model’s predictions and uncertainty estimates. This visualization technique enhances the in-terpretability of our ensemble model and supports more informed clinical decision-making. Through extensive experiments on three Chest CT datasets, we have demonstrated the effectiveness of our approach in estimating uncertainty under domain shifting. Our results provide valuable insights into the reliability and specificity of deep ensemble uncertainty predictions in medical image analysis. Our Uncertainty-Aware Ensemble (UAE) approach can enable reliable and transparent predictions for safety-critical medical applications. J. Ben Tamo, Micky C. Nnamdi, Lea Lesbats, Wenqi Shi 0002, Yishan Zhong, May D. Wang |
BIBM | 6 |
| 2023 | Explainable synthetic image generation to improve risk assessment of rare pediatric heart transplant rejectionabstractExpert microscopic analysis of cells obtained from frequent heart biopsies is vital for early detection of pediatric heart transplant rejection to prevent heart failure. Detection of this rare condition is prone to low levels of expert agreement due to the difficulty of identifying subtle rejection signs within biopsy samples. The rarity of pediatric heart transplant rejection also means that very few gold-standard images are available for developing machine learning models. To solve this urgent clinical challenge, we developed a deep learning model to automatically quantify rejection risk within digital images of biopsied tissue using an explainable synthetic data augmentation approach. We developed this explainable AI framework to illustrate how our progressive and inspirational generative adversarial network models distinguish between normal tissue images and those containing cellular rejection signs. To quantify biopsy-level rejection risk, we first detect local rejection features using a binary image classifier trained with expert-annotated and synthetic examples. We converted these local predictions into a biopsy-wide rejection score via an interpretable histogram-based approach. Our model significantly improves upon prior works with the same dataset with an area under the receiver operating curve (AUROC) of 98.84% for the local rejection detection task and 95.56% for the biopsy-rejection prediction task. A biopsy-level sensitivity of 83.33% makes our approach suitable for early screening of biopsies to prioritize expert analysis. Our framework provides a solution to rare medical imaging challenges currently limited by small datasets. Felipe O. Giuste, Ryan Sequeira, Vikranth Keerthipati, Peter Lais, Ali Mirzazadeh, Arshawn Mohseni, Yuanda Zhu, Wenqi Shi 0002, Benoit Marteau, Yishan Zhong, Li Tong 0001, Bibhuti Das 0002, Bahig M. Shehata, Shriprasad R. Deshpande, May D. Wang |
J. Biomed. Informatics | 15 |
| 2022 | Attention-based Automated Chest CT Image Segmentation Method of COVID-19 Lung InfectionabstractAccording to the World Health Organization, Artificial Intelligence (AI) technology may assist in COVID-19 management. However, existing image segmentation using AI suffers from a lack of accuracy and explainability, which prevents its adoption in actual clinical practice. In this paper, we investigated an attention-based image segmentation method for COVID-19 CT imaging with enhanced interpretation capabilities. Specifically, we developed U-Net architecture-based for segmentation with attention coefficients to produce a salient feature map. We use the DICE score and accuracy to perform a comprehensive model evaluation. We compared to other well-known methods such as Light U-Net, COPLE-Net, and Res U-Net and demonstrated that attention U-Net is superior for COVID-19 segmentation tasks in terms of performance and explainability. We also developed the tool as a web-application with a graphic user interface with the goal to translate this AI-driven clinical decision-support system for real-world clinical use. Beom J. Lee, Sarkis T. Martirosyan, Zaid Khan 0003, Han Y. Chiu, Wenqi Shi 0002, Felipe O. Giuste, Yishan Zhong, Jimin Sun, May D. Wang |
BIBE | 10 |
| 2022 | Interpretable Evaluation of Diabetic Retinopathy Grade Regarding Eye Color Fundus ImagesabstractThis paper reports an interpretable automated grading system for diabetic retinopathy using color fundus images. First, we develop shallow learners as baselines. Second, we pre-train deep neural networks to extract high-dimensional features and complex patterns from fundus images and utilize ensemble models to do automatic grading. Then we develop several explainable artificial intelligence models to visualize the extracted deep features and to interpret the predicted outcomes. We investigate the robustness of our system over two publicly available diabetic retinopathy fundus imaging datasets. In addition, we displayed both local and global explainable results to further illustrate the clinical decision-making process with deep models. The innovations of our work include (1) using ensemble models to boost the performance of diabetic retinopathy grading system, and (2) providing transparency of ensemble models using explainable artificial intelligence. The result has shown the potential to improve the effectiveness and accessibility of diabetic retinopathy screening in clinical practice and research settings. Jieh Sheng Hsu, Noaima Bari, Xu Qiu, Malvika Viswanathan, Wenqi Shi 0002, Felipe O. Giuste, Yishan Zhong, Jimin Sun, May D. Wang |
BIBE | 10 |
| 2022 | Multi-Modal Deep Learning Models for Alzheimer's Disease Prediction Using MRI and EHRabstractAlzheimer's Disease (AD) is an irreversible and progressive neurodegenerative disorder with three stages: cognitively normal (CN), mild cognitive impairment (MCI), and clinical dementia. Progression and stage prediction of dementia plays an important role in prognosis and treatment. In this work, we developed a multi-modal AD progress prediction model that integrates magnetic resonance imaging (MRI) and electronic health record (EHR) to classify patients into three stages: CN, MCI, and AD. We trained deep auto-encoder to extract features from EHR data, and ResNet and 3D U-Net for MRI imaging data. We developed an entropy-based weighted sum classification method to integrate the classification results from each individual modality to generate final prediction. We experimented on Alzheimer's Disease Neuroimaging Initiative (ADNI) data to demonstrate that the multi-modality integration model outperforms single modality models in accuracy, precision, recall, and F1 scores. In addition, our model achieves competitive performance in comparison with other state-of-the-art multi-modality integration methods on AD progression prediction. Sathvik S. Prabhu, John A. Berkebile, Neha Rajagopalan, Renjie Yao, Wenqi Shi 0002, Felipe O. Giuste, Yishan Zhong, Jimin Sun, May D. Wang |
BIBE | 9 |
| 2022 | Development of Machine Learning Regression Model for COVID-19 Drug Target PredictionabstractThere is a perennial need to identify novel, effective therapeutic agents to combat rising infections. Recently, prediction of therapeutic targets to decrease the impact of COVID-19 has posed an urgent challenge requiring innovative solutions. Successful identification of novel drug-target combinations may greatly facilitate drug development. To meet this need, we developed a COVID-19 drug target prediction model using machine learning approaches to quickly identify drug candidates for 18 COVID-19 protein targets. Specifically, we analyzed the performance of three prediction models to predict drug-target docking scores, which represents the strength of interactions between ligands and proteins. Docking scores were predicted for 300,457 molecules on 18 different COVID-19 related protein docking targets. Our proposed approach achieved a competitive performance with $\mathrm{R}^{2}$=0.69,MAE=0.285, MSE=0.627. In addition, we identify chemical structures associated with stronger binding affinities across target binding sites. We believe our work could potentially save pharmaceutical companies significant resources, especially during the early stages of drug development. Alexandra Zamitalo, Qingtong Xie, Mayar Allam, Phinu Philip, Wenqi Shi 0002, Felipe O. Giuste, Benoit Marteau, Mio Murakoso, May D. Wang |
BIBM | 9 |
| 2022 | Proposing Causal Sequence of Death by Neural Machine Translation in Public Health InformaticsabstractEach year there are nearly 57 million deaths worldwide, with over 2.7 million in the United States. Timely, accurate and complete death reporting is critical for public health, especially during the COVID-19 pandemic, as institutions and government agencies rely on death reports to formulate responses to communicable diseases. Unfortunately, determining the causes of death is challenging even for experienced physicians. The novel coronavirus and its variants may further complicate the task, as physicians and experts are still investigating COVID-related complications. To assist physicians in accurately reporting causes of death, an advanced Artificial Intelligence (AI) approach is presented to determine a chronically ordered sequence of conditions that lead to death (named as the causal sequence of death), based on decedent's last hospital discharge record. The key design is to learn the causal relationship among clinical codes and to identify death-related conditions. There exist three challenges: different clinical coding systems, medical domain knowledge constraint, and data interoperability. First, we apply neural machine translation models with various attention mechanisms to generate sequences of causes of death. We use the BLEU (BiLingual Evaluation Understudy) score with three accuracy metrics to evaluate the quality of generated sequences. Second, we incorporate expert-verified medical domain knowledge as constraints when generating the causal sequences of death. Lastly, we develop a Fast Healthcare Interoperability Resources (FHIR) interface that demonstrates the usability of this work in clinical practice. Our results match the state-of-art reporting and can assist physicians and experts in public health crisis such as the COVID-19 pandemic. Yuanda Zhu, Ying Sha, Mai Li, Ryan Hoffman, May D. Wang |
IEEE J. Biomed. Health Informatics | 6 |
| 2021 | A Gradient-based Approach for Explaining Multimodal Deep Learning ClassifiersabstractIn recent years, more biomedical studies have begun to use multimodal data to improve model performance. Many studies have used ablation for explainability, which requires the modification of input data. This can create out-of-distribution samples and lead to incorrect explanations. To avoid this problem, we propose using a gradient-based feature attribution approach, called layer-wise relevance propagation (LRP), to explain the importance of modalities both locally and globally for the first time. We demonstrate the feasibility of the approach with sleep stage classification as our use-case and train a 1-D convolutional neural network with electroencephalogram (EEG), electrooculogram (EOG), and electromyogram (EMG) data. We also analyze the relationship of our local explainability results with clinical and demographic variables to determine whether they affect our classifier. Across all samples, EEG is the most important modality, followed by EOG and EMG. For individual sleep stages, EEG and EOG have higher relevance for awake and non-rapid eye movement 1 (NREM1). EOG is most important for REM, and EEG is most relevant for NREM2-NREM3. Also, LRP gives consistent levels of importance to each modality for the correctly classified samples across folds but inconsistent levels of importance for incorrectly classified samples. Our statistical analyses suggest that medication has a significant effect upon patterns learned for EEG and EOG NREM2 and that subject sex and age significantly affects the EEG and EOG patterns learned, respectively. Our results demonstrate the viability of gradient-based approaches for explaining multimodal electrophysiology classifiers and suggest their generalizability for other multimodal classification domains. Charles A. Ellis, Rongen Zhang, Vince D. Calhoun, Darwin A. Carbajal, Robyn L. Miller, May D. Wang |
BIBE | 6 |
| 2021 | A Novel Local Ablation Approach for Explaining Multimodal ClassifiersabstractWith the growing use of multimodal data for deep learning classification in healthcare research, more studies are presenting explainability methods for insight into multimodal classifiers. Among these studies, few utilize local explainability methods, which can provide (1) insight into the classification of samples over time and (2) better understanding of the effects of demographic and clinical variables upon patterns learned by classifiers. To the best of our knowledge, we present the first local explainability approach for insight into the importance of each modality to the classification of samples over time. Our approach uses ablation, and we demonstrate how it can show the importance of each modality to the correct classification of each class. We further present a novel analysis that explores the effects of demographic and clinical variables upon the multimodal patterns learned by the classifier. As a use-case, we train a convolutional neural network for automated sleep staging with electroencephalogram (EEG), electrooculogram (EOG), and electromyogram (EMG) data. We find that EEG is the most important modality across most stages, though EOG is particularly important for non-rapid eye movement stage 1. Further, we identify significant relationships between the local explanations and subject age, sex, and state of medication which suggest that the classifier learned features associated with these variables across multiple modalities and correctly classified samples. Our novel explainability approach has implications for many fields involving multimodal classification. Moreover, our examination of the degree to which demographic and clinical variables may affect classifiers could provide direction for future studies in automated biomarker discovery. Charles A. Ellis, Rongen Zhang, Vince D. Calhoun, Darwin A. Carbajal, Mohammad S. Eslampanah Sendi, May D. Wang, Robyn L. Miller |
BIBE | 6 |
| 2021 | A FHIR-compliant Application for Multi-Site and Multi-Modality Pediatric Scoliosis Patient RehabilitationabstractScoliosis is a spinal curvature that most frequently affects adolescents. Posterior spinal fusion surgery is required to correct the deformity in patients with severe scoliosis. Surgeons frequently use radiographic measurements and patient reported outcomes to aid in surgical treatment and monitor patient rehabilitation. Shriners Hospitals for Children is a large healthcare system caring for a significant percentage of pediatric patients with scoliosis. Surgeons from SHC-Greenville and SHC-Lexington have recorded data from more than 1,000 individual scoliosis patients. However, these collected data are usually dispersed across individual healthcare sites, necessitating the development of an integrated clinical data repository for data sharing and management. In this paper, we established a standardized research data repository with FHIR resources to harmonize multi-modal patient data from multiple clinical sites. Additionally, a FHIR-compliant application with a web-based user interface was prototyped to enable clinicians and researchers to access scoliosis patient data within our integrated and standardized research repository. Patient cohort definitions can be used to search these records using the same FHIR application. This standardized data-sharing framework and healthcare information system can be applied to multi-site and multimodality studies for clinical and research purposes, with the ultimate goal of improving the quality of patient care. Wenqi Shi 0002, Felipe O. Giuste, Yuanda Zhu, Ashley M. Carpenter, Henry J. Iwinski, Coleman Hilton, J. Michael Wattenbarger, May D. Wang |
BIBM | 8 |
| 2021 | COVID-19 Automatic Diagnosis With Radiographic Imaging: Explainable Attention Transfer Deep Neural NetworksabstractResearchers seek help from deep learning methods to alleviate the enormous burden of reading radiological images by clinicians during the COVID-19 pandemic. However, clinicians are often reluctant to trust deep models due to their black-box characteristics. To automatically differentiate COVID-19 and community-acquired pneumonia from healthy lungs in radiographic imaging, we propose an explainable attention-transfer classification model based on the knowledge distillation network structure. The attention transfer direction always goes from the teacher network to the student network. Firstly, the teacher network extracts global features and concentrates on the infection regions to generate attention maps. It uses a deformable attention module to strengthen the response of infection regions and to suppress noise in irrelevant regions with an expanded reception field. Secondly, an image fusion module combines attention knowledge transferred from teacher network to student network with the essential information in original input. While the teacher network focuses on global features, the student branch focuses on irregularly shaped lesion regions to learn discriminative features. Lastly, we conduct extensive experiments on public chest X-ray and CT datasets to demonstrate the explainability of the proposed architecture in diagnosing COVID-19. Wenqi Shi 0002, Li Tong 0001, Yuanda Zhu, May D. Wang |
IEEE J. Biomed. Health Informatics | 4 |
| 2020 | Mitigating Patient-to-Patient Variation in EEG Seizure Detection using Meta Transfer LearningabstractElectroencephalogram (EEG) signals can be used for seizure detection, but the seizure patterns found in between patient's EEGs can have significant variations. Specifically, focal spikes in patient-specific channels as well as other patient specific patterns can strongly indicate seizure activity. Manual diagnosis on these markers leads to inconsistent interrater agreement and poor detection accuracy. Previous automation attempts have ignored patient specific approaches but fail to generalize to previously unseen patients. To reduce subjectivity in manual diagnosis, we propose an automatic seizure detection pipeline that includes quality control, preprocessing, and meta transfer learning for both feature extraction and classification. To mitigate the inter-patient seizure pattern variation, we adapt Meta UPdate Strategy (MUPS) for four-class classification on the world's largest public seizure dataset of EEGs, Temple University Seizure Corpus (TUSZ). Different from existing works on binary seizure detection, we use the non-seizure samples and the top three most frequent seizure types for seizure detection. Our experiments show that the meta transfer learning approach achieves macro-F1 of 0.5103 and AUC of 0.6792, which outperforms the baseline learners (shallow and deep) by mitigating patient-to patient variations. We demonstrate the effectiveness of meta transfer learning in feature extraction and classification for multi-class seizure detection. Yuanda Zhu, Mohammed Saqib, Elizabeth Ham, Sami Belhareth, Ryan Hoffman, May D. Wang |
BIBE | 6 |
| 2020 | Graph Convolutional Neural Networks to Classify Whole Slide ImagesabstractConventional biomedical measurement and tests used for patient diagnosis are increasingly digitized in the current clinical practice, which provides great opportunities for researchers to build smart computer-aided decision support systems using advanced data analytics. In this paper, we discuss how to build such systems for whole slide images (WSIs), the digitization of entire histology slides used by pathologists. While conventional techniques use grid-style-tile-based processing strategy with extracted features aggregated based on first-order statistics to detect regions of interest such as cancerous regions, they ignore the spatial relationships between nearby tiles: the closer the two tiles are spatially, the more similar they are in disease conditions. To capture such spatial proximity information, we present a novel application of graph convolutional networks (GCNs) to analyze WSIs. We model each tile of a WSI as a node in a graph, and apply GCNs to the resulting graph to aggregate the features for final classification. We use two histopathological image classification datasets, a breast cancer pathological dataset of 58 total samples with 36 benign ones, and another colon cancer dataset of 100 H&E images with 49 benign ones. The diagnosis accuracy achieved by GCN is consistently better than that achieved by conventional methods. The experiments results showed the potential of graph-based neural networks to improve biomedical imaging data analysis. Roshan Konda, May D. Wang |
COMPSAC | 3 |
| 2020 | Generating Region of Interests for Invasive Breast Cancer in Histopathological Whole-Slide-ImageabstractThe detection of the region of interests (ROIs) on Whole Slide Images (WSIs) is one of the primary steps in computer-aided cancer diagnosis and grading. Early and accurate identification of invasive cancer regions in WSI is critical in the improvement of breast cancer diagnosis and further improvements in patient survival rates. However, invasive cancer ROI segmentation is a challenging task on WSI because of the low contrast of invasive cancer cells and their high similarity in terms of appearance, to non-invasive regions. In this paper, we propose a CNN based architecture for generating ROIs through segmentation. The network tackles the constraints of data-driven learning and working with very low-resolution WSI data in the detection of invasive breast cancer. Our proposed approach is based on transfer learning and the use of dilated convolutions. We propose a highly modified version of U-Net based auto-encoder, which takes as input an entire WSI with a resolution of 320×320. The network was trained on low-resolution WSI from four different data cohorts and has been tested for inter as well as intra- dataset variance. The proposed architecture shows significant improvements in terms of accuracy for the detection of invasive breast cancer regions. Shreyas Malakarjun Patil, Li Tong 0001, May D. Wang |
COMPSAC | 3 |
| 2020 | Regularization of Deep Neural Networks for EEG Seizure Detection to Mitigate OverfittingabstractSeizure detection is a major goal for simplifying the workflow of clinicians working on EEG records. Current algorithms can only detect seizures effectively for patients already presented to the classifier. These algorithms are hard to generalize outside the initial training set without proper regularization and fail to capture seizures from the larger population. We proposed a data processing pipeline for seizure detection on an intra-patient dataset from the world's largest public EEG seizure corpus. We created spatially and session invariant features by forcing our networks to rely less on exact combinations of channels and signal amplitudes, but instead to learn dependencies towards seizure detection. For comparison, the baseline results without any additional regularization on a deep learning model achieved an F1 score of 0.544. By using random rearrangements of channels on each minibatch to force the network to generalize to other combinations of channels, we increased the F1 score to 0.629. By using random rescale of the data within a small range, we further increased the F1 score to 0.651 for our best model. Additionally, we applied adversarial multi-task learning and achieved similar results. We observed that session and patient specific dependencies were causing overfitting of deep neural networks, and the most overfitting models learnt features specific only to the EEG data presented. Thus, we created networks with regularization that the deep learning did not learn patient and session-specific features. We are the first to use random rearrangement, random rescale, and adversarial multitask learning to regularize intra-patient seizure detection and have increased sensitivity to 0.86 comparing to baseline study. Mohammed Saqib, Yuanda Zhu, May D. Wang, Brett K. Beaulieu-Jones |
COMPSAC | 3 |
| 2020 | An Information Theoretic Learning for Causal Direction IdentificationabstractCausal inference has been one of the central problems in many data science research and one of the most important problem in causal inference is to draw causal conclusions using observation data. In this paper, we focus on learning causal relationships between variables using observation data. We proposed novel scoring method based on mutual information and corresponding learning algorithms. We then showed the consistency of our models under mild conditions and also validated the effectiveness of our approach in benchmarking experiments. May D. Wang |
COMPSAC | 2 |
| 2020 | Training Confidence-Calibrated Classifier via Distributionally Robust LearningabstractSupervised learning via empirical risk minimization, despite its solid theoretical foundations, faces a major challenge in generalization capability, which limits its application in real-world data science problems. In particular, current models fail to distinguish in-distribution and out-of-distribution and give over confident predictions for out-of-distribution samples. In this paper, we propose an distributionally robust learning method to train classifiers via solving an unconstrained minimax game between an adversary test distribution and a hypothesis. We showed the theoretical generalization performance guarantees, and empirically, our learned classifier when coupled with thresholded detectors, can efficiently detect out-of-distribution samples. May D. Wang |
COMPSAC | 2 |
| 2020 | Graph Convolutional Neural Networks to Classify Whole Slide ImagesabstractWhole slide images (WSIs) are the digitization of histology slides and are increasingly used by pathologists to detect the cancerous regions and make diagnosis for patients. Recent machine learning and deep learning algorithms have shown great success in automated classification of WSIs. However, given the computational challenge associated with processing high resolution WSIs, conventional techniques rely on a patch-based approach, and subsequently aggregate extracted features using firs-order statistics. However, as cancerous regions are clustered together, such an approach ignores the spatial relationships within each tile. Here, we present a novel application of graph convolutional networks (GCNs) to analyze WSIs. GCNs are powerful deep neural networks for modelling node relationships in a graph. To capture the spatial information, we model each tile as a node in the graph, and aggregate features by applying GCNs to the graph. The classification results on real-world histopathology datasets shows improved performance over conventional methods, and highlights the potential of graph-based methods in biomedical data analytics. Roshan Konda, May D. Wang |
ICASSP | 3 |
| 2020 | Editorial Special Issue on "AI-Driven Informatics, Sensing, Imaging and Big Data Analytics for Fighting the COVID-19 Pandemic"abstractThe papers in this special section focuses on artificial intelligent-driven informatics, sensing, imaging and big data analytics in dealing with the COVID-19 pandemic. Amir A. Amini, Wei Chen 0015, Giancarlo Fortino, Ye Li 0002, Yi Pan 0001, May D. Wang |
IEEE J. Biomed. Health Informatics | 6 |
| 2019 | Hybrid Modeling of Ebola PropagationabstractThe Ebola virus disease (EVD) epidemic that occurred in West Africa between 2014-16 resulted in over 28,000 cases and 11,000 deaths - one of the deadliest to date. A generalized model of the spatiotemporal progression of EVD for Liberia, Guinea, and Sierra Leone in 2014-16 remains elusive. There is also a disconnect in the literature on which interventions are most effective in curbing disease progression. To solve these two key issues, we designed a hybrid agent-based and compartmental model that switches from one paradigm to the other on a stochastic threshold. We modeled disease progression with promising accuracy using WHO datasets. Cyrus Tanade, Nathanael Pate, Elianna Paljug, Ryan Hoffman, May D. Wang |
BIBE | 5 |
| 2019 | A Translational Pipeline for Overall Survival Prediction of Breast Cancer Patients by Decision-Level Integration of Multi-Omics DataabstractBreast cancer is the most prevalent and among the most deadly cancers in females. Patients with breast cancer have highly variable survival rates, indicating a need to identify prognostic biomarkers. By integrating multi-omics data (e.g., gene expression, DNA methylation, miRNA expression, and copy number variations (CNVs)), it is likely to improve the accuracy of patient survival predictions compared to prediction using single modality data. Therefore, we propose to develop a machine learning pipeline using decision-level integration of multi-omics tumor data from The Cancer Genome Atlas (TCGA) to predict the overall survival of breast cancer patients. With multi-omics data consisting of gene expression, methylation, miRNA expression, and CNVs, the top performing model predicted survival with an accuracy of 85% and area under the curve (AUC) of 87%. Furthermore, the model was able to identify which modalities best contributed to prediction performance, identifying methylation, miRNA, and gene expression as the best integrated classification combination. Our method not only recapitulated several breast cancer-specific prognostic biomarkers that were previously reported in the literature but also yielded several novel biomarkers. Further analysis of these biomarkers could lend insight into the molecular mechanisms that lead to poor survival. Jonathan Mitchel, Kevin Chatlin, Li Tong 0001, May D. Wang |
BIBM | 4 |
| 2019 | Improving Classification of Breast Cancer by Utilizing the Image Pyramids of Whole-Slide Imaging and Multi-scale Convolutional Neural NetworksabstractWhole-slide imaging (WSI) is the digitization of conventional glass slides. Automatic computer-aided diagnosis (CAD) based on WSI enables digital pathology and the integration of pathology with other data like genomic biomarkers. Numerous computational algorithms have been developed for WSI, with most of them taking the image patches cropped from the highest resolution as the input. However, these models exploit only the local information within each patch and lost the connections between the neighboring patches, which may contain important context information. In this paper, we propose a novel multi-scale convolutional network (ConvNet) to utilize the built-in image pyramids of WSI. For the concentric image patches cropped at the same location of different resolution levels, we hypothesize the extra input images from lower magnifications will provide context information to enhance the prediction of patch images. We build corresponding ConvNets for feature representation and then combine the extracted features by 1) late fusion: concatenation or averaging the feature vectors before performing classification, 2) early fusion: merge the ConvNet feature maps. We have applied the multi-scale networks to a benchmark breast cancer WSI dataset. Extensive experiments have demonstrated that our multiscale networks utilizing the WSI image pyramids can achieve higher accuracy for the classification of breast cancer. The late fusion method by taking the average of feature vectors reaches the highest accuracy (81.50%), which is promising for the application of multi-scale analysis of WSI. Li Tong 0001, Ying Sha, May D. Wang |
COMPSAC (1) | 3 |
| 2019 | CAESNet: Convolutional AutoEncoder based Semi-supervised Network for improving multiclass classification of endomicroscopic imagesabstractOBJECTIVE: This article presents a novel method of semisupervised learning using convolutional autoencoders for optical endomicroscopic images. Optical endomicroscopy (OE) is a newly emerged biomedical imaging modality that can support real-time clinical decisions for the grade of dysplasia. To enable real-time decision making, computer-aided diagnosis (CAD) is essential for its high speed and objectivity. However, traditional supervised CAD requires a large amount of training data. Compared with the limited number of labeled images, we can collect a larger number of unlabeled images. To utilize these unlabeled images, we have developed a Convolutional AutoEncoder based Semi-supervised Network (CAESNet) for improving the classification performance. MATERIALS AND METHODS: We applied our method to an OE dataset collected from patients undergoing endoscope-based confocal laser endomicroscopy procedures for Barrett's esophagus at Emory Hospital, which consists of 429 labeled images and 2826 unlabeled images. Our CAESNet consists of an encoder with 5 convolutional layers, a decoder with 5 transposed convolutional layers, and a classification network with 2 fully connected layers and a softmax layer. In the unsupervised stage, we first update the encoder and decoder with both labeled and unlabeled images to learn an efficient feature representation. In the supervised stage, we further update the encoder and the classification network with only labeled images for multiclass classification of the OE images. RESULTS: Our proposed semisupervised method CAESNet achieves the best average performance for multiclass classification of OE images, which surpasses the performance of supervised methods including standard convolutional networks and convolutional autoencoder network. CONCLUSIONS: Our semisupervised CAESNet can efficiently utilize the unlabeled OE images, which improves the diagnosis and decision making for patients with Barrett's esophagus. Li Tong 0001, May D. Wang |
J. Am. Medical Informatics Assoc. | 3 |
| 2019 | Guest Editorial on the Special Issue on Informatics on Biomedical Data Learning, Reasoning, and RepresentationabstractThe papers in this special section was presented at the 8th ACM-BCB Conference that was held in August 2017 in Boston, MA. Tamer Kahveci, Giuseppe Pozzi, Amarda Shehu, May D. Wang |
IEEE J. Biomed. Health Informatics | 4 |
| 2019 | Novel Data Imputation for Multiple Types of Missing Data in Intensive Care UnitsabstractThe diversity and number of parameters monitored in an intensive care unit (ICU) make the resulting databases highly susceptible to quality issues, such as missing information and erroneous data entry, which adversely affect the downstream processing and predictive modeling. Missing data interpolation and imputation techniques, such as multiple imputation, expectation maximization, and hot-deck imputation techniques do not account for the type of missing data, which can lead to bias. In our study, we first model the missing data as three types: "neglectable" also known as a.k.a "missing completely at random," "recoverable" a.k.a. "missing at random," and "not easily recoverable" a.k.a. "missing not at random." We then design imputation techniques for each type of missing data. We use a publicly available database (MIMIC II) to demonstrate how these imputations perform with random forests for prediction. Our results indicate that these novel imputation techniques outperformed standard mean filling techniques and expectation maximization with a statistical significance p ≤ 0.01 in predicting ICU mortality. Janani Venugopalan, Nikhil Chanani, Kevin O. Maher, May D. Wang |
IEEE J. Biomed. Health Informatics | 4 |
| 2018 | Variance Regularized Counterfactual Risk Minimization via Variational Divergence MinimizationabstractOff-policy learning, the task of evaluating and improving policies using historic data collected from a logging policy, is important because on-policy evaluation is usually expensive and has adverse impacts. One of the major challenge of off-policy learning is to derive counterfactual estimators that also has low variance and thus low generalization error. In this work, inspired by learning bounds for importance sampling problems, we present a new counterfactual learning principle for off-policy learning with bandit feedbacks. Our method regularizes the generalization error by minimizing the distribution divergence between the logging policy and the new policy, and removes the need for iterating through all training samples to compute sample variance regularization in prior work. With neural network policies, our end-to-end training algorithms using variational divergence minimization showed significant improvement over conventional baseline algorithms and is also consistent with our theoretical results. May D. Wang |
ICML | 2 |
| 2018 | LncADeep: an ab initio lncRNA identification and functional annotation tool based on deep learningabstractMotivation: To characterize long non-coding RNAs (lncRNAs), both identifying and functionally annotating them are essential to be addressed. Moreover, a comprehensive construction for lncRNA annotation is desired to facilitate the research in the field. Results: We present LncADeep, a novel lncRNA identification and functional annotation tool. For lncRNA identification, LncADeep integrates intrinsic and homology features into a deep belief network and constructs models targeting both full- and partial-length transcripts. For functional annotation, LncADeep predicts a lncRNA's interacting proteins based on deep neural networks, using both sequence and structure information. Furthermore, LncADeep integrates KEGG and Reactome pathway enrichment analysis and functional module detection with the predicted interacting proteins, and provides the enriched pathways and functional modules as functional annotations for lncRNAs. Test results show that LncADeep outperforms state-of-the-art tools, both for lncRNA identification and lncRNA-protein interaction prediction, and then presents a functional interpretation. We expect that LncADeep can contribute to identifying and annotating novel lncRNAs. Availability and implementation: LncADeep is freely available for academic use at http://cqb.pku.edu.cn/ZhuLab/lncadeep/ and https://github.com/cyang235/LncADeep/. Supplementary information: Supplementary data are available at Bioinformatics online. Cheng Yang 0001, Longshu Yang, Haoling Xie, Chengjiu Zhang, May D. Wang, Huaiqiu Zhu |
Bioinform. | 6 |
| 2018 | Intelligent Mortality Reporting With FHIRabstractOne pressing need in the area of public health is timely, accurate, and complete reporting of deaths and the diseases or conditions leading up to them. Fast Healthcare Interoperability Resources (FHIR) is a new HL7 interoperability standard for electronic health record, while Sustainable Medical Applications and Reusable Technologies (SMART)-on-FHIR enables third-party app development that can work "out of the box." This paper demonstrates the feasibility of developing SMART-on-FHIR applications that enables medical professionals to perform timely and accurate death reporting within multiple different USA State jurisdictions. We explored how the information on a standard certificate of death can be mapped to resources defined in the FHIR standard Draft Standard for Trial Use Version 2 and common profiles. We also demonstrated analytics for potentially improving the accuracy and completeness of mortality reporting data. Ryan Hoffman, Janani Venugopalan, Paula Braun, May D. Wang |
IEEE J. Biomed. Health Informatics | 5 |
| 2017 | Models for Predicting Stage in Head and Neck Squamous Cell Carcinoma Using Proteomic and Transcriptomic DataabstractLate diagnosis is one of the reasons that head and neck squamous cell carcinoma (HNSCC) patients experience relative five-year survival rates ranging from 40%-66%. The molecular-level differences between early and advanced stage HNSCC may provide insight into therapeutic targets and strategies. Previous bioinformatics studies have shown mixed or limited results in identifying gene and protein markers and in developing models for discriminating between early and advanced stage HNSCC. Thus, we have investigated models for HNSCC stage prediction using RNAseq and reverse phase protein array data from The Cancer Genome Atlas and The Cancer Proteome Atlas. We systematically assessed individual and ensemble binary classifiers, using filter and wrapper feature selection methods, to develop several well-performing models. In particular, integrated models harnessing both data types consistently resulted in better performance. This study identifies informative protein and gene feature sets which may increase understanding of HNSCC progression. Chanchala Kaddi, May D. Wang |
IEEE J. Biomed. Health Informatics | 2 |
| 2016 | A multi-modal graph-based semi-supervised pipeline for predicting cancer survivalabstractCancer survival prediction is an active area of research that can help prevent unnecessary therapies and improve patient's quality of life. Gene expression profiling is being widely used in cancer studies to discover informative biomarkers that aid predict different clinical endpoint prediction. We use multiple modalities of data derived from RNA deep-sequencing (RNA-seq) to predict survival of cancer patients. Despite the wealth of information available in expression profiles of cancer tumors, fulfilling the aforementioned objective remains a big challenge, for the most part, due to the paucity of data samples compared to the high dimension of the expression profiles. As such, analysis of transcriptomic data modalities calls for state-of-the-art big-data analytics techniques that can maximally use all the available data to discover the relevant information hidden within a significant amount of noise. In this paper, we propose a pipeline that predicts cancer patients' survival by exploiting the structure of the input (manifold learning) and by leveraging the unlabeled samples using Laplacian support vector machines, a graph-based semi supervised learning (GSSL) paradigm. We show that under certain circumstances, no single modality per se will result in the best accuracy and by fusing different models together via a stacked generalization strategy, we may boost the accuracy synergistically. We apply our approach to two cancer datasets and present promising results. We maintain that a similar pipeline can be used for predictive tasks where labeled samples are expensive to acquire. Hamid Reza Hassanzadeh, John H. Phan, May D. Wang |
BIBM | 3 |
| 2016 | DeeperBind: Enhancing prediction of sequence specificities of DNA binding proteinsabstractTranscription factors (TFs) are macromolecules that bind to cis-regulatory specific sub-regions of DNA promoters and initiate transcription. Finding the exact location of these binding sites (aka motifs) is important in a variety of domains such as drug design and development. To address this need, several in vivo and in vitro techniques have been developed so far that try to characterize and predict the binding specificity of a protein to different DNA loci. The major problem with these techniques is that they are not accurate enough in prediction of the binding affinity and characterization of the corresponding motifs. As a result, downstream analysis is required to uncover the locations where proteins of interest bind. Here, we propose DeeperBind, a long short term recurrent convolutional network for prediction of protein binding specificities with respect to DNA probes. DeeperBind can model the positional dynamics of probe sequences and hence reckons with the contributions made by individual sub-regions in DNA sequences, in an effective way. Moreover, it can be trained and tested on datasets containing varying-length sequences. We apply our pipeline to the datasets derived from protein binding microarrays (PBMs), an in-vitro high-throughput technology for quantification of protein-DNA binding preferences, and present promising results. To the best of our knowledge, this is the most accurate pipeline that can predict binding specificities of DNA sequences from the data produced by high-throughput technologies through utilization of the power of deep learning for feature generation and positional dynamics modeling. Hamid Reza Hassanzadeh, May D. Wang |
BIBM | 2 |
| 2016 | Guest Editorial: MobiHealth 2014, IEEE HealthCom 2014, and IEEE BHI 2014abstractThe papers in this special section were presented at three well-known conferences organized in 2014: EAI Mobihealth, IEEE HealthCom, and IEEE Biomedical and Health Informatics. EAI Mobihealth is an annually organized conference, which started in 2010, to address the demands of the rapidly evolving disciplines of wireless communications, mobile computing, and sensing technologies in healthcare. The IEEE-Healthcom is held every year since 1999 in different countries in Asia, Europe, and in America. It aims at bringing together interested parties working in the field of healthcare to exchange ideas, discuss innovative and emerging solutions, and develop collaborations. The IEEE Biomedical Health Informatics Conference started in 2013 and is organized every year providing the forum to showcase enabling technologies of computing, devices, imaging, sensors, and systems that optimize the acquisition, transmission, processing, storage, retrieval, visualization, and analysis of medical data. The aim of this special section is to present an overview of recent advances in sensing technologies, monitoring of patients, security and privacy of data transfer, provision of collaborative environments, data gathering and analysis from various sources, and predictive models, which all finally target the best strategy for patient monitoring and treatment. Metin Akay, Gouenou Coatrieux, Yang Hao 0001, Dimitrios I. Fotiadis, Andrew F. Laine, Benny P. L. Lo, Konstantina S. Nikita, Norbert Noury, Joel J. P. C. Rodrigues, May D. Wang |
IEEE J. Biomed. Health Informatics | 10 |
| 2015 | Integration of multimodal RNA-seq data for prediction of kidney cancer survivalabstractKidney cancer is of prominent concern in modern medicine. Predicting patient survival is critical to patient awareness and developing a proper treatment regimens. Previous prediction models built upon molecular feature analysis are limited to just gene expression data. In this study we investigate the difference in predicting five year survival between unimodal and multimodal analysis of RNA-seq data from gene, exon, junction, and isoform modalities. Our preliminary findings report higher predictive accuracy-as measured by area under the ROC curve (AUC)-for multimodal learning when compared to unimodal learning with both support vector machine (SVM) and k-nearest neighbor (KNN) methods. The results of this study justify further research on the use of multimodal RNA-seq data to predict survival for other cancer types using a larger sample size and additional machine learning methods. Matt Schwartz, Martin Park, John H. Phan, May D. Wang |
BIBM | 4 |
| 2015 | Guest Editorial EMBC 2014abstractThe ten papers from this special sectoin were presented at the 36th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC’14). Walter G. Besio, Leslie Ying, Jie Liang 0002, Nigel H. Lovell, Carmen C. Y. Poon, May D. Wang |
IEEE J. Biomed. Health Informatics | 6 |
| 2014 | Removing Batch Effects From Histopathological Images for Enhanced Cancer DiagnosisabstractResearchers have developed computer-aided decision support systems for translational medicine that aim to objectively and efficiently diagnose cancer using histopathological images. However, the performance of such systems is confounded by nonbiological experimental variations or "batch effects" that can commonly occur in histopathological data, especially when images are acquired using different imaging devices and patient samples. This is even more problematic in large-scale studies in which cross-laboratory sharing of large volumes of data is necessary. Batch effects can change quantitative morphological image features and decrease the prediction performance. Using four batches of renal tumor images, we compare one image-level and five feature-level batch effect removal methods. Principal component variation analysis shows that batch is a large source of variance in image features. Results show that feature-level normalization methods reduce batch-contributed variance to almost zero. Moreover, feature-level normalization, especially ComBatN, improves cross-batch and combined-batch prediction performance. Compared to no normalization, ComBatN improves performance in 83% and 90% of cross-batch and combined-batch prediction models, respectively. Sonal Kothari, John H. Phan, Todd H. Stokes, Adeboye O. Osunkoya, Andrew N. Young, May D. Wang |
IEEE J. Biomed. Health Informatics | 6 |
| 2014 | Guest Editorial: Computational Solutions to Large-Scale Data Management and Analysis in Translational and Personalized MedicineabstractApplying engineering precepts to biological systems has spawn the field of systems biology to investigate a network of interacting components, including the coordination of internal systems of living organisms such as endocrine, nervous, and respiratory with gene and gene product expression, and behavior and environmental factors, and understand how these components together contribute to the disease initiation and progression, biological development, and health. Proceeding from systems biology, systems medicine incorporates complex and dynamic biochemical, physiological, and environmental interactions between all components of disease and health that sustain living organisms. The current special issue includes a selected number of papers presented at the 12th IEEE International Conference on BioInformatics and BioEngineering (BIBE 2012), Nov. 11-13, 2012 under a special session with the same theme, in addition to papers submitted following an open call for papers. The Special Issue presents experiences as well as technological and scientific developments stemming from some flagship projects funded by the EU under the FP7 framework programme aiming to bring together researchers working in the fields of infrastructures and technologies for integrative biomedical research, ICT for predictive and translational medicine and the VPH community at large. A total of 15 papers are included under the following scientific subdomains: 1) mHealth,Wearable Systems and Telemonitoring Services (five papers), 2) Medical Imaging (four papers), and 3) Computational Biology (six papers). Manolis Tsiknakis, Vasilis J. Promponas, Norbert Graf 0001, May D. Wang, Stephen T. C. Wong, Nikolaos G. Bourbakis, Constantinos S. Pattichis |
IEEE J. Biomed. Health Informatics | 4 |
| 2013 | A fast least-squares algorithm for population inferenceabstractBACKGROUND: Population inference is an important problem in genetics used to remove population stratification in genome-wide association studies and to detect migration patterns or shared ancestry. An individual's genotype can be modeled as a probabilistic function of ancestral population memberships, Q, and the allele frequencies in those populations, P. The parameters, P and Q, of this binomial likelihood model can be inferred using slow sampling methods such as Markov Chain Monte Carlo methods or faster gradient based approaches such as sequential quadratic programming. This paper proposes a least-squares simplification of the binomial likelihood model motivated by a Euclidean interpretation of the genotype feature space. This results in a faster algorithm that easily incorporates the degree of admixture within the sample of individuals and improves estimates without requiring trial-and-error tuning. RESULTS: We show that the expected value of the least-squares solution across all possible genotype datasets is equal to the true solution when part of the problem has been solved, and that the variance of the solution approaches zero as its size increases. The Least-squares algorithm performs nearly as well as Admixture for these theoretical scenarios. We compare least-squares, Admixture, and FRAPPE for a variety of problem sizes and difficulties. For particularly hard problems with a large number of populations, small number of samples, or greater degree of admixture, least-squares performs better than the other methods. On simulated mixtures of real population allele frequencies from the HapMap project, Admixture estimates sparsely mixed individuals better than Least-squares. The least-squares approach, however, performs within 1.5% of the Admixture error. On individual genotypes from the HapMap project, Admixture and least-squares perform qualitatively similarly and within 1.2% of each other. Significantly, the least-squares approach nearly always converges 1.5- to 6-times faster. CONCLUSIONS: The computational advantage of the least-squares approach along with its good estimation performance warrants further research, especially for very large datasets. As problem sizes increase, the difference in estimation performance between all algorithms decreases. In addition, when prior information is known, the least-squares approach easily incorporates the expected degree of admixture to improve the estimate. R. Mitchell Parry, May D. Wang |
BMC Bioinform. | 2 |
| 2013 | Assessing the impact of human genome annotation choice on RNA-seq expression estimatesabstractBACKGROUND: Genome annotation is a crucial component of RNA-seq data analysis. Much effort has been devoted to producing an accurate and rational annotation of the human genome. An annotated genome provides a comprehensive catalogue of genomic functional elements. Currently, at least six human genome annotations are publicly available, including AceView Genes, Ensembl Genes, H-InvDB Genes, RefSeq Genes, UCSC Known Genes, and Vega Genes. Characteristics of these annotations differ because of variations in annotation strategies and information sources. When performing RNA-seq data analysis, researchers need to choose a genome annotation. However, the effect of genome annotation choice on downstream RNA-seq expression estimates is still unclear. This study (1) investigates the effect of different genome annotations on RNA-seq quantification and (2) provides guidelines for choosing a genome annotation based on research focus. RESULTS: We define the complexity of human genome annotations in terms of the number of genes, isoforms, and exons. This definition facilitates an investigation of potential relationships between complexity and variations in RNA-seq quantification. We apply several evaluation metrics to demonstrate the impact of genome annotation choice on RNA-seq expression estimates. In the mapping stage, the least complex genome annotation, RefSeq Genes, appears to have the highest percentage of uniquely mapped short sequence reads. In the quantification stage, RefSeq Genes results in the most stable expression estimates in terms of the average coefficient of variation over all genes. Stable expression estimates in the quantification stage translate to accurate statistics for detecting differentially expressed genes. We observe that RefSeq Genes produces the most accurate fold-change measures with respect to a ground truth of RT-qPCR gene expression estimates. CONCLUSIONS: Based on the observed variations in the mapping, quantification, and differential expression calling stages, we demonstrate that the selection of human genome annotation results in different gene expression estimates. When conducting research that emphasizes reproducible and robust gene expression estimates, a less complex genome annotation may be preferred. However, simpler genome annotations may limit opportunities for identifying or characterizing novel transcriptional or regulatory mechanisms. When conducting research that aims to be more exploratory, a more complex genome annotation may be preferred. Po-Yen Wu, John H. Phan, May D. Wang |
BMC Bioinform. | 3 |
| 2013 | Biomedical imaging informatics in the era of precision medicine: progress, challenges, and opportunitiesabstractBiomedical informatics is the interdisciplinary field that studies and pursues the effective uses of biomedical data, information, and knowledge for scientific inquiry, problem solving, and decision making, motivated by efforts to improve human health.1 ,2 Not only do biomedical informaticians study and develop theories, methods, and processes for the generation, manipulation, and sharing of biomedical data, but they also investigate how to model and reason on these data in order to effect beneficial change in the healthcare enterprise. In addition, an important aspect associated with developments in this field is the consideration of social and behavioral sciences in the design and evaluation of technical solutions. As a subfield of biomedical informatics, biomedical imaging informatics (BMII) encompasses all of the aforementioned aspects from the perspective of imaging. BMII has emerged as one of the fastest growing research areas in recent years given the evolution of techniques in molecular imaging, anatomical imaging, and functional imaging and advancements in imaging biomarker generation. Developments have also been accelerated by efforts to realize precision medicine,3 which necessitates a multiscale understanding of diseases that integrate insights in areas such as radiology, pathology, and genetics. This focus issue highlights the growing impact of BMII, demonstrating the increasing breadth of imaging modalities (eg, optical, molecular, in addition to traditional diagnostic modalities) and the diversity of specialties that depend on imaging information (eg, dermatology, pathology, surgery).
Early efforts in BMII can be traced to the 1980s when the rise in radiological imaging techniques such as CT and MRI necessitated a digital, filmless approach to acquiring and interpreting images. The ability to acquire and distribute images electronically using picture archiving and communication systems (PACS) spawned a variety of applications aimed at improving radiological practice, research, and education. Imaging informatics efforts resulted in the development of specialized standardized …
Correspondence to Dr William Hsu, Medical Imaging Informatics (MII) Group, Department of Radiological Sciences, UCLA David Geffen School of Medicine, Los Angeles, CA 90024, USA; willhsu{at}mii.ucla.edu William Hsu, Mia K. Markey, May D. Wang |
J. Am. Medical Informatics Assoc. | 3 |
| 2013 | Review: Pathology imaging informatics for quantitative analysis of whole-slide imagesabstractOBJECTIVES: With the objective of bringing clinical decision support systems to reality, this article reviews histopathological whole-slide imaging informatics methods, associated challenges, and future research opportunities. TARGET AUDIENCE: This review targets pathologists and informaticians who have a limited understanding of the key aspects of whole-slide image (WSI) analysis and/or a limited knowledge of state-of-the-art technologies and analysis methods. SCOPE: First, we discuss the importance of imaging informatics in pathology and highlight the challenges posed by histopathological WSI. Next, we provide a thorough review of current methods for: quality control of histopathological images; feature extraction that captures image properties at the pixel, object, and semantic levels; predictive modeling that utilizes image features for diagnostic or prognostic applications; and data and information visualization that explores WSI for de novo discovery. In addition, we highlight future research directions and discuss the impact of large public repositories of histopathological data, such as the Cancer Genome Atlas, on the field of pathology informatics. Following the review, we present a case study to illustrate a clinical decision support system that begins with quality control and ends with predictive modeling for several cancer endpoints. Currently, state-of-the-art software tools only provide limited image processing capabilities instead of complete data analysis for clinical decision-making. We aim to inspire researchers to conduct more research in pathology imaging informatics so that clinical decision support can become a reality. Sonal Kothari, John H. Phan, Todd H. Stokes, May D. Wang |
J. Am. Medical Informatics Assoc. | 4 |
| 2013 | Multivariate Hypergeometric Similarity MeasureabstractWe propose a similarity measure based on the multivariate hypergeometric distribution for the pairwise comparison of images and data vectors. The formulation and performance of the proposed measure are compared with other similarity measures using synthetic data. A method of piecewise approximation is also implemented to facilitate application of the proposed measure to large samples. Example applications of the proposed similarity measure are presented using mass spectrometry imaging data and gene expression microarray data. Results from synthetic and biological data indicate that the proposed measure is capable of providing meaningful discrimination between samples, and that it can be a useful tool for identifying potentially related samples in large-scale biological data sets. Chanchala Kaddi, R. Mitchell Parry, May D. Wang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2012 | Reverse engineering biomolecular systems using -omic data: challenges, progress and opportunitiesabstractRecent advances in high-throughput biotechnologies have led to the rapid growing research interest in reverse engineering of biomolecular systems (REBMS). 'Data-driven' approaches, i.e. data mining, can be used to extract patterns from large volumes of biochemical data at molecular-level resolution while 'design-driven' approaches, i.e. systems modeling, can be used to simulate emergent system properties. Consequently, both data- and design-driven approaches applied to -omic data may lead to novel insights in reverse engineering biological systems that could not be expected before using low-throughput platforms. However, there exist several challenges in this fast growing field of reverse engineering biomolecular systems: (i) to integrate heterogeneous biochemical data for data mining, (ii) to combine top-down and bottom-up approaches for systems modeling and (iii) to validate system models experimentally. In addition to reviewing progress made by the community and opportunities encountered in addressing these challenges, we explore the emerging field of synthetic biology, which is an exciting approach to validate and analyze theoretical system models directly through experimental synthesis, i.e. analysis-by-synthesis. The ultimate goal is to address the present and future challenges in reverse engineering biomolecular systems (REBMS) using integrated workflow of data mining, systems modeling and synthetic biology. Chang F. Quo, Chanchala Kaddi, John H. Phan, Amin Zollanvari, Mingqing Xu, May D. Wang, Gil Alterovitz |
Briefings Bioinform. | 6 |
| 2012 | Win percentage: a novel measure for assessing the suitability of machine classifiers for biological problemsabstractBACKGROUND: Selecting an appropriate classifier for a particular biological application poses a difficult problem for researchers and practitioners alike. In particular, choosing a classifier depends heavily on the features selected. For high-throughput biomedical datasets, feature selection is often a preprocessing step that gives an unfair advantage to the classifiers built with the same modeling assumptions. In this paper, we seek classifiers that are suitable to a particular problem independent of feature selection. We propose a novel measure, called "win percentage", for assessing the suitability of machine classifiers to a particular problem. We define win percentage as the probability a classifier will perform better than its peers on a finite random sample of feature sets, giving each classifier equal opportunity to find suitable features. RESULTS: First, we illustrate the difficulty in evaluating classifiers after feature selection. We show that several classifiers can each perform statistically significantly better than their peers given the right feature set among the top 0.001% of all feature sets. We illustrate the utility of win percentage using synthetic data, and evaluate six classifiers in analyzing eight microarray datasets representing three diseases: breast cancer, multiple myeloma, and neuroblastoma. After initially using all Gaussian gene-pairs, we show that precise estimates of win percentage (within 1%) can be achieved using a smaller random sample of all feature pairs. We show that for these data no single classifier can be considered the best without knowing the feature set. Instead, win percentage captures the non-zero probability that each classifier will outperform its peers based on an empirical estimate of performance. CONCLUSIONS: Fundamentally, we illustrate that the selection of the most suitable classifier (i.e., one that is more likely to perform better than its peers) not only depends on the dataset and application but also on the thoroughness of feature selection. In particular, win percentage provides a single measurement that could assist users in eliminating or selecting classifiers for their particular application. R. Mitchell Parry, John H. Phan, May D. Wang |
BMC Bioinform. | 3 |
| 2012 | Cardiovascular Genomics: A Biomarker Identification PipelineabstractGenomic biomarkers are essential for understanding the underlying molecular basis of human diseases such as cardiovascular disease. In this review, we describe a biomarker identification pipeline for cardiovascular disease, which includes 1) high-throughput genomic data acquisition, 2) preprocessing and normalization of data, 3) exploratory analysis, 4) feature selection, 5) classification, and 6) interpretation and validation of candidate biomarkers. We review each step in the pipeline, presenting current and widely used bioinformatics methods. Furthermore, we analyze several publicly available cardiovascular genomics datasets to illustrate the pipeline. Finally, we summarize the current challenges and opportunities for further research. John H. Phan, Chang F. Quo, May D. Wang |
IEEE Trans. Inf. Technol. Biomed. | 3 |
| 2011 | Hypergeometric Similarity Measure for Spatial Analysis in Tissue Imaging Mass SpectrometryabstractTissue imaging mass spectrometry (TIMS) is a data-intensive technique for spatial biochemical analysis. TIMS contributes both molecular and spatial information to tissue analysis. We propose and evaluate a similarity measure, based on the hypergeometric distribution, for comparing m/z images from TIMS datasets, with the goal of identifying m/z values with similar spatial distributions. We compare the formulation and properties of the proposed method with those of other similarity measures, and examine the performance of each measure on synthetic and biological data. This study demonstrates that the proposed hypergeometric similarity measure is effective in identifying similar m/z images, and may be a useful addition to current methods in TIMS data analysis. Chanchala Kaddi, R. Mitchell Parry, May D. Wang |
BIBM | 3 |
| 2011 | Histological Image Feature Mining Reveals Emergent Diagnostic Properties for Renal CancerabstractComputer-aided histological image classification systems are important for making objective and timely cancer diagnostic decisions. These systems use combinations of image features that quantify a variety of image properties. Because researchers tend to validate their diagnostic systems on specific cancer endpoints, it is difficult to predict which image features will perform well given a new cancer endpoint. In this paper, we define a comprehensive set of common image features (consisting of 12 distinct feature subsets) that quantify a variety of image properties. We use a data-mining approach to determine which feature subsets and image properties emerge as part of an "optimal" diagnostic model when applied to specific cancer endpoints. Our goal is to assess the performance of such comprehensive image feature sets for application to a wide variety of diagnostic problems. We perform this study on 12 endpoints including 6 renal tumor subtype endpoints and 6 renal cancer grade endpoints. Keywords-histology, image mining, computer-aided diagnosis. Sonal Kothari, John H. Phan, Andrew N. Young, May D. Wang |
BIBM | 4 |
| 2011 | Biological Interpretation of Model-Reference Adaptive Control in a Mass Action Kinetics Metabolic Pathway ModelabstractThe usefulness of control theory to model robustness in metabolic pathways is limited because controller properties and their implications on pathway regulation are unclear. Using sphingolipid biosynthesis in response to single-gene overexpression as a case study, we apply model-reference adaptive control (MRAC) to model regulation in a mass action kinetics pathway model and report on its properties. Tracking error between treated cells (plant) and wild type (reference) is reduced in 9 of 10 system variables compared to using mass action kinetics only. This result is robust when system parameters are perturbed. Furthermore, we interpret control dynamics to infer potential regulatory interactions. Some observations are consistent with independent studies on the effects of the same experimental treatment, while others represent novel hypotheses that may be tested to yield additional biological insight. The usefulness and interpretation of MRAC to model metabolic pathway regulation is shown where plant dynamics approach the reference. Chang F. Quo, May D. Wang |
BIBM | 2 |
| 2011 | An Analysis of Scale and Rotation Invariance in the Bag-of-Features Method for Histopathological Image Classification
S. Hussain Raza 0001, R. Mitchell Parry, Richard A. Moffitt, Andrew N. Young, May D. Wang |
MICCAI (3) | 5 |
| 2011 | caCORRECT2: Improving the accuracy and reliability of microarray data in the presence of artifactsabstractBACKGROUND: In previous work, we reported the development of caCORRECT, a novel microarray quality control system built to identify and correct spatial artifacts commonly found on Affymetrix arrays. We have made recent improvements to caCORRECT, including the development of a model-based data-replacement strategy and integration with typical microarray workflows via caCORRECT's web portal and caBIG grid services. In this report, we demonstrate that caCORRECT improves the reproducibility and reliability of experimental results across several common Affymetrix microarray platforms. caCORRECT represents an advance over state-of-art quality control methods such as Harshlighting, and acts to improve gene expression calculation techniques such as PLIER, RMA and MAS5.0, because it incorporates spatial information into outlier detection as well as outlier information into probe normalization. The ability of caCORRECT to recover accurate gene expressions from low quality probe intensity data is assessed using a combination of real and synthetic artifacts with PCR follow-up confirmation and the affycomp spike in data. The caCORRECT tool can be accessed at the website: http://cacorrect.bme.gatech.edu. RESULTS: We demonstrate that (1) caCORRECT's artifact-aware normalization avoids the undesirable global data warping that happens when any damaged chips are processed without caCORRECT; (2) When used upstream of RMA, PLIER, or MAS5.0, the data imputation of caCORRECT generally improves the accuracy of microarray gene expression in the presence of artifacts more than using Harshlighting or not using any quality control; (3) Biomarkers selected from artifactual microarray data which have undergone the quality control procedures of caCORRECT are more likely to be reliable, as shown by both spike in and PCR validation experiments. Finally, we present a case study of the use of caCORRECT to reliably identify biomarkers for renal cell carcinoma, yielding two diagnostic biomarkers with potential clinical utility, PRKAB1 and NNMT. CONCLUSIONS: caCORRECT is shown to improve the accuracy of gene expression, and the reproducibility of experimental results in clinical application. This study suggests that caCORRECT will be useful to clean up possible artifacts in new as well as archived microarray data. Richard A. Moffitt, Qiqin Yin-Goen, Todd H. Stokes, R. Mitchell Parry, James H. Torrance, John H. Phan, Andrew N. Young, May D. Wang |
BMC Bioinform. | 8 |
| 2009 | Development of a High Resolution 3D Infant Stomach Model for Surgical Planning
Qaiser Chaudry, S. Hussain Raza 0001, Jeonggyu Lee, Mark Wulkan, May D. Wang |
CAIP | 6 |
| 2008 | Improving renal cell carcinoma classification by automatic region of interest selectionabstractIn this paper, we present an improved automated system for classification of pathological image data of renal cell carcinoma. The task of analyzing tissue biopsies, generally performed manually by expert pathologists, is extremely challenging due to the variability in the tissue morphology, the preparation of tissue specimen, and the image acquisition process. Due to the complexity of this task and heterogeneity of patient tissue, this process suffers from inter-observer and intra-observer variability. In continuation of our previous work, which proposed a knowledge-based automated system, we observe that real life clinical biopsy images which contain necrotic regions and glands significantly degrade the classification process. Following the pathologist's technique of focusing on selected region of interest (ROI), we propose a simple ROI selection process which automatically rejects the glands and necrotic regions thereby improving the classification accuracy. We were able to improve the classification accuracy from 90% to 95% on a significantly heterogeneous image data set using our technique. Qaiser Chaudry, S. Hussain Raza 0001, Yachna Sharma, Andrew N. Young, May D. Wang |
BIBE | 5 |
| 2008 | Using spiral intensity profile to quantify head and neck cancerabstractDuring the analysis of microscopy images, researchers locate regions of interest (ROI) and extract relevant information within it. Identifying the ROI is mostly done manually and subjectively by pathologists. Computer algorithms could help in reducing their workload and improve reproducibility. In particular, we want to assess the validity of the folic acid receptor as a biomarker for head and neck cancer. We are only interested in folic acid receptors appearing in cancerous tissue. Therefore, the first step is to segment images into cancerous and noncancerous regions. We propose to use a spiral intensity profile for segmentation of light microscopy images. Many algorithms identify objects in an image by considering pixel intensity and spatial information separately. Our algorithm integrates intensity and spatial information by considering the change, or profile, of pixel intensity in a spiral fashion. Using a spiral intensity profile can also perform segmentation at different scales from cancer regions to nuclei cluster to individual nuclei. We compared our algorithm with manually segmented image and obtained a specificity of 83.7% and sensitivity of 61.1%. Spiral intensity profiles can be used as a feature to improve other segmentation algorithms. Segmentation of cancerous images at different scales allows effective quantification of folic acid receptor inside cancerous regions, nuclei clusters, or individual cells. Koon Yin Kong, Yachna Sharma, S. Hussain Raza 0001, Susan Muller, May D. Wang |
BIBE | 6 |
| 2008 | Matrix factorization techniques for analysis of imaging mass spectrometry dataabstractImaging mass spectrometry is a method for understanding the molecular distribution in a two-dimensional sample. This method is effective for a wide range of molecules, but generates a large amount of data. It is difficult to extract important information from these large datasets manually and automated methods for discovering important spatial and spectral features are needed. Independent component analysis and non-negative matrix factorization are explained and explored as tools for identifying underlying factors in the data. These techniques are compared and contrasted with principle component analysis, the more standard analysis tool. Independent component analysis and non-negative matrix factorization are found to be more effective analysis methods. A mouse cerebellum dataset is used for testing. Peter W. Siy, Richard A. Moffitt, R. Mitchell Parry, Yanfeng Chen, M. Cameron Sullards, Alfred H. Merrill, May D. Wang |
BIBE | 8 |
| 2008 | Quantitative analysis of numerical solvers for oscillatory biomolecular system modelsabstractBACKGROUND: This article provides guidelines for selecting optimal numerical solvers for biomolecular system models. Because various parameters of the same system could have drastically different ranges from 10(-15) to 10(10), the ODEs can be stiff and ill-conditioned, resulting in non-unique, non-existing, or non-reproducible modeling solutions. Previous studies have not examined in depth how to best select numerical solvers for biomolecular system models, which makes it difficult to experimentally validate the modeling results. To address this problem, we have chosen one of the well-known stiff initial value problems with limit cycle behavior as a test-bed system model. Solving this model, we have illustrated that different answers may result from different numerical solvers. We use MATLAB numerical solvers because they are optimized and widely used by the modeling community. We have also conducted a systematic study of numerical solver performances by using qualitative and quantitative measures such as convergence, accuracy, and computational cost (i.e. in terms of function evaluation, partial derivative, LU decomposition, and "take-off" points). The results show that the modeling solutions can be drastically different using different numerical solvers. Thus, it is important to intelligently select numerical solvers when solving biomolecular system models. RESULTS: The classic Belousov-Zhabotinskii (BZ) reaction is described by the Oregonator model and is used as a case study. We report two guidelines in selecting optimal numerical solver(s) for stiff, complex oscillatory systems: (i) for problems with unknown parameters, ode45 is the optimal choice regardless of the relative error tolerance; (ii) for known stiff problems, both ode113 and ode15s are good choices under strict relative tolerance conditions. CONCLUSIONS: For any given biomolecular model, by building a library of numerical solvers with quantitative performance assessment metric, we show that it is possible to improve reliability of the analytical modeling, which in turn can improve the efficiency and effectiveness of experimental validations of these models. Also, our study can be extended to study a variety of molecular-level system models for human disease diagnosis and therapeutic treatment. Chang F. Quo, May D. Wang |
BMC Bioinform. | 2 |
| 2008 | ArrayWiki: an enabling technology for sharing public microarray data repositories and meta-analysesabstractBACKGROUND: A survey of microarray databases reveals that most of the repository contents and data models are heterogeneous (i.e., data obtained from different chip manufacturers), and that the repositories provide only basic biological keywords linking to PubMed. As a result, it is difficult to find datasets using research context or analysis parameters information beyond a few keywords. For example, to reduce the "curse-of-dimension" problem in microarray analysis, the number of samples is often increased by merging array data from different datasets. Knowing chip data parameters such as pre-processing steps (e.g., normalization, artefact removal, etc), and knowing any previous biological validation of the dataset is essential due to the heterogeneity of the data. However, most of the microarray repositories do not have meta-data information in the first place, and do not have a a mechanism to add or insert this information. Thus, there is a critical need to create "intelligent" microarray repositories that (1) enable update of meta-data with the raw array data, and (2) provide standardized archiving protocols to minimize bias from the raw data sources. RESULTS: To address the problems discussed, we have developed a community maintained system called ArrayWiki that unites disparate meta-data of microarray meta-experiments from multiple primary sources with four key features. First, ArrayWiki provides a user-friendly knowledge management interface in addition to a programmable interface using standards developed by Wikipedia. Second, ArrayWiki includes automated quality control processes (caCORRECT) and novel visualization methods (BioPNG, Gel Plots), which provide extra information about data quality unavailable in other microarray repositories. Third, it provides a user-curation capability through the familiar Wiki interface. Fourth, ArrayWiki provides users with simple text-based searches across all experiment meta-data, and exposes data to search engine crawlers (Semantic Agents) such as Google to further enhance data discovery. CONCLUSIONS: Microarray data and meta information in ArrayWiki are distributed and visualized using a novel and compact data storage format, BioPNG. Also, they are open to the research community for curation, modification, and contribution. By making a small investment of time to learn the syntax and structure common to all sites running MediaWiki software, domain scientists and practioners can all contribute to make better use of microarray technologies in research and medical practices. ArrayWiki is available at http://www.bio-miblab.org/arraywiki. Todd H. Stokes, J. T. Torrance, Henry Li, May D. Wang |
BMC Bioinform. | 4 |
| 2007 | Microtubule Dynamics Classification Using a Statistical Model of the Movement of Outer TipsabstractA new method is proposed for tracking the dynamics of microtubules. It combines a salient point extraction mechanism for segmenting plus-end tips, a robust tracking method capable of locating the trajectories of a large number of feature points, and a classification algorithm capable of determining if the level of activity of a given microtubule video is typical of that of a treated or a control cell. Our method does not rely on the precise tracking of a single microtubule like many previous works, but instead focuses on the generalized movement of ending tips as a whole, which gives a more statistically reliable interpretation of the movement of microtubules. The proposed algorithm is tested using twenty videos of breast cancer microtubules -ten are treated with Taxol and ten are control. We are able to correctly classify those test videos 85% of the time, which is comparable in accuracy, but uses a less complex algorithm than other algorithms. Christopher Alberti, Jean-Philippe Villaréal, Delano Billingsley, Koon Yin Kong, Adam I. Marcus, Paraskevi Giannakakou, May D. Wang |
BIBE | 7 |
| 2007 | Review of Systems Biology Simulation Tools for Translational ResearchabstractSystems biology models and simulation tools are critical components for bridging molecular biology with predictive medicine. We report a systematic comparison of popular simulation tools, including CellDesigner, COPASI and VirtualCell, to facilitate translational research in genomics, proteomics and systems biology. Different tools evaluating the same model may produce dissimilar results. This inconsistency is a roadblock to developing patient-customized disease progression models which reduce uncertainty in clinical decisions. We implement existing molecular-level SBML and CellML and compare simulation results with published data. Preliminary results suggest some tools perform better in terms of numerical stability to determine true model behavior. Furthermore, we uncover several worrying issues: (1) disparities between tools in terms of solver algorithms and language format, (2) lack of interactivity between users and tools, (3) lack of standardization for systems biology modeling languages and (4) need for models addressing specific pressing clinical objectives such as cancer disease progression. Melissa Freedenberg, Chanchala Kaddi, Chang F. Quo, May D. Wang |
BIBE | 4 |
| 2007 | Bio-Nano-Info Integration for Personalized MedicineabstractEvery disease has genetic and molecular basis. For example, in 2005, cancer became the number one killer in the USA for people under the age of 85. It is estimated that 1.3 million people will be diagnosed with cancer and more than 560,000 people will die each year. The underlying reasons for these statistics include the biological complexity of cancer as a disease which we are just now beginning to understand. The human genome project and other advanced technologies such as bionanotechnologies bring new hope to patient care because they can potentially address the disease on the molecular level. These new technologies, when linked to an individual patient's molecular profile, can provide personalized and predictive early detection, diagnosis, prognosis tracking, and novel targeted therapies. In concert with development of these advanced technologies we must develop ways to validate the novel biotechnologies; methods to analyze the high volume of data coming from genomics, molecular imaging, and bionanotechnologies; and methods to interpret these data and make relevant predictions for patient care. This workshop keynote lecture will focus on how linking molecular biology, with advances in nanotechnology and information technology, to speed up the discovery and development process and clinical translation that leads to the advances in patient care. Specifically, the workshop keynote lecture will cover topics in ontology, data mining, data management, and image analysis that enable biomarker-based diagnosis, molecular imaging probe design, and therapeutic development. Eric Jakobsson, May D. Wang, Linda K. Molnar |
BIBE | 2 |
| 2007 | Multivariate Analysis of Imaging Mass Spectrometry DataabstractImaging mass spectrometry can be used to reveal spatial distributions of multiple molecular species in a 2D biological sample. Due to the large amount of data produced by this technology, it is difficult and time-consuming to manually extract meaningful results from imaging mass spectrometry experimentation. We have developed and implemented an original approach to easily and consistently process mass spectrometry imaging data with the goal of automatically identifying interesting regions of molecule expression. Based on multivariate analysis techniques such as principal component analysis, the system allows researchers to conveniently define and visualize spatial regions based on spectral similarity. Features of our system are demonstrated on mouse cerebellum data. E. R. Muir, I. J. Ndiour, N. A. LeGoasduff, Richard A. Moffitt, M. Cameron Sullards, Alfred H. Merrill, Yanfeng Chen, May D. Wang |
BIBE | 9 |
| 2007 | Evolving Biological Behavior in Gene-Based Cellular SimulationsabstractCellular automata (CA) have long been capable of producing life-like behavior such as complexity, communication and self-replication using simple rules. Despite these properties, CA and other discrete simulations have failed to achieve real-world utility in cancer research or developmental biology, largely because they do not conform to a rules model which is understandable by clinicians and biologists. We present a method to generate CA with a desired phenotypic behavior within a biologically-based family of rule sets modeling simple gene regulation in a cell cycle signaling pathway. Designing CA within this biological context ensures the interpretability of any emergent results, thus opening the door for applications in biomedicine such as tumor growth and angiogenesis. Rule sets are encoded in intuitive genome structures, which are co-evolved using a Genetic Algorithm (GA) with a fitness function chosen to reward Wolfram's Type IV behavior. Results show the ability to generate interpretable type IV behavior in just a few hours on a desktop PC. This work is expected to have many applications including systems biology and cancer research. John H. Phan, Richard A. Moffitt, Todd H. Stokes, May D. Wang |
BIBE | 4 |
| 2007 | Estimating Classification Error to Identify Biomarkers in Time Series Expression DataabstractOne of the primary objectives in the study of human diseases is the development of accurate and early diagnostic tests using molecular profiling technology. These investigations usually focus on feature selection with the goal of building a classifier using only the most clinically relevant features. With time series molecular profiles, each patient's assay contains observations measured at several time points. Using traditional time series classification methods, we can only determine a patient's diagnosis after obtaining all time points, eliminating the possibility of early diagnosis. This problem can be alleviated by dividing the time series into smaller overlapping sub-series. Unfortunately, these sub-series are not independent and identically distributed (iid). Consequently, when we estimate classification error for feature selection using traditional methods, we may encounter estimation bias. In response, we have developed a novel method that ranks time series biomarkers using specialized blocked error estimation methods designed to reduce estimation bias. Our investigation applies special cross validation and bootstrap methods, including h-block, hv-block cross validation, and blocked bootstrap to synthetic and clinical time series data. Results indicate a clear decrease in estimation bias using these methods on synthetic time series data. Similar results for a drug treatment dataset show further evidence that these blocked algorithms can improve biomarker identification. John H. Phan, May D. Wang |
BIBE | 2 |
| 2007 | Computer Aided Histopathological Classification of Cancer SubtypesabstractIn this paper we present the results of our effort to develop a computer aided diagnosis system for pathological imaging data using renal cell carcinoma as a case study. Traditionally, cancer diagnosis is performed by an expert pathologist studying biopsy tissue under a microscope. Due to the complex nature of the task and the heterogeneity of patient tissue, these methods are not only time consuming but also suffer from subjective variability. To improve the repeatability and accuracy of the diagnosis process, a computational diagnosis system is proposed here. In this paper we report that with our novel knowledge-based methodology, we are able to achieve high level of classification accuracy (98%) when trying to classify 64 images (n=64) using a simple Bayesian classifier based on 8 extracted features and complete-leave-one-out cross-validation. This methodology is implemented in MATLAB and is expected to aid pathologists in the clinical setting to diagnose renal cell carcinoma as well as other types of cancer. Sohaib Waheed, Richard A. Moffitt, Qaiser Chaudry, Andrew N. Young, May D. Wang |
BIBE | 5 |
| 2007 | Using Particle Filter to Track and Model Microtubule DynamicsabstractWe propose to use particle filter [1], along with active contour [2] to track and model the plus-end tips of microtubules in confocal microscopy. Microtubules are polymers that change between states of growth, shortening, and pause. These events are critical to many cellular functions and are targets for successful cancer chemotherapy agents like Taxol. However, analyses are performed manually by researchers in most cases. Hence there is a need for a rapid and efficient quantification algorithm. In this paper, we propose to uses particle filter to track microtubule dynamics. While there are other algorithms that track microtubule movements, none of them uses inter-frame information. In our system, we use an open active contour to segment individual microtubule in each frame. Particle filter is used to track microtubule movements using information from previous frame. A simple motion and observation model is used to model the motion of microtubule movement. We show some of the results using MCF-7 breast cancer cell lines captured using fluorescent confocal microscopy and conclude that adding particle filter improves the accuracy of the system. Koon Yin Kong, Adam I. Marcus, Paraskevi Giannakakou, May D. Wang |
ICIP (5) | 4 |
| 2006 | Ceulular Imaging Data Analysis: Mircotubule Dynamics in Living CellabstractMicrotubules are dynamic polymers that rapidly transition between states of growth, shortening, and pause. These dynamic events are critical for studying cellular processes such as the cancer drug effectiveness study. Typically, these events are quantified by imaging microtubule movements over time, which results in large data sets that require rigorous quantitative analysis. In most cases, the analysis was performed manually by the researcher. This process is tedious and prone to error and becomes a bottleneck in modern cancer research. Thus, an efficient, reliable, and rapid quantification method is in critical need. In this paper, we describe open contour-based tracking methods to automatically segment and track microtubule movements. We redefine the internal energy terms specifically for open snake, and examine different external energy terms for locating the end points of a microtubule. This algorithm has been validated using simulated images, untreated MCF-7 breast cancer cell lines, and cells treated with the microtubule-targeting chemotherapeutic agent, Taxol. Koon Yin Kong, Adam I. Marcus, Jin Young Hong, Paraskevi Giannakakou, May D. Wang |
ICIP | 5 |
| 2004 | High speed processing of biomedical images using programmable gpuabstractIn this paper, we report our research results on high speed processing of large size biomedical images. The biomedical images usually contain various shapes of bioorganism. To accurately quantify these objects, shape-independent image processing techniques are needed. One of such techniques is level set (LS) method. However, its application to large size images is constrained by the extremely high computational cost. This is because a set of numerical simulations has to be performed repeatedly on every pixel of an image and the general-purpose central processing unit (CPU) has only one execution core and limited memory bandwidth. Thus, we researched for techniques that can perform a large number of iterative tasks effectively. As a result, we designed and developed a graphics processing unit (GPU) based level set (LS) algorithm, GPU-LS, to process large size of biomedical images. In this paper, we report how we perform LS processing utilizing features of advanced graphics hardware. Comparing CPU-LS and GPU-LS, we have achieved average 12-13 times increase in processing speed. Jin Young Hong, May D. Wang |
ICIP | 2 |
| 2003 | Development of Gene Ontology Tool for Biological Interpretation of Genomic and Proteomic Data
Weimin Feng, Geoffrey Wang, Barry Zeeberg, Kejiao Guo, Anthony T. Fojo, David W. Kane, William C. Reinhold, Samir Lababidi, John N. Weinstein, May D. Wang |
AMIA | 10 |
| 2003 | Development of A Biomedical Imaging Informatics System for Diagnosis and Treatment Planning
Geoffrey Wang, May D. Wang |
AMIA | 2 |
| 2003 | Volumetric Medical Image Compression and Reconstruction for Interactive Visualization in Surgical PlanningabstractSummary form only given. Different compression schemes that take advantage of interpolation methods for 3D volumetric reconstruction of medical imaging data are discussed. A volume composition system that uses different compression schemes combined with interpolation algorithms to facilitate the rapid visualization of volumetric images is developed. The quality and speed were compared to evaluate the best choice for volume rendering data compression by region of interest (ROI), discrete cosine transform (DCT) and ROI combined with DCT techniques. Abhirup Patra, May D. Wang |
DCC | 2 |