Georgia D. Tourassi

dblp:69/4501 · DBLP profile ↗
← Back
45ranked-venue papers
8as first author
10since 2021 · last 2026
0000-0002-9418-9638ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 30 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 17 · 3 first-authorDatabases, data management, data science and information retrieval · 5 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Position Paper: Artificial Intelligence in Medical Image Analysis: Advances, Clinical Translation, and Emerging Frontiers
abstract
Over the past five years, artificial intelligence (AI) has introduced new models and methods for addressing the challenges associated with the broader adoption of AI models and systems in medicine. This paper reviews recent advances in AI for medical image and video analysis, outlines emerging paradigms, highlights pathways for successful clinical translation, and provides recommendations for future work. Hybrid Convolutional Neural Network (CNN) Transformer architectures now deliver state-of-the-art results in segmentation, classification, reconstruction, synthesis, and registration. Foundation and generative AI models enable the use of transfer learning to smaller datasets with limited ground truth. Federated learning supports privacy-preserving collaboration across institutions. Explainable and trustworthy AI approaches have become essential to foster clinician trust, ensure regulatory compliance, and facilitate ethical deployment. Together, these developments pave the way for integrating AI into radiology, pathology, and wider healthcare workflows.
Andreas Panayides, Hao Chen 0011, Nenad Filipovic, Tijana Geroski, Junlin Hou, Karim Lekadir, Kostas Marias, George K. Matsopoulos, Giorgos Papanastasiou, Pinaki Sarder, Georgia D. Tourassi, Sotirios A. Tsaftaris, Huazhu Fu, Efthyvoulos C. Kyriacou, Christos P. Loizou, Michalis E. Zervakis, Joel H. Saltz, Farah Shamout, Ken C. L. Wong, Jianhua Yao 0001, Amir A. Amini, Dimitrios I. Fotiadis, Constantinos S. Pattichis, Marios S. Pattichis
IEEE J. Biomed. Health Informatics11
2025 Guest Editorial: Transforming Healthcare and Medicine With Biomedical Informatics and Emerging AI
Bobak Mortazavi, Yu-Chiao Chiu, Arun Das 0001, Georgia D. Tourassi, Björn M. Eskofier
IEEE J. Biomed. Health Informatics5
2024 Powering Progress in Leadership Computing in the Era of Generative AI and Energy Constraints
abstract
The advent of exascale computing has unlocked unprecedented opportunities for scientific discovery and technological advancement. As we push the boundaries of computational power, we find ourselves at the intersection of two critical challenges: harnessing the transformative potential of generative AI and navigating the growing demands on energy consumption. In this presentation I will describe the Oak Ridge National Laboratory’s journey to exascale computing, highlighting the remarkable achievements made possible by the Frontier supercomputing across various scientific domains. I will highlight the intricate interplay between large-scale modeling, simulation, and the growing field of generative AI, showcasing how these technologies can be seamlessly interwoven to tackle complex scientific problems and drive innovation. However, the energy consumption of these cutting-edge systems poses significant challenges that demand our attention and ingenuity. I will discuss the strategies and best practices we are implementing at the Oak Ridge Leadership Computing Facility to manage and optimize energy efficiency, ensuring the sustainability of our computing infrastructure while trying to solve the most pressing scientific and technical challenges facing humanity.
Georgia D. Tourassi
SIGSIM-PADS1
2024 Deep learning uncertainty quantification for clinical text classification
abstract
INTRODUCTION: Machine learning algorithms are expected to work side-by-side with humans in decision-making pipelines. Thus, the ability of classifiers to make reliable decisions is of paramount importance. Deep neural networks (DNNs) represent the state-of-the-art models to address real-world classification. Although the strength of activation in DNNs is often correlated with the network's confidence, in-depth analyses are needed to establish whether they are well calibrated. METHOD: In this paper, we demonstrate the use of DNN-based classification tools to benefit cancer registries by automating information extraction of disease at diagnosis and at surgery from electronic text pathology reports from the US National Cancer Institute (NCI) Surveillance, Epidemiology, and End Results (SEER) population-based cancer registries. In particular, we introduce multiple methods for selective classification to achieve a target level of accuracy on multiple classification tasks while minimizing the rejection amount-that is, the number of electronic pathology reports for which the model's predictions are unreliable. We evaluate the proposed methods by comparing our approach with the current in-house deep learning-based abstaining classifier. RESULTS: Overall, all the proposed selective classification methods effectively allow for achieving the targeted level of accuracy or higher in a trade-off analysis aimed to minimize the rejection rate. On in-distribution validation and holdout test data, with all the proposed methods, we achieve on all tasks the required target level of accuracy with a lower rejection rate than the deep abstaining classifier (DAC). Interpreting the results for the out-of-distribution test data is more complex; nevertheless, in this case as well, the rejection rate from the best among the proposed methods achieving 97% accuracy or higher is lower than the rejection rate based on the DAC. CONCLUSIONS: We show that although both approaches can flag those samples that should be manually reviewed and labeled by human annotators, the newly proposed methods retain a larger fraction and do so without retraining-thus offering a reduced computational cost compared with the in-house deep learning-based abstaining classifier.
Alina Peluso, Ioana Danciu, Hong-Jun Yoon, Jamaludin Mohd-Yusof, Tanmoy Bhattacharya 0001, Adam Spannaus, Noah Schaefferkoetter, Eric B. Durbin, Xiao-Cheng Wu, Antoinette Stroup, Jennifer A. Doherty, Stephen M. Schwartz, Charles Wiggins, Linda Coyle, Lynne Penberthy, Georgia D. Tourassi, Shang Gao 0008
J. Biomed. Informatics16
2023 Guest Editorial Advancing Biomedical Discovery and Healthcare Delivery Through Digital Technology
abstract
Digital technology has had a significant impact on biomedical sciences and healthcare delivery, not only in terms of theoretical and practical contributions to involved disciplines but also due to the profound and rapid changes in medical research and in medical care applications, as well as in the management organization after the recent COVID-19 epidemic. In this regard, several elements stand out from the 11 selected contributions published in this Special Issue.
Sergio Cerutti, Björn M. Eskofier, Georgia D. Tourassi
IEEE J. Biomed. Health Informatics3
2023 Evaluation of pre-training large language models on leadership-class supercomputers
Junqi Yin, Sajal Dash, John Gounley, Feiyi Wang, Georgia D. Tourassi
J. Supercomput.5
2022 Class imbalance in out-of-distribution datasets: Improving the robustness of the TextCNN for the classification of rare cancer types
abstract
In the last decade, the widespread adoption of electronic health record documentation has created huge opportunities for information mining. Natural language processing (NLP) techniques using machine and deep learning are becoming increasingly widespread for information extraction tasks from unstructured clinical notes. Disparities in performance when deploying machine learning models in the real world have recently received considerable attention. In the clinical NLP domain, the robustness of convolutional neural networks (CNNs) for classifying cancer pathology reports under natural distribution shifts remains understudied. In this research, we aim to quantify and improve the performance of the CNN for text classification on out-of-distribution (OOD) datasets resulting from the natural evolution of clinical text in pathology reports. We identified class imbalance due to different prevalence of cancer types as one of the sources of performance drop and analyzed the impact of previous methods for addressing class imbalance when deploying models in real-world domains. Our results show that our novel class-specialized ensemble technique outperforms other methods for the classification of rare cancer types in terms of macro F1 scores. We also found that traditional ensemble methods perform better in top classes, leading to higher micro F1 scores. Based on our findings, we formulate a series of recommendations for other ML practitioners on how to build robust models with extremely imbalanced datasets in biomedical NLP applications.
Kevin De Angeli, Shang Gao 0008, Ioana Danciu, Eric B. Durbin, Xiao-Cheng Wu, Antoinette Stroup, Jennifer A. Doherty, Stephen M. Schwartz, Charles Wiggins, Mark Damesyn, Linda Coyle, Lynne Penberthy, Georgia D. Tourassi, Hong-Jun Yoon
J. Biomed. Informatics13
2022 A Keyword-Enhanced Approach to Handle Class Imbalance in Clinical Text Classification
abstract
Recent applications ofdeep learning have shown promising results for classifying unstructured text in the healthcare domain. However, the reliability of models in production settings has been hindered by imbalanced data sets in which a small subset of the classes dominate. In the absence of adequate training data, rare classes necessitate additional model constraints for robust performance. Here, we present a strategy for incorporating short sequences of text (i.e. keywords) into training to boost model accuracy on rare classes. In our approach, we assemble a set of keywords, including short phrases, associated with each class. The keywords are then used as additional data during each batch of model training, resulting in a training loss that has contributions from both raw data and keywords. We evaluate our approach on classification of cancer pathology reports, which shows a substantial increase in model performance for rare classes. Furthermore, we analyze the impact of keywords on model output probabilities for bigrams, providing a straightforward method to identify model difficulties for limited training data.
Andrew E. Blanchard, Shang Gao 0008, Hong-Jun Yoon, James Blair Christian, Eric B. Durbin, Xiao-Cheng Wu, Antoinette Stroup, Jennifer A. Doherty, Stephen M. Schwartz, Charles Wiggins, Linda Coyle, Lynne Penberthy, Georgia D. Tourassi
IEEE J. Biomed. Health Informatics13
2021 Deep active learning for classifying cancer pathology reports
abstract
BACKGROUND: Automated text classification has many important applications in the clinical setting; however, obtaining labelled data for training machine learning and deep learning models is often difficult and expensive. Active learning techniques may mitigate this challenge by reducing the amount of labelled data required to effectively train a model. In this study, we analyze the effectiveness of 11 active learning algorithms on classifying subsite and histology from cancer pathology reports using a Convolutional Neural Network as the text classification model. RESULTS: We compare the performance of each active learning strategy using two differently sized datasets and two different classification tasks. Our results show that on all tasks and dataset sizes, all active learning strategies except diversity-sampling strategies outperformed random sampling, i.e., no active learning. On our large dataset (15K initial labelled samples, adding 15K additional labelled samples each iteration of active learning), there was no clear winner between the different active learning strategies. On our small dataset (1K initial labelled samples, adding 1K additional labelled samples each iteration of active learning), marginal and ratio uncertainty sampling performed better than all other active learning techniques. We found that compared to random sampling, active learning strongly helps performance on rare classes by focusing on underrepresented classes. CONCLUSIONS: Active learning can save annotation cost by helping human annotators efficiently and intelligently select which samples to label. Our results show that a dataset constructed using effective active learning techniques requires less than half the amount of labelled data to achieve the same performance as a dataset constructed using random sampling.
Kevin De Angeli, Shang Gao 0008, Mohammed M. Alawad, Hong-Jun Yoon, Noah Schaefferkoetter, Xiao-Cheng Wu, Eric B. Durbin, Jennifer A. Doherty, Antoinette Stroup, Linda Coyle, Lynne Penberthy, Georgia D. Tourassi
BMC Bioinform.12
2021 Limitations of Transformers on Clinical Text Classification
abstract
Bidirectional Encoder Representations from Transformers (BERT) and BERT-based approaches are the current state-of-the-art in many natural language processing (NLP) tasks; however, their application to document classification on long clinical texts is limited. In this work, we introduce four methods to scale BERT, which by default can only handle input sequences up to approximately 400 words long, to perform document classification on clinical texts several thousand words long. We compare these methods against two much simpler architectures - a word-level convolutional neural network and a hierarchical self-attention network - and show that BERT often cannot beat these simpler baselines when classifying MIMIC-III discharge summaries and SEER cancer pathology reports. In our analysis, we show that two key components of BERT - pretraining and WordPiece tokenization - may actually be inhibiting BERT's performance on clinical text classification tasks where the input document is several thousand words long and where correctly identifying labels may depend more on identifying a few key words or phrases rather than understanding the contextual meaning of sequences of text.
Shang Gao 0008, Mohammed M. Alawad, M. Todd Young, John Gounley, Noah Schaefferkoetter, Hong-Jun Yoon, Xiao-Cheng Wu, Eric B. Durbin, Jennifer A. Doherty, Antoinette Stroup, Linda Coyle, Georgia D. Tourassi
IEEE J. Biomed. Health Informatics12
2020 Automatic extraction of cancer registry reportable information from free-text pathology reports using multitask convolutional neural networks
abstract
OBJECTIVE: We implement 2 different multitask learning (MTL) techniques, hard parameter sharing and cross-stitch, to train a word-level convolutional neural network (CNN) specifically designed for automatic extraction of cancer data from unstructured text in pathology reports. We show the importance of learning related information extraction (IE) tasks leveraging shared representations across the tasks to achieve state-of-the-art performance in classification accuracy and computational efficiency. MATERIALS AND METHODS: Multitask CNN (MTCNN) attempts to tackle document information extraction by learning to extract multiple key cancer characteristics simultaneously. We trained our MTCNN to perform 5 information extraction tasks: (1) primary cancer site (65 classes), (2) laterality (4 classes), (3) behavior (3 classes), (4) histological type (63 classes), and (5) histological grade (5 classes). We evaluated the performance on a corpus of 95 231 pathology documents (71 223 unique tumors) obtained from the Louisiana Tumor Registry. We compared the performance of the MTCNN models against single-task CNN models and 2 traditional machine learning approaches, namely support vector machine (SVM) and random forest classifier (RFC). RESULTS: MTCNNs offered superior performance across all 5 tasks in terms of classification accuracy as compared with the other machine learning models. Based on retrospective evaluation, the hard parameter sharing and cross-stitch MTCNN models correctly classified 59.04% and 57.93% of the pathology reports respectively across all 5 tasks. The baseline models achieved 53.68% (CNN), 46.37% (RFC), and 36.75% (SVM). Based on prospective evaluation, the percentages of correctly classified cases across the 5 tasks were 60.11% (hard parameter sharing), 58.13% (cross-stitch), 51.30% (single-task CNN), 42.07% (RFC), and 35.16% (SVM). Moreover, hard parameter sharing MTCNNs outperformed the other models in computational efficiency by using about the same number of trainable parameters as a single-task CNN. CONCLUSIONS: The hard parameter sharing MTCNN offers superior classification accuracy for automated coding support of pathology documents across a wide range of cancers and multiple information extraction tasks while maintaining similar training and inference time as those of a single task-specific model.
Mohammed M. Alawad, Shang Gao 0008, John X. Qiu, Hong-Jun Yoon, James Blair Christian, Lynne Penberthy, Brent J. Mumphrey, Xiao-Cheng Wu, Linda Coyle, Georgia D. Tourassi
J. Am. Medical Informatics Assoc.10
2020 Accelerated training of bootstrap aggregation-based deep information extraction systems from cancer pathology reports
Hong-Jun Yoon, Hilda B. Klasky, John Gounley, Mohammed M. Alawad, Shang Gao 0008, Eric B. Durbin, Xiao-Cheng Wu, Antoinette Stroup, Jennifer A. Doherty, Linda Coyle, Lynne Penberthy, James Blair Christian, Georgia D. Tourassi
J. Biomed. Informatics13
2020 Knowledge Graph-Enabled Cancer Data Analytics
abstract
Cancer registries collect unstructured and structured cancer data for surveillance purposes which provide important insights regarding cancer characteristics, treatments, and outcomes. Cancer registry data typically (1) categorize each reportable cancer case or tumor at the time of diagnosis, (2) contain demographic information about the patient such as age, gender, and location at time of diagnosis, (3) include planned and completed primary treatment information, and (4) may contain survival outcomes. As structured data is being extracted from various unstructured sources, such as pathology reports, radiology reports, medical records, and stored for reporting and other needs, the associated information representing a reportable cancer is constantly expanding and evolving. While some popular analytic approaches including SEER*Stat and SAS exist, we provide a knowledge graph approach to organizing cancer registry data. Our approach offers unique advantages for timely data analysis and presentation and visualization of valuable information. This knowledge graph approach semantically enriches the data, and easily enables linking with third-party data which can help explain variation in cancer incidence patterns, disparities, and outcomes. We developed a prototype knowledge graph based on the Louisiana Tumor Registry dataset. We present the advantages of the knowledge graph approach by examining: i) scenario-specific queries, ii) links with openly available external datasets, iii) schema evolution for iterative analysis, and iv) data visualization. Our results demonstrate that this graph based solution can perform complex queries, improve query run-time performance by up to 76%, and more easily conduct iterative analyses to enhance researchers' understanding of cancer registry data.
S. M. Shamimul Hasan, Donna R. Rivera, Xiao-Cheng Wu, Eric B. Durbin, James Blair Christian, Georgia D. Tourassi
IEEE J. Biomed. Health Informatics6
2019 Adversarial Training for Privacy-Preserving Deep Learning Model Distribution
abstract
Collaboration among cancer registries is essential to develop accurate, robust, and generalizable deep learning models for automated information extraction from cancer pathology reports. Sharing data presents a serious privacy issue, especially in biomedical research and healthcare delivery domains. Distributing pretrained deep learning (DL) models has been proposed to avoid critical data sharing. However, there is growing recognition that collaboration among clinical institutes through DL model distribution exposes new security and privacy vulnerabilities. These vulnerabilities increase in natural language processing (NLP) applications, in which the dataset vocabulary with word vector representations needs to be associated with the other model parameters. In this paper, we propose a novel privacy-preserving DL model distribution across cancer registries for information extraction from cancer pathology reports with privacy and confidentiality considerations. The proposed approach exploits the adversarial training framework to distinguish private features from shared features among different datasets. It only shares registry-invariant model parameters, without sharing raw data nor registry-specific model parameters among cancer registries. Thus, it protects both the data and the trained model simultaneously. We compare our proposed approach to single-registry models, and a model trained on centrally hosted data from different cancer registries. The results show that the proposed approach significantly outperforms the single-registry models and achieves statistically indistinguishable micro and macro F1-score as compared to the centralized model.
Mohammed M. Alawad, Shang Gao 0008, Xiao-Cheng Wu, Eric B. Durbin, Linda Coyle, Lynne Penberthy, Georgia D. Tourassi
IEEE BigData7
2019 Information Extraction from Cancer Pathology Reports with Graph Convolution Networks for Natural Language Texts
abstract
Graph-of-words is a flexible and efficient text representation which addresses well-known challenges, such as word ordering and variation of expressions, to natural language processing. In this paper, we consider the latest graph-based convolutional neural network technique, the Text GraphConvolutional Network (Text GCN), in the context of performingclassification tasks on free-form natural language texts. To do this, we designed a study of multi-task information extraction from medical text documents. We implemented multi-task learning in the Text GCN, performed hyperparameter optimization, and measured the clinical task performances. We evaluated micro and macro-F1 scores of four information extraction tasks,including subsite, laterality, behavior, and histological grades from cancer pathology reports. The scores for the Text GCN significantly outperformed our previous studies with convolutional neural networks, suggesting that the Text GCN model is superior to traditional models in task performance.
Hong-Jun Yoon, John Gounley, M. Todd Young, Georgia D. Tourassi
IEEE BigData4
2019 Classifying cancer pathology reports with hierarchical self-attention networks
abstract
We introduce a deep learning architecture, hierarchical self-attention networks (HiSANs), designed for classifying pathology reports and show how its unique architecture leads to a new state-of-the-art in accuracy, faster training, and clear interpretability. We evaluate performance on a corpus of 374,899 pathology reports obtained from the National Cancer Institute's (NCI) Surveillance, Epidemiology, and End Results (SEER) program. Each pathology report is associated with five clinical classification tasks - site, laterality, behavior, histology, and grade. We compare the performance of the HiSAN against other machine learning and deep learning approaches commonly used on medical text data - Naive Bayes, logistic regression, convolutional neural networks, and hierarchical attention networks (the previous state-of-the-art). We show that HiSANs are superior to other machine learning and deep learning text classifiers in both accuracy and macro F-score across all five classification tasks. Compared to the previous state-of-the-art, hierarchical attention networks, HiSANs not only are an order of magnitude faster to train, but also achieve about 1% better relative accuracy and 5% better relative macro F-score.
Shang Gao 0008, John X. Qiu, Mohammed M. Alawad, Jacob D. Hinkle, Noah Schaefferkoetter, Hong-Jun Yoon, James Blair Christian, Paul A. Fearn, Lynne Penberthy, Xiao-Cheng Wu, Linda Coyle, Georgia D. Tourassi, Arvind Ramanathan
Artif. Intell. Medicine12
2019 Guest Editorial: AI Enabled Connected Health Informatics
abstract
The articles in this special section provide a snapshot of the latest research advancements in all aspects of connected health and informatic systems where artificial intelligence has been evident, including sensing, transfer, storage and analytics of biomedical data. These articles capture the end-to-end view of solutions that use automated informatics to address single ormultiple scenarios of health engineering such as primary care, preventive care, predictive technologies, hospitalization, home care, and occupational health.
Shuayb Zarar, Georgia D. Tourassi, Chris D. Nugent
IEEE J. Biomed. Health Informatics2
2018 Retrofitting Word Embeddings with the UMLS Metathesaurus for Clinical Information Extraction
abstract
Deep learning has surged in popularity and proven to be effective for various artificial intelligence applications including information extraction from cancer pathology reports. Since word representation is a core unit that enables deep learning algorithms to understand words and be able to perform NLP, this representation must include as much information as possible to help these algorithms achieve high classification performance. Therefore, in this work in addition to the distributional information of words in large sized corpora, we use UMLS vocabulary resources to enrich the vector space representation of words with the semantic relations between words. These resources provide many terminologies pertaining to cancer. The refined word embeddings are used with a convolutional neural (CNN) model to extract four data elements from cancer pathology reports; ICD-O-3 tumor topography codes, tumor laterality, behavior, and histological grade. We observed that using UMLS vocabulary resources to enrich word embeddings of CNN models consistently outperformed CNN models without pre-training word embeddings and even with pre-trained word embeddings on a domain specific corpus across all four tasks. The results show marginal improvement on the laterality task, but a significant improvement on the other tasks, especially for the macro-f score. Specifically, the improvements are 3%, 13%, and 15% for tumor site, histological grade, and behavior tasks, respectively. This approach is encouraging to enrich word embeddings with more clinical data resources to be used for information abstraction tasks from clinical pathology reports.
Mohammed M. Alawad, S. M. Shamimul Hasan, James Blair Christian, Georgia D. Tourassi
IEEE BigData4
2018 CAT: computer aided triage improving upon the Bayes risk through ε-refusal triage rules
abstract
BACKGROUND: Manual extraction of information from electronic pathology (epath) reports to populate the Surveillance, Epidemiology, and End Result (SEER) database is labor intensive. Systematizing the data extraction automatically using machine-learning (ML) and natural language processing (NLP) is desirable to reduce the human labor required to populate the SEER database and to improve the timeliness of the data. This enables scaling up registry efficiency and collection of new data elements. To ensure the integrity, quality, and continuity of the SEER data, the misclassification error of ML and NPL algorithms needs to be negligible. Current algorithms fail to achieve the precision of human experts who can bring additional information in their assessments. Differences in registry format and the desire to develop a common information extraction platform further complicate the ML/NLP tasks. The purpose of our study is to develop triage rules to partially automate registry workflow to improve the precision of the auto-extracted information. RESULTS: This paper presents a mathematical framework to improve the precision of a classifier beyond that of the Bayes classifier by selectively classifying item that are most likely to be correct. This results in a triage rule that only classifies a subset of the item. We characterize the optimal triage rule and demonstrate its usefulness in the problem of classifying cancer site from electronic pathology reports to achieve a desired precision. CONCLUSIONS: From the mathematical formalism, we propose a heuristic estimate for triage rule based on post-processing the soft-max output from standard machine learning algorithms. We show, in test cases, that the triage rule significantly improve the classification accuracy.
Nicolas W. Hengartner, Leticia Cuellar, Xiao-Cheng Wu, Georgia D. Tourassi, John X. Qiu, James Blair Christian, Tanmoy Bhattacharya 0001
BMC Bioinform.4
2018 Scalable deep text comprehension for Cancer surveillance on high-performance computing
abstract
BACKGROUND: Deep Learning (DL) has advanced the state-of-the-art capabilities in bioinformatics applications which has resulted in trends of increasingly sophisticated and computationally demanding models trained by larger and larger data sets. This vastly increased computational demand challenges the feasibility of conducting cutting-edge research. One solution is to distribute the vast computational workload across multiple computing cluster nodes with data parallelism algorithms. In this study, we used a High-Performance Computing environment and implemented the Downpour Stochastic Gradient Descent algorithm for data parallelism to train a Convolutional Neural Network (CNN) for the natural language processing task of information extraction from a massive dataset of cancer pathology reports. We evaluated the scalability improvements using data parallelism training and the Titan supercomputer at Oak Ridge Leadership Computing Facility. To evaluate scalability, we used different numbers of worker nodes and performed a set of experiments comparing the effects of different training batch sizes and optimizer functions. RESULTS: We found that Adadelta would consistently converge at a lower validation loss, though requiring over twice as many training epochs as the fastest converging optimizer, RMSProp. The Adam optimizer consistently achieved a close 2nd place minimum validation loss significantly faster; using a batch size of 16 and 32 allowed the network to converge in only 4.5 training epochs. CONCLUSIONS: We demonstrated that the networked training process is scalable across multiple compute nodes communicating with message passing interface while achieving higher classification accuracy compared to a traditional machine learning algorithm.
John X. Qiu, Hong-Jun Yoon, Kshitij Srivastava, Thomas P. Watson, James Blair Christian, Arvind Ramanathan, Xiao-Cheng Wu, Paul A. Fearn, Georgia D. Tourassi
BMC Bioinform.9
2018 Hierarchical attention networks for information extraction from cancer pathology reports
abstract
OBJECTIVE: We explored how a deep learning (DL) approach based on hierarchical attention networks (HANs) can improve model performance for multiple information extraction tasks from unstructured cancer pathology reports compared to conventional methods that do not sufficiently capture syntactic and semantic contexts from free-text documents. MATERIALS AND METHODS: Data for our analyses were obtained from 942 deidentified pathology reports collected by the National Cancer Institute Surveillance, Epidemiology, and End Results program. The HAN was implemented for 2 information extraction tasks: (1) primary site, matched to 12 International Classification of Diseases for Oncology topography codes (7 breast, 5 lung primary sites), and (2) histological grade classification, matched to G1-G4. Model performance metrics were compared to conventional machine learning (ML) approaches including naive Bayes, logistic regression, support vector machine, random forest, and extreme gradient boosting, and other DL models, including a recurrent neural network (RNN), a recurrent neural network with attention (RNN w/A), and a convolutional neural network. RESULTS: Our results demonstrate that for both information tasks, HAN performed significantly better compared to the conventional ML and DL techniques. In particular, across the 2 tasks, the mean micro and macro F-scores for the HAN with pretraining were (0.852,0.708), compared to naive Bayes (0.518, 0.213), logistic regression (0.682, 0.453), support vector machine (0.634, 0.434), random forest (0.698, 0.508), extreme gradient boosting (0.696, 0.522), RNN (0.505, 0.301), RNN w/A (0.637, 0.471), and convolutional neural network (0.714, 0.460). CONCLUSIONS: HAN-based DL models show promise in information abstraction tasks within unstructured clinical pathology reports.
Shang Gao 0008, Michael T. Young, John X. Qiu, Hong-Jun Yoon, James Blair Christian, Paul A. Fearn, Georgia D. Tourassi, Arvind Ramanathan
J. Am. Medical Informatics Assoc.7
2018 Deep Learning for Automated Extraction of Primary Sites From Cancer Pathology Reports
abstract
Pathology reports are a primary source of information for cancer registries which process high volumes of free-text reports annually. Information extraction and coding is a manual, labor-intensive process. In this study, we investigated deep learning and a convolutional neural network (CNN), for extracting ICD-O-3 topographic codes from a corpus of breast and lung cancer pathology reports. We performed two experiments, using a CNN and a more conventional term frequency vector approach, to assess the effects of class prevalence and inter-class transfer learning. The experiments were based on a set of 942 pathology reports with human expert annotations as the gold standard. CNN performance was compared against a more conventional term frequency vector space approach. We observed that the deep learning models consistently outperformed the conventional approaches in the class prevalence experiment, resulting in micro- and macro-F score increases of up to 0.132 and 0.226, respectively, when class labels were well populated. Specifically, the best performing CNN achieved a micro-F score of 0.722 over 12 ICD-O-3 topography codes. Transfer learning provided a consistent but modest performance boost for the deep learning methods but trends were contingent on the CNN method and cancer site. These encouraging results demonstrate the potential of deep learning for automated abstraction of pathology reports.
John X. Qiu, Hong-Jun Yoon, Paul A. Fearn, Georgia D. Tourassi
IEEE J. Biomed. Health Informatics4
2017 Leveraging Large-Scale Computing for Population Information Integration, Analysis, and Modeling
Jessica A. Boten, Donna R. Rivera, Madhumita Myneni, Georgia D. Tourassi, Tanmoy Bhattacharya 0001, Ana Paula de Oliveira Sales, Thomas S. Brettin, Paul A. Fearn, Lynne Penberthy
AMIA4
2017 Energy efficient stochastic-based deep spiking neural networks for sparse datasets
abstract
With large deep neural networks (DNNs) necessary to solve complex and data-intensive problems, energy efficiency is a key bottleneck for effectively deploying DL in the real world. Deep spiking NNs have gained much research attention recently due to the interest in building biological neural networks and the availability of neuromorphic platforms, which can be orders of magnitude more energy efficient compared to CPUs and GPUs. Although spiking NNs have proven to be an efficient technique for solving many machine learning and computer vision problems, to the best of our knowledge, this is the first attempt to adapt spiking NNs to sparse datasets. In this paper, we study the behaviour of spiking NNs in handling NLP datasets and the sparsity in their data representation. Then, we propose a novel framework for spiking NN using the concept of stochastic computing. Specifically, instead of generating spike trains with firing rates proportional to the intensity of each value in the feature set separately, the whole feature set is treated as a distribution function and a stochastic spiking train that follow this distribution is generated. This framework reduces the connectivity between NN layers from O(N) to O(log N). Also, it encodes input data differently and make suitable to handle sparse datasets. Finally, the framework achieves high energy efficiency since it uses Integrate and Fire neurons same as conventional spiking NNs. The results show that our proposed stochastic-based SNN achieves nearly the same accuracy as the original DNN on MNIST dataset, and it has better performance than state-of-the-art SNN. Besides that stochastic-based SNN is energy efficient, where the fully connected DNN, the conventional SNN, and the data normalized SNN consume 38.24, 1.83, and 1.85-times more energy than the stochastic-based SNN, respectively. For sparse datasets, including IMDb and In-House clinical datasets, stochastic-based SNN achieves performance comparable to that of the conventional DNN. However, the conventional spiking NN has a significant decline in classification accuracy.
Mohammed M. Alawad, Hong-Jun Yoon, Georgia D. Tourassi
IEEE BigData3
2017 Deep learning enabled national cancer surveillance
abstract
Pathology reports are a primary source of information for cancer registries which process high volumes of free-text reports annually. Information extraction and coding is a manual, labor-intensive process. In this talk I will discuss the latest deep learning technology, presenting both theoretical and practical perspectives that are relevant to natural language processing of clinical pathology reports. Using different deep learning architectures, I will present benchmark studies for various information extraction tasks and discuss their importance in supporting a comprehensive and scalable national cancer surveillance program.
Georgia D. Tourassi
IEEE BigData1
2016 The utility of web mining for epidemiological research: studying the association between parity and cancer risk
abstract
BACKGROUND: The World Wide Web has emerged as a powerful data source for epidemiological studies related to infectious disease surveillance. However, its potential for cancer-related epidemiological discoveries is largely unexplored. METHODS: Using advanced web crawling and tailored information extraction procedures, the authors automatically collected and analyzed the text content of 79 394 online obituary articles published between 1998 and 2014. The collected data included 51 911 cancer (27 330 breast; 9470 lung; 6496 pancreatic; 6342 ovarian; 2273 colon) and 27 483 non-cancer cases. With the derived information, the authors replicated a case-control study design to investigate the association between parity (i.e., childbearing) and cancer risk. Age-adjusted odds ratios (ORs) with 95% confidence intervals (CIs) were calculated for each cancer type and compared to those reported in large-scale epidemiological studies. RESULTS: Parity was found to be associated with a significantly reduced risk of breast cancer (OR = 0.78, 95% CI, 0.75-0.82), pancreatic cancer (OR = 0.78, 95% CI, 0.72-0.83), colon cancer (OR = 0.67, 95% CI, 0.60-0.74), and ovarian cancer (OR = 0.58, 95% CI, 0.54-0.62). Marginal association was found for lung cancer risk (OR = 0.87, 95% CI, 0.81-0.92). The linear trend between increased parity and reduced cancer risk was dramatically more pronounced for breast and ovarian cancer than the other cancers included in the analysis. CONCLUSION: This large web-mining study on parity and cancer risk produced findings very similar to those reported with traditional observational studies. It may be used as a promising strategy to generate study hypotheses for guiding and prioritizing future epidemiological studies.
Georgia D. Tourassi, Hong-Jun Yoon, Songhua Xu, Xuesong Han
J. Am. Medical Informatics Assoc.1
2016 A novel web informatics approach for automated surveillance of cancer mortality trends
Georgia D. Tourassi, Hong-Jun Yoon, Songhua Xu
J. Biomed. Informatics1
2014 Extracting Patient Demographics and Personal Medical Information from Online Health Forums
Yang Liu 0046, Songhua Xu, Hong-Jun Yoon, Georgia D. Tourassi
AMIA4
2014 A user-oriented web crawler for selectively acquiring online content in e-health research
abstract
MOTIVATION: Life stories of diseased and healthy individuals are abundantly available on the Internet. Collecting and mining such online content can offer many valuable insights into patients' physical and emotional states throughout the pre-diagnosis, diagnosis, treatment and post-treatment stages of the disease compared with those of healthy subjects. However, such content is widely dispersed across the web. Using traditional query-based search engines to manually collect relevant materials is rather labor intensive and often incomplete due to resource constraints in terms of human query composition and result parsing efforts. The alternative option, blindly crawling the whole web, has proven inefficient and unaffordable for e-health researchers. RESULTS: We propose a user-oriented web crawler that adaptively acquires user-desired content on the Internet to meet the specific online data source acquisition needs of e-health researchers. Experimental results on two cancer-related case studies show that the new crawler can substantially accelerate the acquisition of highly relevant online content compared with the existing state-of-the-art adaptive web crawling technology. For the breast cancer case study using the full training set, the new method achieves a cumulative precision between 74.7 and 79.4% after 5 h of execution till the end of the 20-h long crawling session as compared with the cumulative precision between 32.8 and 37.0% using the peer method for the same time period. For the lung cancer case study using the full training set, the new method achieves a cumulative precision between 56.7 and 61.2% after 5 h of execution till the end of the 20-h long crawling session as compared with the cumulative precision between 29.3 and 32.4% using the peer method. Using the reduced training set in the breast cancer case study, the cumulative precision of our method is between 44.6 and 54.9%, whereas the cumulative precision of the peer method is between 24.3 and 26.3%; for the lung cancer case study using the reduced training set, the cumulative precisions of our method and the peer method are, respectively, between 35.7 and 46.7% versus between 24.1 and 29.6%. These numbers clearly show a consistently superior accuracy of our method in discovering and acquiring user-desired online content for e-health research. AVAILABILITY AND IMPLEMENTATION: The implementation of our user-oriented web crawler is freely available to non-commercial users via the following Web site: http://bsec.ornl.gov/AdaptiveCrawler.shtml. The Web site provides a step-by-step guide on how to execute the web crawler implementation. In addition, the Web site provides the two study datasets including manually labeled ground truth, initial seeds and the crawling results reported in this article.
Songhua Xu, Hong-Jun Yoon, Georgia D. Tourassi
Bioinform.3
2013 Research and applications: Investigating the link between radiologists' gaze, diagnostic decision, and image content
abstract
OBJECTIVE: To investigate machine learning for linking image content, human perception, cognition, and error in the diagnostic interpretation of mammograms. METHODS: Gaze data and diagnostic decisions were collected from three breast imaging radiologists and three radiology residents who reviewed 20 screening mammograms while wearing a head-mounted eye-tracker. Image analysis was performed in mammographic regions that attracted radiologists' attention and in all abnormal regions. Machine learning algorithms were investigated to develop predictive models that link: (i) image content with gaze, (ii) image content and gaze with cognition, and (iii) image content, gaze, and cognition with diagnostic error. Both group-based and individualized models were explored. RESULTS: By pooling the data from all readers, machine learning produced highly accurate predictive models linking image content, gaze, and cognition. Potential linking of those with diagnostic error was also supported to some extent. Merging readers' gaze metrics and cognitive opinions with computer-extracted image features identified 59% of the readers' diagnostic errors while confirming 97.3% of their correct diagnoses. The readers' individual perceptual and cognitive behaviors could be adequately predicted by modeling the behavior of others. However, personalized tuning was in many cases beneficial for capturing more accurately individual behavior. CONCLUSIONS: There is clearly an interaction between radiologists' gaze, diagnostic decision, and image content which can be modeled with machine learning algorithms.
Georgia D. Tourassi, Sophie Voisin, Vincent C. Paquit, Elizabeth A. Krupinski
J. Am. Medical Informatics Assoc.1
2012 The effect of class imbalance on case selection for case-based classifiers: An empirical study in the context of medical decision support
Jordan M. Malof, Maciej A. Mazurowski, Georgia D. Tourassi
Neural Networks3
2011 Mutual information-based template matching scheme for detection of breast masses: From mammography to digital breast tomosynthesis
Maciej A. Mazurowski, Joseph Y. Lo, Brian P. Harrawood, Georgia D. Tourassi
J. Biomed. Informatics4
2009 The effect of class imbalance on case selection for case-based classifiers, with emphasis on computer-aided diagnosis systems
abstract
In this paper, the effect of class imbalance in the case base of a case-based classifier is investigated as it pertains to case base reduction and the resulting classifier performance. A k-nearest neighbor algorithm is used as a classifier and the random mutation hill climbing (RMHC) algorithm is used for case base reduction. The effects at various levels of positive class prevalence are tested in a binary classification problem. The results indicate that class imbalance is detrimental to both case base reduction and classifier performance. Selection with RMHC generally improves the classification performance regardless of the case base prevalence.
Jordan M. Malof, Maciej A. Mazurowski, Georgia D. Tourassi
IJCNN3
2009 Evaluating classifiers: Relation between area under the receiver operator characteristic curve and overall accuracy
abstract
In this study, we investigated the relation between two popular classifier performance measures: area under the receiver operator characteristic curve and overall accuracy. We also evaluated the impact of class imbalance and number of examples in test set on this relation. We perform a set of experiments in which we train multiple neural networks and test them in various, well controlled conditions. The experimental results show that given a large and balanced test set, increase in one performance measure is a very good indicator of increase in the other measure. Furthermore increasing the total number of examples, while keeping the positive class prevalence constant generally increases the correlation between the two measures. Our results also indicate that increasing the extent of class imbalance in the test set has a detrimental effect on this correlation.
Maciej A. Mazurowski, Georgia D. Tourassi
IJCNN2
2008 Training neural network classifiers for medical decision making: The effects of imbalanced datasets on classification performance
Maciej A. Mazurowski, Piotr A. Habas, Jacek M. Zurada, Joseph Y. Lo, Jay A. Baker, Georgia D. Tourassi
Neural Networks6
2007 Case-base reduction for a computer assisted breast cancer detection system using genetic algorithms
abstract
A knowledge-based computer assisted decision (KB-CAD) system is a case-based reasoning system previously proposed for breast cancer detection. Although it was demonstrated to be very effective for the diagnostic problem, it was also shown to be computationally expensive due to the use of mutual information between images as a similarity measure. Here, the authors propose to alleviate this drawback by reducing the case-base size. The problem is formalized and a genetic algorithm is utilized as an optimization tool. Appropriate for the problem representation and operators are presented and discussed. A clinically relevant index of the area under the receiver operator characteristic curve is used as a measure of the system performance during the optimization and testing stages. Experimental results show that application of the proposed method can significantly reduce the case-base size while the classification performance of the KB-CAD, in fact, increases.
Maciej A. Mazurowski, Piotr A. Habas, Georgia D. Tourassi, Jacek M. Zurada
IEEE Congress on Evolutionary Computation3
2007 Bilateral Breast Volume Asymmetry in Screening Mammograms as a Potential Marker of Breast Cancer: Preliminary Experience
abstract
The biological concept of bilateral symmetry as a marker of developmental stability and good health is well established. Although most individuals deviate slightly from perfect symmetry, humans are essentially considered bilaterally symmetrical. Studies have shown that if an individual is exposed to genetic mutations or environmental stresses, the homeostatic mechanisms that maintain symmetry of paired structures (such as breasts) tend to break down. Consequently, increased fluctuating asymmetry of paired structures could be an indicator of poor health. This preliminary study tested if bilateral morphological breast asymmetry in screening mammograms correlates with the presence of breast cancer. Following the biological definition of breast asymmetry in terms of volume, we applied automated computer algorithms for screening mammograms that segment the breast region and then measure each segmented breast's volume. These parameters were measured separately for each breast in each mammographic view (CC and MLO). Then, the normalized absolute differences of these parameters were investigated as measurements of fluctuating asymmetry (FA). Based on 268 cancer cases and 82 normal cases from the DDSM database, we observed that cancer patients demonstrate statistically significantly higher fluctuating asymmetry in their screening mammograms than patients with normal screening studies. Using an artificial neural network to combine FA measurements from both views along with the patient's age and breast parenchymal density resulted in an ROC area of 0.80plusmn0.03. These results suggest that bilateral breast volume asymmetry estimated in screening mammograms should be studied as a risk factor for breast cancer.
Nevine H. Eltonsy, Adel Said Elmaghraby, Georgia D. Tourassi
ICIP (5)3
2007 Impact of Low Class Prevalence on the Performance Evaluation of Neural Network Based Classifiers: Experimental Study in the Context of Computer-Assisted Medical Diagnosis
abstract
This paper presents an experimental study on the impact of low class prevalence on the neural network based classifier performance as measured using receiver operator characteristic (ROC) analysis. Two methods of dealing with the problem are investigated: oversampling and undersampling in the context of varying the class prevalence and the size of training datasets with uncorrelated and correlated features. The results show that the class imbalance can significantly decrease the classifier performance especially in the case of small training datasets. Furthermore, the oversampling method is shown to be more effective than the undersampling method in compensating the class imbalance. Statistically significant differences, however, are observed only in the cases with large total number of samples and very low prevalence.
Maciej A. Mazurowski, Piotr A. Habas, Georgia D. Tourassi, Jacek M. Zurada
IJCNN3
2007 Stacked Generalization in Computer-Assisted Decision Systems: Empirical Comparison of Data Handling Schemes
abstract
Computer-assisted decision (CAD) systems are becoming increasingly popular for the diagnostic interpretation of radiologic images. These CAD systems often involve the stacked generalization of several different decision models. Combining decision models is a common meta-analysis strategy to improve upon the diagnostic performance of each individual model. This study investigates how different data handling schemes may affect the performance evaluation of CAD systems that rely on stacked generalization. The study is based on a multistage CAD system for the detection of masses in screening mammograms. The CAD system consists of a series of knowledge-based modules that operate at Level 0 capturing morphological as well as multiscale textural information. Then, the knowledge-based predictions are combined with a Level 1 classifier. The study shows that a leave-one-out sampling scheme appears to be an effective and relatively unbiased strategy for the estimation of the overall performance of a CAD system that is based on stacked generalization. However, extra caution should be placed on the complexity of the Level 1 combiner. When the available dataset is relatively small, a relatively simple learning system such as a backpropagation neural network with very few hidden nodes is preferable to avoid optimistically biased estimates of diagnostic performance.
Georgia D. Tourassi, Jonathan L. Jesneck, Maciej A. Mazurowski, Piotr A. Habas
IJCNN1
2007 A Concentric Morphology Model for the Detection of Masses in Mammography
abstract
We propose a technique for the automated detection of malignant masses in screening mammography. The technique is based on the presence of concentric layers surrounding a focal area with suspicious morphological characteristics and low relative incidence in the breast region. Mammographic locations with high concentration of concentric layers with progressively lower average intensity are considered suspicious deviations from normal parenchyma. The multiple concentric layers (MCLs) technique was trained and tested using the craniocaudal views of 270 mammographic cases with biopsy proven malignant masses from the digital database of screening mammography. One-half of the available cases were used for optimizing the parameters of the detection algorithm. The remaining cases were used for testing. During testing, malignant masses were detected with 92%, 88%, and 81% sensitivity at 5.4, 2.4, and 0.6 false positive marks per image. Testing on 82 normal screening mammograms showed a false positive rate of 5.0, 1.7, and 0.2 marks per image at the previously reported operating points. Furthermore, additional evaluation on 135 benign cases produced a significantly lower detection rate for benign masses (61.6%, 58.3%, and 43.7% at 5.1, 2.8, and 1.2 false positives per image, respectively). Overall, MCL is a promising computer-assisted detection strategy for screening mammograms to identify malignant masses while maintaining the detection rate of benign masses considerably lower.
Nevine H. Eltonsy, Georgia D. Tourassi, Adel Said Elmaghraby
IEEE Trans. Medical Imaging2
2005 Estimation of generalized entropies with sample spacing
Mark P. Wachowiak, Renata Wachowiak-Smolíkova, Georgia D. Tourassi, Adel Said Elmaghraby
Pattern Anal. Appl.3
2005 Estimation of generalized entropies with sample spacing
Mark P. Wachowiak, Renata Wachowiak-Smolíkova, Georgia D. Tourassi, Adel Said Elmaghraby
Pattern Anal. Appl.3
2003 Self-organizing map for cluster analysis of a breast cancer database
Mia K. Markey, Joseph Y. Lo, Georgia D. Tourassi, Carey E. Floyd Jr.
Artif. Intell. Medicine3
2000 Fractal Texture Analysis of Perfusion Lung Scans
Georgia D. Tourassi, Erik D. Frederick, Neal F. Vittitoe, R. Edward Coleman
Comput. Biomed. Res.1
1999 A constraint satisfaction neural network for medical diagnosis
abstract
The objective of this study was to explore how a constraint satisfaction neural network (CSNN) can be used for medical diagnostic tasks. The study is based on a database of 500 patients who underwent breast biopsy at Duke University Medical Center due to suspicious mammographic findings. A CSNN was developed and evaluated to predict the biopsy result from the patient's mammographic findings. The diagnostic performance of the CSNN network was compared to a traditional backpropagation neural network and a case-based-reasoning algorithm by means of receiver operating characteristics analysis. The study demonstrates (i) how CSNNs can be applied to medical diagnostic tasks and, (ii) how they can be utilized to extract meaningful clinical information regarding underlying relationships among medical findings and associated diagnoses.
Georgia D. Tourassi, Carey E. Floyd Jr., Joseph Y. Lo
IJCNN1