EDBT 2026 Demo / reviewers in the wild / expert
Daniel L. Rubin
dblp:35/5410
· DBLP profile ↗
104ranked-venue papers
12as first author
19since 2021 · last 2026
0000-0001-5057-4369ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 88 · 11 first-author · 13 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 since 2021Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Extraction of distant recurrence sites for breast cancer patients from free-text clinical notes using large language models
Madhu Babu Sikha, Amara Tariq, Allison W. Kurian, Kevin C. Ward, Theresa H. M. Keegan, Daniel L. Rubin, Imon Banerjee |
J. Biomed. Informatics | 6 |
| 2024 | Mirrored X-Net: Joint classification and contrastive learning for weakly supervised GA segmentation in SD-OCT
Zexuan Ji, Xiao Ma 0011, Theodore Leng, Daniel L. Rubin, Qiang Chen 0004 |
Pattern Recognit. | 4 |
| 2023 | RaLEs: a Benchmark for Radiology Language EvaluationsabstractThe radiology report is the main form of communication between radiologists and other clinicians. Prior work in natural language processing in radiology reports has shown the value of developing methods tailored for individual tasks such as identifying reports with critical results or disease detection. Meanwhile, English and biomedical natural language understanding benchmarks such as the General Language Understanding and Evaluation as well as Biomedical Language Understanding and Reasoning Benchmark have motivated the development of models that can be easily adapted to address many tasks in those domains. Here, we characterize the radiology report as a distinct domain and introduce RaLEs, the Radiology Language Evaluations, as a benchmark for natural language understanding and generation in radiology. RaLEs is comprised of seven natural language understanding and generation evaluations including the extraction of anatomical and disease entities and their relations, procedure selection, and report summarization. We characterize the performance of models designed for the general, biomedical, clinical and radiology domains across these tasks. We find that advances in the general and biomedical domains do not necessarily translate to radiology, and that improved models from the general domain can perform comparably to smaller clinical-specific models. The limited performance of existing pre-trained models on RaLEs highlights the opportunity to improve domain-specific self-supervised models for natural language processing in radiology. We propose RaLEs as a benchmark to promote and track the development of such domain-specific radiology language models. Juanma Zambrano Chaves, Nandita Bhaskhar, Maayane Attias, Jean-Benoit Delbrouck, Daniel L. Rubin, Andreas M. Loening, Curt Langlotz, Akshay Chaudhari |
NeurIPS | 5 |
| 2023 | Semi-Supervised Learning for Sparsely-Labeled Sequential Data: Application to Healthcare Video ProcessingabstractLabeled data is a critical resource for training and evaluating machine learning models. However, many real-life datasets are only partially labeled. We propose a semi-supervised machine learning training strategy to improve event detection performance on sequential data, such as video recordings, when only sparse labels are available, such as event start times without their corresponding end times. Our method uses noisy guesses of the events’ end times to train event detection models. Depending on how conservative these guesses are, mislabeled samples may be introduced into the training set. We further propose a mathematical model for explaining and estimating the evolution of the classification performance for increasingly noisier end time estimates. We show that neural networks can improve their detection performance by leveraging more training data with less conservative approximations despite the higher proportion of incorrect labels. We adapt sequential versions of CIFAR-10 and MNIST, and use the Berkeley MHAD and HMBD51 video datasets to empirically evaluate our method, and find that our risk-tolerant strategy outperforms conservative estimates by 3.5 points of mean average precision for CIFAR, 30 points for MNIST, 3 points for MHAD, and 14 points for HMBD51. Then, we leverage the proposed training strategy to tackle a real-life application: processing continuous video recordings of epilepsy patients, and show that our method outperforms baseline labeling methods by 17 points of average precision, and reaches a classification performance similar to that of fully supervised models. We share part of the code for this article at the following repository: fpgdubost/CIFAR-10-Sparsely-Labeled-Sequential-Data. Florian Dubost, Erin Hong, Siyi Tang, Nandita Bhaskhar, Christopher Lee-Messer, Daniel L. Rubin |
WACV | 6 |
| 2023 | ATCON: Attention Consistency for Vision ModelsabstractAttention–or attribution–maps methods are methods designed to highlight regions of the model’s input that were discriminative for its predictions. However, different attention maps methods can highlight different regions of the input, with sometimes contradictory explanations for a prediction. This effect is exacerbated when the training set is small. This indicates that either the model learned incorrect representations or that the attention maps methods did not accurately estimate the model’s representations. We propose an unsupervised fine-tuning method that optimizes the consistency of attention maps and show that it improves both classification performance and the quality of attention maps. We propose an implementation for two state-of-the-art attention computation methods, Grad-CAM and Guided Backpropagation, which relies on an input masking technique. We also show results on Grad-CAM and Integrated Gradients in an ablation study. We evaluate this method on our own dataset of event detection in continuous video recordings of hospital patients aggregated and curated for this work. As a sanity check, we also evaluate the proposed method on PASCAL VOC and SVHN. With the proposed method, with small training sets, we achieve a 6.6 points lift of F1 score over the baselines on our video dataset, a 2.9 point lift of F1 score on PASCAL, and a 1.8 points lift of mean Intersection over Union over Grad-CAM for weakly supervised detection on PASCAL. Those improved attention maps may help clinicians better understand vision model predictions and ease the deployment of machine learning systems into clinical care. We share part of the code for this article at the following repository: https://github.com/alimirzazadeh/SemisupervisedAttention. Ali Mirzazadeh, Florian Dubost, Maxwell Pike, Krish Maniar, Max Zuo, Christopher Lee-Messer, Daniel L. Rubin |
WACV | 7 |
| 2023 | Clinical outcome prediction using observational supervision with electronic health records and audit logs
Nandita Bhaskhar, Wui Ip, Jonathan H. Chen, Daniel L. Rubin |
J. Biomed. Informatics | 4 |
| 2023 | Predicting 30-Day All-Cause Hospital Readmission Using Multimodal Spatiotemporal Graph Neural NetworksabstractReduction in 30-day readmission rate is an important quality factor for hospitals as it can reduce the overall cost of care and improve patient post-discharge outcomes. While deep-learning-based studies have shown promising empirical results, several limitations exist in prior models for hospital readmission prediction, such as: (a) only patients with certain conditions are considered, (b) do not leverage data temporality, (c) individual admissions are assumed independent of each other, which ignores patient similarity, (d) limited to single modality or single center data. In this study, we propose a multimodal, spatiotemporal graph neural network (MM-STGNN) for prediction of 30-day all-cause hospital readmission, which fuses in-patient multimodal, longitudinal data and models patient similarity using a graph. Using longitudinal chest radiographs and electronic health records from two independent centers, we show that MM-STGNN achieved an area under the receiver operating characteristic curve (AUROC) of 0.79 on both datasets. Furthermore, MM-STGNN significantly outperformed the current clinical reference standard, LACE+ (AUROC = 0.61), on the internal dataset. For subset populations of patients with heart disease, our model significantly outperformed baselines, such as gradient-boosting and Long Short-Term Memory models (e.g., AUROC improved by 3.7 points in patients with heart disease). Qualitative interpretability analysis indicated that while patients' primary diagnoses were not explicitly used to train the model, features crucial for model prediction may reflect patients' diagnoses. Our model could be utilized as an additional clinical decision aid during discharge disposition and triaging high-risk patients for closer post-discharge follow-up for potential preventive measures. Siyi Tang, Amara Tariq, Jared Dunnmon, Praneetha Elugunti, Daniel L. Rubin, Bhavik N. Patel, Imon Banerjee |
IEEE J. Biomed. Health Informatics | 6 |
| 2023 | Label-Efficient Self-Supervised Federated Learning for Tackling Data Heterogeneity in Medical ImagingabstractThe collection and curation of large-scale medical datasets from multiple institutions is essential for training accurate deep learning models, but privacy concerns often hinder data sharing. Federated learning (FL) is a promising solution that enables privacy-preserving collaborative learning among different institutions, but it generally suffers from performance deterioration due to heterogeneous data distributions and a lack of quality labeled data. In this paper, we present a robust and label-efficient self-supervised FL framework for medical image analysis. Our method introduces a novel Transformer-based self-supervised pre-training paradigm that pre-trains models directly on decentralized target task datasets using masked image modeling, to facilitate more robust representation learning on heterogeneous data and effective knowledge transfer to downstream models. Extensive empirical results on simulated and real-world medical imaging non-IID federated datasets show that masked image modeling with Transformers significantly improves the robustness of models against various degrees of data heterogeneity. Notably, under severe data heterogeneity, our method, without relying on any additional pre-training data, achieves an improvement of 5.06%, 1.53% and 4.58% in test accuracy on retinal, dermatology and chest X-ray classification compared to the supervised baseline with ImageNet pre-training. In addition, we show that our federated self-supervised pre-training methods yield models that generalize better to out-of-distribution data and perform more effectively when fine-tuning with limited labeled data, compared to existing FL algorithms. The code is available at https://github.com/rui-yan/SSL-FL. Liangqiong Qu, Qingyue Wei, Shih-Cheng Huang, Liyue Shen, Daniel L. Rubin, Lei Xing 0001, Yuyin Zhou |
IEEE Trans. Medical Imaging | 6 |
| 2022 | Graph-based Fusion Modeling and Explanation for Disease Trajectory Prediction
Amara Tariq, Siyi Tang, Hifza Sakhi, Leo A. Celi, Janice M. Newsome, Daniel L. Rubin, Hari Trivedi, Judy Gichoya, Bhavik N. Patel, Imon Banerjee |
AMIA | 6 |
| 2022 | Rethinking Architecture Design for Tackling Data Heterogeneity in Federated LearningabstractFederated learning is an emerging research paradigm enabling collaborative training of machine learning models among different organizations while keeping data private at each institution. Despite recent progress, there remain fundamental challenges such as the lack of convergence and the potential for catastrophic forgetting across real-world heterogeneous devices. In this paper, we demonstrate that self-attention-based architectures (e.g., Transformers) are more robust to distribution shifts and hence improve federated learning over heterogeneous data. Concretely, we conduct the first rigorous empirical investigation of different neural architectures across a range of federated algorithms, real-world benchmarks, and heterogeneous data splits. Our experiments show that simply replacing convolutional networks with Transformers can greatly reduce catastrophic forgetting of previous devices, accelerate convergence, and reach a better global model, especially when dealing with heterogeneous data. We release our code and pretrained models to encourage future exploration in robust architectures as an alternative to current research efforts on the optimization front. Liangqiong Qu, Yuyin Zhou, Paul Pu Liang, Yingda Xia, Ehsan Adeli-Mosabbeb, Li Fei-Fei 0001, Daniel L. Rubin |
CVPR | 8 |
| 2022 | Self-Supervised Graph Neural Networks for Improved Electroencephalographic Seizure Analysis
Siyi Tang, Jared Dunnmon, Khaled Saab 0002, Qianying Huang, Florian Dubost, Daniel L. Rubin, Christopher Lee-Messer |
ICLR | 7 |
| 2022 | Opportunistic Incidence Prediction of Multiple Chronic Diseases from Abdominal CT Imaging Using Multi-task Learning
Louis Blankemeier, Isabel Gallegos, Juan Manuel Zambrano Chaves, David J. Maron, Alexander T. Sandhu, Fátima Rodriguez, Daniel L. Rubin, Bhavik N. Patel, Marc H. Willis, Robert D. Boutin, Akshay Chaudhari |
MICCAI (8) | 7 |
| 2022 | A weakly supervised model for the automated detection of adverse events using clinical notes
Josh Sanyal, Daniel L. Rubin, Imon Banerjee |
J. Biomed. Informatics | 2 |
| 2022 | SplitAVG: A Heterogeneity-Aware Federated Deep Learning Method for Medical ImagingabstractFederated learning is an emerging research paradigm for enabling collaboratively training deep learning models without sharing patient data. However, the data from different institutions are usually heterogeneous across institutions, which may reduce the performance of models trained using federated learning. In this study, we propose a novel heterogeneity-aware federated learning method, SplitAVG, to overcome the performance drops from data heterogeneity in federated learning. Unlike previous federated methods that require complex heuristic training or hyper parameter tuning, our SplitAVG leverages the simple network split and feature map concatenation strategies to encourage the federated model training an unbiased estimator of the target data distribution. We compare SplitAVG with seven state-of-the-art federated learning methods, using centrally hosted training data as the baseline on a suite of both synthetic and real-world federated datasets. We find that the performance of models trained using all the comparison federated learning methods degraded significantly with the increasing degrees of data heterogeneity. In contrast, SplitAVG method achieves comparable results to the baseline method under all heterogeneous settings, that it achieves 96.2% of the accuracy and 110.4% of the mean absolute error obtained by the baseline in a diabetic retinopathy binary classification dataset and a bone age prediction dataset, respectively, on highly heterogeneous data partitions. We conclude that SplitAVG method can effectively overcome the performance drops from variability in data distributions across institutions. Experimental results also show that SplitAVG can be adapted to different base convolutional neural networks (CNNs) and generalized to various types of medical imaging tasks. The code is publicly available at https://github.com/zm17943/SplitAVG. Miao Zhang 0030, Liangqiong Qu, Praveer Singh, Jayashree Kalpathy-Cramer, Daniel L. Rubin |
IEEE J. Biomed. Health Informatics | 5 |
| 2021 | A Fusion NLP Model for the Inference of Standardized Thyroid Nodule Malignancy Scores from Radiology Report Text
Thiago Santos, Omar Kallas, Janice M. Newsome, Daniel L. Rubin, Judy Gichoya, Imon Banerjee |
AMIA | 4 |
| 2021 | Observational Supervision for Medical Image Classification Using Gaze Data
Khaled Saab 0002, Sarah M. Hooper, Nimit Sharad Sohoni, Jupinder Parmar, Brian Pogatchnik, Sen Wu 0002, Jared Dunnmon, Hongyang R. Zhang, Daniel L. Rubin, Christopher Ré |
MICCAI (2) | 9 |
| 2021 | Comparison of segmentation-free and segmentation-dependent computer-aided diagnosis of breast masses on a public mammography datasetabstractPURPOSE: To compare machine learning methods for classifying mass lesions on mammography images that use predefined image features computed over lesion segmentations to those that leverage segmentation-free representation learning on a standard, public evaluation dataset. METHODS: We apply several classification algorithms to the public Curated Breast Imaging Subset of the Digital Database for Screening Mammography (CBIS-DDSM), in which each image contains a mass lesion. Segmentation-free representation learning techniques for classifying lesions as benign or malignant include both a Bag-of-Visual-Words (BoVW) method and a Convolutional Neural Network (CNN). We compare classification performance of these techniques to that obtained using two different segmentation-dependent approaches from the literature that rely on specific combinations of end classifiers (e.g. linear discriminant analysis, neural networks) and predefined features computed over the lesion segmentation (e.g. spiculation measure, morphological characteristics, intensity metrics). RESULTS: values of 0.73 for a segmentation-free BoVW method, 0.86 for a segmentation-free CNN method, 0.75 for a segmentation-dependent linear discriminant analysis of Rubber-Band Straightening Transform features, and 0.58 for a hybrid rule-based neural network classification using a small number of hand-designed features. CONCLUSIONS: We find that malignancy classification performance on the CBIS-DDSM dataset using segmentation-free BoVW features is comparable to that of the best segmentation-dependent methods we study, but also observe that a common segmentation-free CNN model substantially and significantly outperforms each of these (p < 0.05). These results reinforce recent findings suggesting that representation learning techniques such as BoVW and CNNs are advantageous for mammogram analysis because they do not require lesion segmentation, the quality and specific characteristics of which can vary substantially across datasets. We further observe that segmentation-dependent methods achieve performance levels on CBIS-DDSM inferior to those achieved on the original evaluation datasets reported in the literature. Each of these findings reinforces the need for standardization of datasets, segmentation techniques, and model implementations in performance assessments of automated classifiers for medical imaging. Rebecca Sawyer Lee, Jared Dunnmon, Ann He, Siyi Tang, Christopher Ré, Daniel L. Rubin |
J. Biomed. Informatics | 6 |
| 2021 | An integrated time adaptive geographic atrophy prediction model for SD-OCT images
Yuhan Zhang 0001, Zexuan Ji, Sijie Niu, Theodore Leng, Daniel L. Rubin, Songtao Yuan, Qiang Chen 0004 |
Medical Image Anal. | 6 |
| 2021 | Learning Domain-Agnostic Visual Representation for Computational Pathology Using Medically-Irrelevant Style Transfer AugmentationabstractSuboptimal generalization of machine learning models on unseen data is a key challenge which hampers the clinical applicability of such models to medical imaging. Although various methods such as domain adaptation and domain generalization have evolved to combat this challenge, learning robust and generalizable representations is core to medical image understanding, and continues to be a problem. Here, we propose STRAP (Style TRansfer Augmentation for histoPathology), a form of data augmentation based on random style transfer from non-medical style sources such as artistic paintings, for learning domain-agnostic visual representations in computational pathology. Style transfer replaces the low-level texture content of an image with the uninformative style of randomly selected style source image, while preserving the original high-level semantic content. This improves robustness to domain shift and can be used as a simple yet powerful tool for learning domain-agnostic representations. We demonstrate that STRAP leads to state-of-the-art performance, particularly in the presence of domain shifts, on two particular classification tasks in computational pathology. Our code is available at https://github.com/rikiyay/style-transfer-for-digital-pathology. Rikiya Yamashita, Jin Long, Snikitha Banda, Jeanne Shen, Daniel L. Rubin |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Cancer Treatment Classification with Electronic Medical Health Records (Student Abstract)abstractWe built a natural language processing (NLP) language model that can be used to extract cancer treatment information using structured and unstructured electronic medical records (EMR). Our work appears to be the first that combines EMR and NLP for treatment identification. Jiaming Zeng, Imon Banerjee, Michael Gensheimer 0001, Daniel L. Rubin |
AAAI | 4 |
| 2020 | Accounting for data variability in multi-institutional distributed deep learning for medical imagingabstractOBJECTIVES: Sharing patient data across institutions to train generalizable deep learning models is challenging due to regulatory and technical hurdles. Distributed learning, where model weights are shared instead of patient data, presents an attractive alternative. Cyclical weight transfer (CWT) has recently been demonstrated as an effective distributed learning method for medical imaging with homogeneous data across institutions. In this study, we optimize CWT to overcome performance losses from variability in training sample sizes and label distributions across institutions. MATERIALS AND METHODS: Optimizations included proportional local training iterations, cyclical learning rate, locally weighted minibatch sampling, and cyclically weighted loss. We evaluated our optimizations on simulated distributed diabetic retinopathy detection and chest radiograph classification. RESULTS: Proportional local training iteration mitigated performance losses from sample size variability, achieving 98.6% of the accuracy attained by centrally hosting in the diabetic retinopathy dataset split with highest sample size variance across institutions. Locally weighted minibatch sampling and cyclically weighted loss both mitigated performance losses from label distribution variability, achieving 98.6% and 99.1%, respectively, of the accuracy attained by centrally hosting in the diabetic retinopathy dataset split with highest label distribution variability across institutions. DISCUSSION: Our optimizations to CWT improve its capability of handling data variability across institutions. Compared to CWT without optimizations, CWT with optimizations achieved performance significantly closer to performance from centrally hosting. CONCLUSION: Our work is the first to identify and address challenges of sample size and label distribution variability in simulated distributed deep learning for medical imaging. Future work is needed to address other sources of real-world data variability. Niranjan Balachandar, Ken Chang, Jayashree Kalpathy-Cramer, Daniel L. Rubin |
J. Am. Medical Informatics Assoc. | 4 |
| 2020 | Corrigendum to: Accounting for data variability in multi-institutional distributed deep learning for medical imagingabstractJournal of the American Medical Informatics Association, 27(5), 2020, 700–708; doi: 10.1093/jamia/ocaa017 Jayashree Kalpathy-Cramer and Daniel L Rubin have been added to the list of contributing authors. Niranjan Balachandar, Ken Chang, Jayashree Kalpathy-Cramer, Daniel L. Rubin |
J. Am. Medical Informatics Assoc. | 4 |
| 2020 | Natural Language Generation Model for Mammography Reports SimulationabstractExtending the size of labeled corpora of medical reports is a major step towards a successful training of machine learning algorithms. Simulating new text reports is a key solution for reports augmentation, which extends the cohort size. However, text generation in the medical domain is challenging because it needs to preserve both content and style that are typical for real reports, without risking the patients' privacy. In this paper, we present a conditioned LSTM-RNN architecture for simulating realistic mammography reports. We evaluated the performance by analyzing the characteristics of the simulated reports and classifying them into benign and malignant classes. An average classification AUC was calculated over two distinct test sets. A qualitative analysis was also performed in which a masked radiologist classified 0.75 of the simulated reports as real reports, showing that both the style and content of the simulated reports were similar to real reports. Finally, we compared our RNN-LSTM generative model with Markov Random Fields. The RNN-LSTM provided significantly better and more stable performance than MRFs ( , Wilcoxon). Assaf Hoogi, Arjun Mishra, Francisco Gimenez, Jeffrey Dong, Daniel L. Rubin |
IEEE J. Biomed. Health Informatics | 5 |
| 2020 | MS-CAM: Multi-Scale Class Activation Maps for Weakly-Supervised Segmentation of Geographic Atrophy Lesions in SD-OCT ImagesabstractAs one of the most critical characteristics in advanced stage of non-exudative Age-related Macular Degeneration (AMD), Geographic Atrophy (GA) is one of the significant causes of sustained visual acuity loss. Automatic localization of retinal regions affected by GA is a fundamental step for clinical diagnosis. In this paper, we present a novel weakly supervised model for GA segmentation in Spectral-Domain Optical Coherence Tomography (SD-OCT) images. A novel Multi-Scale Class Activation Map (MS-CAM) is proposed to highlight the discriminatory significance regions in localization and detail descriptions. To extract available multi-scale features, we design a Scaling and UpSampling (SUS) module to balance the information content between features of different scales. To capture more discriminative features, an Attentional Fully Connected (AFC) module is proposed by introducing the attention mechanism into the fully connected operations to enhance the significant informative features and suppress less useful ones. Based on the location cues, the final GA region prediction is obtained by the projection segmentation of MS-CAM. The experimental results on two independent datasets demonstrate that the proposed weakly supervised model outperforms the conventional GA segmentation methods and can produce similar or superior accuracy comparing with fully supervised approaches. The source code has been released and is available on GitHub: https://github.com/ jizexuan/Multi-Scale-Class-Activation-Map-Tensorflow. Xiao Ma 0011, Zexuan Ji, Sijie Niu, Theodore Leng, Daniel L. Rubin, Qiang Chen 0004 |
IEEE J. Biomed. Health Informatics | 5 |
| 2019 | Prediction of Imaging Outcomes from Electronic Health Records: Pulmonary Embolism Case-Study
Imon Banerjee, Miji Sofela, Timothy Amrhein, Daniel L. Rubin, Roham Zamanian, Matthew P. Lungren |
AMIA | 4 |
| 2019 | Detecting unanticipated actions downstream from clinical decision support: a data mining approach
Ron C. Li, Imon Banerjee, Daniel L. Rubin, Jonathan H. Chen |
AMIA | 3 |
| 2019 | Doubly Weak Supervision of Deep Learning Models for Head CT
Khaled Saab 0002, Jared Dunnmon, Roger E. Goldman, Alexander Ratner, Hersh Sagreiya, Christopher Ré, Daniel L. Rubin |
MICCAI (3) | 7 |
| 2019 | Comparative effectiveness of convolutional neural network (CNN) and recurrent neural network (RNN) architectures for radiology text report classification
Imon Banerjee, Yuan Ling, Matthew C. Chen, Sadid A. Hasan, Curt Langlotz, Nathaniel Moradzadeh, Brian E. Chapman, Timothy Amrhein, David A. Mong, Daniel L. Rubin, Oladimeji Farri, Matthew P. Lungren |
Artif. Intell. Medicine | 10 |
| 2019 | Automatic inference of BI-RADS final assessment categories from narrative mammography report findings
Imon Banerjee, Selen Bozkurt, Emel Alkim, Hersh Sagreiya, Allison W. Kurian, Daniel L. Rubin |
J. Biomed. Informatics | 6 |
| 2018 | A Scalable Machine Learning Approach for Inferring Probabilistic US-LI-RADS Categorization
Imon Banerjee, Hailey H. Choi, Terry S. Desser, Daniel L. Rubin |
AMIA | 4 |
| 2018 | An Automated Feature Engineering for Digital Rectal Examination Documentation using Natural Language Processing
Selen Bozkurt, Jung Park In, Kathleen Mary Kan, Michelle Ferrari, Daniel L. Rubin, James D. Brooks, Tina Hernandez-Boussard |
AMIA | 5 |
| 2018 | Unraveling the Molecular Basis of Lung Adenocarcinoma Dedifferentiation and Prognosis by Integrating Omics and Histopathology
Kun-Hsing Yu, Gerald J. Berry, Daniel L. Rubin, Christopher Ré, Russ B. Altman, Michael Snyder 0001 |
AMIA | 3 |
| 2018 | Distributed deep learning networks among institutions for medical imagingabstractObjective: Deep learning has become a promising approach for automated support for clinical diagnosis. When medical data samples are limited, collaboration among multiple institutions is necessary to achieve high algorithm performance. However, sharing patient data often has limitations due to technical, legal, or ethical concerns. In this study, we propose methods of distributing deep learning models as an attractive alternative to sharing patient data. Methods: We simulate the distribution of deep learning models across 4 institutions using various training heuristics and compare the results with a deep learning model trained on centrally hosted patient data. The training heuristics investigated include ensembling single institution models, single weight transfer, and cyclical weight transfer. We evaluated these approaches for image classification in 3 independent image collections (retinal fundus photos, mammography, and ImageNet). Results: We find that cyclical weight transfer resulted in a performance that was comparable to that of centrally hosted patient data. We also found that there is an improvement in the performance of cyclical weight transfer heuristic with a high frequency of weight transfer. Conclusions: We show that distributing deep learning models is an effective alternative to sharing patient data. This finding has implications for any collaborative deep learning study. Ken Chang, Niranjan Balachandar, Carson K. Lam, Darvin Yi, James M. Brown 0001, Andrew Beers, Bruce R. Rosen, Daniel L. Rubin, Jayashree Kalpathy-Cramer |
J. Am. Medical Informatics Assoc. | 8 |
| 2018 | Expanding a radiology lexicon using contextual patterns in radiology reportsabstractObjective: Distributional semantics algorithms, which learn vector space representations of words and phrases from large corpora, identify related terms based on contextual usage patterns. We hypothesize that distributional semantics can speed up lexicon expansion in a clinical domain, radiology, by unearthing synonyms from the corpus. Materials and Methods: We apply word2vec, a distributional semantics software package, to the text of radiology notes to identify synonyms for RadLex, a structured lexicon of radiology terms. We stratify performance by term category, term frequency, number of tokens in the term, vector magnitude, and the context window used in vector building. Results: Ranking candidates based on distributional similarity to a target term results in high curation efficiency: on a ranked list of 775 249 terms, >50% of synonyms occurred within the first 25 terms. Synonyms are easier to find if the target term is a phrase rather than a single word, if it occurs at least 100× in the corpus, and if its vector magnitude is between 4 and 5. Some RadLex categories, such as anatomical substances, are easier to identify synonyms for than others. Discussion: The unstructured text of clinical notes contains a wealth of information about human diseases and treatment patterns. However, searching and retrieving information from clinical notes often suffer due to variations in how similar concepts are described in the text. Biomedical lexicons address this challenge, but are expensive to produce and maintain. Distributional semantics algorithms can assist lexicon curation, saving researchers time and money. Bethany Percha, Yuhao Zhang 0004, Selen Bozkurt, Daniel L. Rubin, Russ B. Altman, Curt Langlotz |
J. Am. Medical Informatics Assoc. | 4 |
| 2018 | The LOINC RSNA radiology playbook - a unified terminology for radiology proceduresabstractObjective: This paper describes the unified LOINC/RSNA Radiology Playbook and the process by which it was produced. Methods: The Regenstrief Institute and the Radiological Society of North America (RSNA) developed a unification plan consisting of six objectives 1) develop a unified model for radiology procedure names that represents the attributes with an extensible set of values, 2) transform existing LOINC procedure codes into the unified model representation, 3) create a mapping between all the attribute values used in the unified model as coded in LOINC (ie, LOINC Parts) and their equivalent concepts in RadLex, 4) create a mapping between the existing procedure codes in the RadLex Core Playbook and the corresponding codes in LOINC, 5) develop a single integrated governance process for managing the unified terminology, and 6) publicly distribute the terminology artifacts. Results: We developed a unified model and instantiated it in a new LOINC release artifact that contains the LOINC codes and display name (ie LONG_COMMON_NAME) for each procedure, mappings between LOINC and the RSNA Playbook at the procedure code level, and connections between procedure terms and their attribute values that are expressed as LOINC Parts and RadLex IDs. We transformed all the existing LOINC content into the new model and publicly distributed it in standard releases. The organizations have also developed a joint governance process for ongoing maintenance of the terminology. Conclusions: The LOINC/RSNA Radiology Playbook provides a universal terminology standard for radiology orders and results. Daniel J. Vreeman, Swapna Abhyankar, Kenneth C. Wang, Chris Carr, Beverly Collins, Daniel L. Rubin, Curt Langlotz |
J. Am. Medical Informatics Assoc. | 6 |
| 2018 | Radiology report annotation using intelligent word embeddings: Applied to multi-institutional chest CT cohort
Imon Banerjee, Matthew C. Chen, Matthew P. Lungren, Daniel L. Rubin |
J. Biomed. Informatics | 4 |
| 2018 | Relevance feedback for enhancing content based image retrieval and automatic prediction of semantic image features: Application to bone tumor radiographs
Imon Banerjee, Camille Kurtz, Alon Edward Devorah, Bao H. Do, Daniel L. Rubin, Christopher F. Beaulieu |
J. Biomed. Informatics | 5 |
| 2018 | Automatic information extraction from unstructured mammography reports using distributed semantics
Anupama Gupta, Imon Banerjee, Daniel L. Rubin |
J. Biomed. Informatics | 3 |
| 2017 | Intelligent Word Embeddings of Free-Text Radiology Reports
Imon Banerjee, Sriraman Madhavan, Roger E. Goldman, Daniel L. Rubin |
AMIA | 4 |
| 2017 | Toward Automated Pre-Biopsy Thyroid Cancer Risk Estimation in Ultrasound
Alfiia Galimzianova, Sean M. Siebert, Aya Kamaya, Terry S. Desser, Daniel L. Rubin |
AMIA | 5 |
| 2017 | Differential Data Augmentation Techniques for Medical Imaging Classification Tasks
Zeshan Hussain, Francisco Gimenez, Darvin Yi, Daniel L. Rubin |
AMIA | 4 |
| 2017 | The LOINC/RSNA Radiology Playbook: A unified terminology for radiology procedures
Daniel J. Vreeman, Ken Wang, Chris Carr, Beverly Collins, Swapna Abhyankar, Jamalynne Deckard, Clement J. McDonald, Daniel L. Rubin, Curt Langlotz |
AMIA | 8 |
| 2017 | Predicting Non-Small Cell Lung Cancer Diagnosis and Prognosis by Fully Automated Microscopic Pathology Image Features
Kun-Hsing Yu, Ce Zhang 0001, Gerald J. Berry, Russ B. Altman, Christopher Ré, Daniel L. Rubin, Michael Snyder 0001 |
AMIA | 6 |
| 2017 | Inferring Generative Model Structure with Static AnalysisabstractObtaining enough labeled data to robustly train complex discriminative models is a major bottleneck in the machine learning pipeline. A popular solution is combining multiple sources of weak supervision using generative models. The structure of these models affects the quality of the training labels, but is difficult to learn without any ground truth labels. We instead rely on weak supervision sources having some structure by virtue of being encoded programmatically. We present Coral, a paradigm that infers generative model structure by statically analyzing the code for these heuristics, thus significantly reducing the amount of data required to learn structure. We prove that Coral's sample complexity scales quasilinearly with the number of heuristics and number of relations identified, improving over the standard sample complexity, which is exponential in n for learning n-th degree relations. Empirically, Coral matches or outperforms traditional structure learning approaches by up to 3.81 F1 points. Using Coral to model dependencies instead of assuming independence results in better performance than a fully supervised model by 3.07 accuracy points when heuristics are used to label radiology data without ground truth labels. Paroma Varma, Bryan Dawei He, Payal Bajaj, Nishith Khandwala, Imon Banerjee, Daniel L. Rubin, Christopher Ré |
NIPS | 6 |
| 2017 | Dynamic strategy for personalized medicine: An application to metastatic breast cancer
Ross D. Shachter, Allison W. Kurian, Daniel L. Rubin |
J. Biomed. Informatics | 4 |
| 2017 | Adaptive local window for level set segmentation of CT and MRI liver lesions
Assaf Hoogi, Christopher F. Beaulieu, Guilherme M. Cunha, Elhamy Heba, Claude B. Sirlin, Sandy Napel, Daniel L. Rubin |
Medical Image Anal. | 7 |
| 2017 | Piecewise convexity of artificial neural networks
Blaine Rister, Daniel L. Rubin |
Neural Networks | 2 |
| 2017 | Robust noise region-based active contour model via local similarity factor for image segmentation
Sijie Niu, Qiang Chen 0004, Luis de Sisternes, Zexuan Ji, Ze Ming Zhou, Daniel L. Rubin |
Pattern Recognit. | 6 |
| 2017 | Volumetric Image Registration From Invariant KeypointsabstractWe present a method for image registration based on 3D scale- and rotation-invariant keypoints. The method extends the scale invariant feature transform (SIFT) to arbitrary dimensions by making key modifications to orientation assignment and gradient histograms. Rotation invariance is proven mathematically. Additional modifications are made to extrema detection and keypoint matching based on the demands of image registration. Our experiments suggest that the choice of neighborhood in discrete extrema detection has a strong impact on image registration accuracy. In head MR images, the brain is registered to a labeled atlas with an average Dice coefficient of 92%, outperforming registration from mutual information as well as an existing 3D SIFT implementation. In abdominal CT images, the spine is registered with an average error of 4.82 mm. Furthermore, keypoints are matched with high precision in simulated head MR images exhibiting lesions from multiple sclerosis. These results were achieved using only affine transforms, and with no change in parameters across a wide variety of medical images. This paper is freely available as a cross-platform software library. Blaine Rister, Mark Horowitz, Daniel L. Rubin |
IEEE Trans. Image Process. | 3 |
| 2017 | A Convolutional Neural Network for Automatic Characterization of Plaque Composition in Carotid UltrasoundabstractCharacterization of carotid plaque composition, more specifically the amount of lipid core, fibrous tissue, and calcified tissue, is an important task for the identification of plaques that are prone to rupture, and thus for early risk estimation of cardiovascular and cerebrovascular events. Due to its low costs and wide availability, carotid ultrasound has the potential to become the modality of choice for plaque characterization in clinical practice. However, its significant image noise, coupled with the small size of the plaques and their complex appearance, makes it difficult for automated techniques to discriminate between the different plaque constituents. In this paper, we propose to address this challenging problem by exploiting the unique capabilities of the emerging deep learning framework. More specifically, and unlike existing works which require a priori definition of specific imaging features or thresholding values, we propose to build a convolutional neural network (CNN) that will automatically extract from the images the information that is optimal for the identification of the different plaque constituents. We used approximately 90 000 patches extracted from a database of images and corresponding expert plaque characterizations to train and to validate the proposed CNN. The results of cross-validation experiments show a correlation of about 0.90 with the clinical assessment for the estimation of lipid core, fibrous cap, and calcified tissue areas, indicating the potential of deep learning for the challenging task of automatic characterization of plaque composition in carotid ultrasound. Karim Lekadir, Alfiia Galimzianova, Àngels Betriu, Maria del Mar Vila, Laura Igual, Daniel L. Rubin, Elvira Fernández, Petia Radeva, Sandy Napel |
IEEE J. Biomed. Health Informatics | 6 |
| 2017 | Adaptive Estimation of Active Contour Parameters Using Convolutional Neural Networks and Texture AnalysisabstractIn this paper, we propose a generalization of the level set segmentation approach by supplying a novel method for adaptive estimation of active contour parameters. The presented segmentation method is fully automatic once the lesion has been detected. First, the location of the level set contour relative to the lesion is estimated using a convolutional neural network (CNN). The CNN has two convolutional layers for feature extraction, which lead into dense layers for classification. Second, the output CNN probabilities are then used to adaptively calculate the parameters of the active contour functional during the segmentation process. Finally, the adaptive window size surrounding each contour point is re-estimated by an iterative process that considers lesion size and spatial texture. We demonstrate the capabilities of our method on a dataset of 164 MRI and 112 CT images of liver lesions that includes low contrast and heterogeneous lesions as well as noisy images. To illustrate the strength of our method, we evaluated it against state of the art CNN-based and active contour techniques. For all cases, our method, as assessed by Dice similarity coefficients, performed significantly better than currently available methods. An average Dice improvement of 0.27 was found across the entire dataset over all comparisons. We also analyzed two challenging subsets of lesions and obtained a significant Dice improvement of 0.24 with our method (p <;0.001, Wilcoxon). Assaf Hoogi, Arjun Subramaniam, Rishi Veerapaneni, Daniel L. Rubin |
IEEE Trans. Medical Imaging | 4 |
| 2016 | Using automatically extracted information from mammography reports for decision-support
Selen Bozkurt, Francisco Gimenez, Elizabeth S. Burnside, Kemal Hakan Gülkesen, Daniel L. Rubin |
J. Biomed. Informatics | 5 |
| 2016 | Toward rapid learning in cancer treatment selection: An analytical engine for practice-based clinical data
Samuel G. Finlayson, Mia A. Levy, Sunil Reddy, Daniel L. Rubin |
J. Biomed. Informatics | 4 |
| 2016 | Automated classification of brain tumor type in whole-slide digital pathology images using local representative tiles
Jocelyn Barker, Assaf Hoogi, Adrien Depeursinge, Daniel L. Rubin |
Medical Image Anal. | 4 |
| 2016 | Improved Patch-Based Automated Liver Lesion Classification by Separate Analysis of the Interior and Boundary RegionsabstractThe bag-of-visual-words (BoVW) method with construction of a single dictionary of visual words has been used previously for a variety of classification tasks in medical imaging, including the diagnosis of liver lesions. In this paper, we describe a novel method for automated diagnosis of liver lesions in portal-phase computed tomography (CT) images that improves over single-dictionary BoVW methods by using an image patch representation of the interior and boundary regions of the lesions. Our approach captures characteristics of the lesion margin and of the lesion interior by creating two separate dictionaries for the margin and the interior regions of lesions ("dual dictionaries" of visual words). Based on these dictionaries, visual word histograms are generated for each region of interest within the lesion and its margin. For validation of our approach, we used two datasets from two different institutions, containing CT images of 194 liver lesions (61 cysts, 80 metastasis, and 53 hemangiomas). The final diagnosis of each lesion was established by radiologists. The classification accuracy for the images from the two institutions was 99% and 88%, respectively, and 93% for a combined dataset. Our new BoVW approach that uses dual dictionaries shows promising results. We believe the benefits of our approach may generalize to other application domains within radiology. Idit Diamant, Assaf Hoogi, Christopher F. Beaulieu, Mustafa Safdari, Eyal Klang, Michal Amitai, Hayit Greenspan, Daniel L. Rubin |
IEEE J. Biomed. Health Informatics | 8 |
| 2016 | A 3-D Riesz-Covariance Texture Model for Prediction of Nodule Recurrence in Lung CTabstractThis paper proposes a novel imaging biomarker of lung cancer relapse from 3-D texture analysis of CT images. Three-dimensional morphological nodular tissue properties are described in terms of 3-D Riesz-wavelets. The responses of the latter are aggregated within nodular regions by means of feature covariances, which leverage rich intra- and inter-variations of the feature space dimensions. When compared to the classical use of the average for feature aggregation, feature covariances preserve spatial co-variations between features. The obtained Riesz-covariance descriptors lie on a manifold governed by Riemannian geometry allowing geodesic measurements and differentiations. The latter property is incorporated both into a kernel for support vector machines (SVM) and a manifold-aware sparse regularized classifier. The effectiveness of the presented models is evaluated on a dataset of 110 patients with non-small cell lung carcinoma (NSCLC) and cancer recurrence information. Disease recurrence within a timeframe of 12 months could be predicted with an accuracy of 81.3-82.7%. The anatomical location of recurrence could be discriminated between local, regional and distant failure with an accuracy of 78.3-93.3%. The obtained results open novel research perspectives by revealing the importance of the nodular regions used to build the predictive models. Pol Cirujeda, Yashin Dicente Cid, Henning Müller, Daniel L. Rubin, Todd A. Aguilera, Billy W. Loo, Maximilian Diehn, Xavier Binefa, Adrien Depeursinge |
IEEE Trans. Medical Imaging | 4 |
| 2015 | Automated Grading of Gliomas using Deep Learning in Digital Pathology Images: A modular approach with ensemble of convolutional neural networks
Mehmet Günhan Ertosun, Daniel L. Rubin |
AMIA | 2 |
| 2015 | ePAD: Leveraging image data in learning healthcare systems
Daniel L. Rubin |
AMIA | 1 |
| 2015 | Polychromatic X-Ray Absorptiometry to Quantify Breast Density Volume, Ratio and their Associated Breast Cancer Risk in Full-Digital Mammography
Luis de Sisternes, Joseph Rothstein, Abra Jeffers, Weiva Sieh, Daniel L. Rubin |
AMIA | 5 |
| 2015 | Probabilistic visual search for masses within mammography images using deep learningabstractWe developed a deep learning-based visual search system for the task of automated search and localization of masses in whole mammography images. The system consists of two modules: a classification engine and a localization engine. It first classifies mammograms as containing a mass or no mass using a deep learning classifier, and then localizes the mass(es) within the image using a regional probabilistic approach based on a deep learning network. We obtained 85% accuracy for the task of identifying images that contain a mass, and we were able to localize 85% of the masses at an average of 0.9 false positives per image. Our system has the advantages of being able to work with an entire mammography image as input without the need for image segmentation or other pre-processing steps, such as cropping or tiling the image, and it is based on deep learning with unsupervised feature discovery, so it does not require pre-defined and hand-crafted image features. Mehmet Günhan Ertosun, Daniel L. Rubin |
BIBM | 2 |
| 2015 | Automatic Classification of Cancer Tumors Using Image Annotations and OntologiesabstractInformation about cancer stage in a patient is crucial when clinicians assess treatment progress. Determining cancer stage is a process that takes into account the description, location, characteristics and possible metastasis of cancerous tumors in a patient. It should follow classification standards, such as TNM Classification of Malignant Tumors. However, in clinical practice, the implementation of this process can be tedious and error-prone and create uncertainty. In order to alleviate these problems, we intend to assist radiologists by providing a second opinion in the evaluation of cancer stage in patients. For doing this, Semantic Web technologies, such as ontologies and reasoning, will be used to automatically classify cancer stages. This classification will use semantic annotations, made by radiologists (using the ePAD tool) and stored in the AIM format, and rules of an ontology representing the TNM standard. The whole process will be validated through a proof of concept with users from the Radiology Dept. of the Stanford University. Edson F. Luque Mamani, Daniel L. Rubin, Dilvan de Abreu Moreira |
CBMS | 2 |
| 2015 | 3D Markup of Radiological Images in ePAD, a Web-Based Image Annotation ToolabstractQuantitative and semantic information about medical images are vital parts of a radiological report. However, current image viewing systems do not record it in a format that permits machine interpretation. The ePAD tool can generate machine-computable image annotations on 2D images as part of a radiologist's routine workflow. The tool has been evaluated in image studies with good results. Since ePAD currently only provides 2D visualization and annotation of images, we developed a plugin to ePAD for the visualization of volumetric image datasets, using the three planes: axial, frontal and sagittal. A study with 6 radiologists was carried out to determine the best interface for also marking 3D ROIs. Video prototypes were created for 3 options: join pixels based on intensity similarity, detect borders around image features, and paint ROIs using a spheric 3D cursor. The 3D cursor was the preferred option. We present these results and also show the final 3D cursor implementation. Dilvan de Abreu Moreira, Cleber Hage, Edson F. Luque Mamani, Debra Willrett, Daniel L. Rubin |
CBMS | 5 |
| 2015 | Automatic abstraction of imaging observations with their characteristics from mammography reportsabstractBACKGROUND: Radiology reports are usually narrative, unstructured text, a format which hinders the ability to input report contents into decision support systems. In addition, reports often describe multiple lesions, and it is challenging to automatically extract information on each lesion and its relationships to characteristics, anatomic locations, and other information that describes it. The goal of our work is to develop natural language processing (NLP) methods to recognize each lesion in free-text mammography reports and to extract its corresponding relationships, producing a complete information frame for each lesion. MATERIALS AND METHODS: We built an NLP information extraction pipeline in the General Architecture for Text Engineering (GATE) NLP toolkit. Sequential processing modules are executed, producing an output information frame required for a mammography decision support system. Each lesion described in the report is identified by linking it with its anatomic location in the breast. In order to evaluate our system, we selected 300 mammography reports from a hospital report database. RESULTS: The gold standard contained 797 lesions, and our system detected 815 lesions (780 true positives, 35 false positives, and 17 false negatives). The precision of detecting all the imaging observations with their modifiers was 94.9, recall was 90.9, and the F measure was 92.8. CONCLUSIONS: Our NLP system extracts each imaging observation and its characteristics from mammography reports. Although our application focuses on the domain of mammography, we believe our approach can generalize to other domains and may narrow the gap between unstructured clinical report text and structured information extraction needed for data mining and decision support. Selen Bozkurt, Jafi A. Lipson, Utku Senol, Daniel L. Rubin |
J. Am. Medical Informatics Assoc. | 4 |
| 2015 | Comparing image search behaviour in the ARRS GoldMiner search engine and a clinical PACS/RIS
Maria De-Arteaga, Ivan Eggel, Bao H. Do, Daniel L. Rubin, Charles E. Kahn Jr., Henning Müller |
J. Biomed. Informatics | 4 |
| 2014 | A Novel Method to Assess Incompleteness of Mammography Report Content
Francisco Gimenez, Yirong Wu, Elizabeth S. Burnside, Daniel L. Rubin |
AMIA | 4 |
| 2014 | A semantic framework for the retrieval of similar radiological images based on medical annotationsabstractImage retrieval approaches can assist radiologists by finding similar images in databases as a means to providing decision support. In general, images are indexed using low-level imaging features, and a distance function is used to find the best matches in the feature space. However, using low-level features to capture the appearance of diseases in images is challenging and the semantic gap between these features and the high-level visual concepts in radiology may impair the system performance. We present a semantic framework that enables retrieving similar images based on high-level semantic image annotations. This framework relies on (1) an automatic approach to predict the annotations as semantic terms from Riesz texture image features and (2) a distance function to compare images considering both texture-based and radiodensity-based similarities among image annotations. Experiments performed on CT images emphasize the relevance of this framework. Camille Kurtz, Adrien Depeursinge, Christopher F. Beaulieu, Daniel L. Rubin |
ICIP | 4 |
| 2014 | Classification of hepatic lesions using the matching metric
Aaron Adcock, Daniel L. Rubin, Gunnar E. Carlsson |
Comput. Vis. Image Underst. | 2 |
| 2014 | A hierarchical knowledge-based approach for retrieving similar medical images described with semantic annotations
Camille Kurtz, Christopher F. Beaulieu, Sandy Napel, Daniel L. Rubin |
J. Biomed. Informatics | 4 |
| 2014 | On combining image-based and ontological semantic dissimilarities for medical image retrieval applications
Camille Kurtz, Adrien Depeursinge, Sandy Napel, Christopher F. Beaulieu, Daniel L. Rubin |
Medical Image Anal. | 5 |
| 2014 | Predicting Visual Semantic Descriptive Terms From Radiological Image Data: Preliminary Results With Liver Lesions in CTabstractWe describe a framework to model visual semantics of liver lesions in CT images in order to predict the visual semantic terms (VST) reported by radiologists in describing these lesions. Computational models of VST are learned from image data using linear combinations of high-order steerable Riesz wavelets and support vector machines (SVM). In a first step, these models are used to predict the presence of each semantic term that describes liver lesions. In a second step, the distances between all VST models are calculated to establish a nonhierarchical computationally-derived ontology of VST containing inter-term synonymy and complementarity. A preliminary evaluation of the proposed framework was carried out using 74 liver lesions annotated with a set of 18 VSTs from the RadLex ontology. A leave-one-patient-out cross-validation resulted in an average area under the ROC curve of 0.853 for predicting the presence of each VST. The proposed framework is expected to foster human-computer synergies for the interpretation of radiological images while using rotation-covariant computational models of VSTs to 1) quantify their local likelihood and 2) explicitly link them with pixel-based image content in the context of a given imaging domain. Adrien Depeursinge, Camille Kurtz, Christopher F. Beaulieu, Sandy Napel, Daniel L. Rubin |
IEEE Trans. Medical Imaging | 5 |
| 2013 | Classifying Benign and Malignant Lung Diseases by Applying Machine Learning Methods to Microscopic Pathology Images
Kun-Hsing Yu, Shanshan Tuo, Daniel L. Rubin |
AMIA | 3 |
| 2013 | Research and applications: Dynamic contrast-enhanced MRI-based biomarkers of therapeutic response in triple-negative breast cancerabstractOBJECTIVE: To predict the response of breast cancer patients to neoadjuvant chemotherapy (NAC) using features derived from dynamic contrast-enhanced (DCE) MRI. MATERIALS AND METHODS: 60 patients with triple-negative early-stage breast cancer receiving NAC were evaluated. Features assessed included clinical data, patterns of tumor response to treatment determined by DCE-MRI, MRI breast imaging-reporting and data system descriptors, and quantitative lesion kinetic texture derived from the gray-level co-occurrence matrix (GLCM). All features except for patterns of response were derived before chemotherapy; GLCM features were determined before and after chemotherapy. Treatment response was defined by the presence of residual invasive tumor and/or positive lymph nodes after chemotherapy. Statistical modeling was performed using Lasso logistic regression. RESULTS: Pre-chemotherapy imaging features predicted all measures of response except for residual tumor. Feature sets varied in effectiveness at predicting different definitions of treatment response, but in general, pre-chemotherapy imaging features were able to predict pathological complete response with area under the curve (AUC)=0.68, residual lymph node metastases with AUC=0.84 and residual tumor with lymph node metastases with AUC=0.83. Imaging features assessed after chemotherapy yielded significantly improved model performance over those assessed before chemotherapy for predicting residual tumor, but no other outcomes. CONCLUSIONS: DCE-MRI features can be used to predict whether triple-negative breast cancer patients will respond to NAC. Models such as the ones presented could help to identify patients not likely to respond to treatment and to direct them towards alternative therapies. Daniel I. Golden, Jafi A. Lipson, Melinda L. Telli, James M. Ford, Daniel L. Rubin |
J. Am. Medical Informatics Assoc. | 5 |
| 2013 | Automated drusen segmentation and quantification in SD-OCT images
Qiang Chen 0004, Theodore Leng, Luoluo Zheng, Lauren Kutzscher, Jeffrey J. Ma, Luis de Sisternes, Daniel L. Rubin |
Medical Image Anal. | 7 |
| 2012 | Automatic Annotation of Radiological Observations in Liver CT Images
Francisco Gimenez, Tiffany Ting Liu, Christopher F. Beaulieu, Daniel L. Rubin, Sandy Napel |
AMIA | 6 |
| 2012 | Automatic classification of mammography reports by BI-RADS breast tissue composition classabstractBecause breast tissue composition partially predicts breast cancer risk, classification of mammography reports by breast tissue composition is important from both a scientific and clinical perspective. A method is presented for using the unstructured text of mammography reports to classify them into BI-RADS breast tissue composition categories. An algorithm that uses regular expressions to automatically determine BI-RADS breast tissue composition classes for unstructured mammography reports was developed. The algorithm assigns each report to a single BI-RADS composition class: 'fatty', 'fibroglandular', 'heterogeneously dense', 'dense', or 'unspecified'. We evaluated its performance on mammography reports from two different institutions. The method achieves >99% classification accuracy on a test set of reports from the Marshfield Clinic (Wisconsin) and Stanford University. Since large-scale studies of breast cancer rely heavily on breast tissue composition information, this method could facilitate this research by helping mine large datasets to correlate breast composition with other covariates. Bethany Percha, Houssam Nassif, Jafi A. Lipson, Elizabeth S. Burnside, Daniel L. Rubin |
J. Am. Medical Informatics Assoc. | 5 |
| 2011 | The Biomedical Resource Ontology (BRO) to enable resource discovery in clinical and translational researchabstractThe biomedical research community relies on a diverse set of resources, both within their own institutions and at other research centers. In addition, an increasing number of shared electronic resources have been developed. Without effective means to locate and query these resources, it is challenging, if not impossible, for investigators to be aware of the myriad resources available, or to effectively perform resource discovery when the need arises. In this paper, we describe the development and use of the Biomedical Resource Ontology (BRO) to enable semantic annotation and discovery of biomedical resources. We also describe the Resource Discovery System (RDS) which is a federated, inter-institutional pilot project that uses the BRO to facilitate resource discovery on the Internet. Through the RDS framework and its associated Biositemaps infrastructure, the BRO facilitates semantic search and discovery of biomedical resources, breaking down barriers and streamlining scientific research that will improve human health. Jessica D. Tenenbaum, Patricia L. Whetzel, Kent Anderson, Charles D. Borromeo, Ivo D. Dinov, Davera Gabriel, Beth A. Kirschner, Barbara Mirel, Timothy D. Morris, Natasha F. Noy, Csongor Nyulas, David Rubenson, Paul R. Saxman, Nancy Whelan, Zachary C. Wright, Brian D. Athey, Michael J. Becich, Geoffrey S. Ginsburg, Mark A. Musen, Kevin A. Smith 0001, Alice F. Tarantal, Daniel L. Rubin, Peter Lyster |
J. Biomed. Informatics | 23 |
| 2011 | A practical method for transforming free-text eligibility criteria into computable criteria
Samson W. Tu, Mor Peleg, Simona Carini, Michael Bobak, Jessica Ross, Daniel L. Rubin, Ida Sim |
J. Biomed. Informatics | 6 |
| 2009 | Semantic Reasoning with Image Annotations for Tumor Assessment
Mia A. Levy, Martin J. O'Connor, Daniel L. Rubin |
AMIA | 3 |
| 2009 | Computational neuroanatomy: ontology-based representation of neural components and connectivityabstractBACKGROUND: A critical challenge in neuroscience is organizing, managing, and accessing the explosion in neuroscientific knowledge, particularly anatomic knowledge. We believe that explicit knowledge-based approaches to make neuroscientific knowledge computationally accessible will be helpful in tackling this challenge and will enable a variety of applications exploiting this knowledge, such as surgical planning. RESULTS: We developed ontology-based models of neuroanatomy to enable symbolic lookup, logical inference and mathematical modeling of neural systems. We built a prototype model of the motor system that integrates descriptive anatomic and qualitative functional neuroanatomical knowledge. In addition to modeling normal neuroanatomy, our approach provides an explicit representation of abnormal neural connectivity in disease states, such as common movement disorders. The ontology-based representation encodes both structural and functional aspects of neuroanatomy. The ontology-based models can be evaluated computationally, enabling development of automated computer reasoning applications. CONCLUSION: Neuroanatomical knowledge can be represented in machine-accessible format using ontologies. Computational neuroanatomical approaches such as described in this work could become a key tool in translational informatics, leading to decision support applications that inform and guide surgical planning and personalized care for neurological disease in the future. Daniel L. Rubin, Ion-Florin Talos, Michael Halle, Mark A. Musen, Ron Kikinis |
BMC Bioinform. | 1 |
| 2009 | Comparison of concept recognizers for building the Open Biomedical AnnotatorabstractThe National Center for Biomedical Ontology (NCBO) is developing a system for automated, ontology-based access to online biomedical resources (Shah NH, et al.: Ontology-driven indexing of public datasets for translational bioinformatics. BMC Bioinformatics 2009, 10(Suppl 2):S1). The system's indexing workflow processes the text metadata of diverse resources such as datasets from GEO and ArrayExpress to annotate and index them with concepts from appropriate ontologies. This indexing requires the use of a concept-recognition tool to identify ontology concepts in the resource's textual metadata. In this paper, we present a comparison of two concept recognizers - NLM's MetaMap and the University of Michigan's Mgrep. We utilize a number of data sources and dictionaries to evaluate the concept recognizers in terms of precision, recall, speed of execution, scalability and customizability. Our evaluations demonstrate that Mgrep has a clear edge over MetaMap for large-scale service oriented applications. Based on our analysis we also suggest areas of potential improvements for Mgrep. We have subsequently used Mgrep to build the Open Biomedical Annotator service. The Annotator service has access to a large dictionary of biomedical terms derived from the United Medical Language System (UMLS) and NCBO ontologies. The Annotator also leverages the hierarchical structure of the ontologies and their mappings to expand annotations. The Annotator service is available to the community as a REST Web service for creating ontology-based annotations of their data. Nigam H. Shah, Nipun Bhatia, Clément Jonquet, Daniel L. Rubin, Annie P. Chiang, Mark A. Musen |
BMC Bioinform. | 4 |
| 2009 | Research Paper: Automated Semantic Indexing of Figure Captions to Improve Radiology Image RetrievalabstractOBJECTIVE: We explored automated concept-based indexing of unstructured figure captions to improve retrieval of images from radiology journals. DESIGN: The MetaMap Transfer program (MMTx) was used to map the text of 84,846 figure captions from 9,004 peer-reviewed, English-language articles to concepts in three controlled vocabularies from the UMLS Metathesaurus, version 2006AA. Sampling procedures were used to estimate the standard information-retrieval metrics of precision and recall, and to evaluate the degree to which concept-based retrieval improved image retrieval. MEASUREMENTS: Precision was estimated based on a sample of 250 concepts. Recall was estimated based on a sample of 40 concepts. The authors measured the impact of concept-based retrieval to improve upon keyword-based retrieval in a random sample of 10,000 search queries issued by users of a radiology image search engine. RESULTS: Estimated precision was 0.897 (95% confidence interval, 0.857-0.937). Estimated recall was 0.930 (95% confidence interval, 0.838-1.000). In 5,535 of 10,000 search queries (55%), concept-based retrieval found results not identified by simple keyword matching; in 2,086 searches (21%), more than 75% of the results were found by concept-based search alone. CONCLUSION: Concept-based indexing of radiology journal figure captions achieved very high precision and recall, and significantly improved image retrieval. Charles E. Kahn Jr., Daniel L. Rubin |
J. Am. Medical Informatics Assoc. | 2 |
| 2008 | Tool Support to Enable Evaluation of the Clinical Response to Treatment
Mia A. Levy, Daniel L. Rubin |
AMIA | 2 |
| 2008 | A Bayesian Classifier for Differentiating Benign versus Malignant Thyroid Nodules using Sonographic Features
Yueyi I. Liu, Aya Kamaya, Terry S. Desser, Daniel L. Rubin |
AMIA | 4 |
| 2008 | FMA-RadLex: An Application Ontology of Radiological Anatomy derived from the Foundational Model of Anatomy Reference Ontology
José L. V. Mejino Jr., Daniel L. Rubin, James F. Brinkley |
AMIA | 2 |
| 2008 | iPad: Semantic Annotation and Markup of Radiological Images
Daniel L. Rubin, Cesar Rodriguez, Priyanka Shah, Christopher F. Beaulieu |
AMIA | 1 |
| 2008 | Biomedical ontologies: a functional perspectiveabstractThe information explosion in biology makes it difficult for researchers to stay abreast of current biomedical knowledge and to make sense of the massive amounts of online information. Ontologies--specifications of the entities, their attributes and relationships among the entities in a domain of discourse--are increasingly enabling biomedical researchers to accomplish these tasks. In fact, bio-ontologies are beginning to proliferate in step with accruing biological data. The myriad of ontologies being created enables researchers not only to solve some of the problems in handling the data explosion but also introduces new challenges. One of the key difficulties in realizing the full potential of ontologies in biomedical research is the isolation of various communities involved: some workers spend their career developing ontologies and ontology-related tools, while few researchers (biologists and physicians) know how ontologies can accelerate their research. The objective of this review is to give an overview of biomedical ontology in practical terms by providing a functional perspective--describing how bio-ontologies can and are being used. As biomedical scientists begin to recognize the many different ways ontologies enable biomedical research, they will drive the emergence of new computer applications that will help them exploit the wealth of research data now at their fingertips. Daniel L. Rubin, Nigam H. Shah, Natasha F. Noy |
Briefings Bioinform. | 1 |
| 2008 | MScanner: a classifier for retrieving Medline citationsabstractBACKGROUND: Keyword searching through PubMed and other systems is the standard means of retrieving information from Medline. However, ad-hoc retrieval systems do not meet all of the needs of databases that curate information from literature, or of text miners developing a corpus on a topic that has many terms indicative of relevance. Several databases have developed supervised learning methods that operate on a filtered subset of Medline, to classify Medline records so that fewer articles have to be manually reviewed for relevance. A few studies have considered generalisation of Medline classification to operate on the entire Medline database in a non-domain-specific manner, but existing applications lack speed, available implementations, or a means to measure performance in new domains. RESULTS: MScanner is an implementation of a Bayesian classifier that provides a simple web interface for submitting a corpus of relevant training examples in the form of PubMed IDs and returning results ranked by decreasing probability of relevance. For maximum speed it uses the Medical Subject Headings (MeSH) and journal of publication as a concise document representation, and takes roughly 90 seconds to return results against the 16 million records in Medline. The web interface provides interactive exploration of the results, and cross validated performance evaluation on the relevant input against a random subset of Medline. We describe the classifier implementation, cross validate it on three domain-specific topics, and compare its performance to that of an expert PubMed query for a complex topic. In cross validation on the three sample topics against 100,000 random articles, the classifier achieved excellent separation of relevant and irrelevant article score distributions, ROC areas between 0.97 and 0.99, and averaged precision between 0.69 and 0.92. CONCLUSION: MScanner is an effective non-domain-specific classifier that operates on the entire Medline database, and is suited to retrieving topics for which many features may indicate relevance. Its web interface simplifies the task of classifying Medline citations, compared to building a pre-filter and classifier specific to the topic. The data sets and open source code used to obtain the results in this paper are available on-line and as supplementary material, and the web interface may be accessed at http://mscanner.stanford.edu. Graham L. Poulter, Daniel L. Rubin, Russ B. Altman, Cathal Seoighe |
BMC Bioinform. | 2 |
| 2008 | A prototype symbolic model of canonical functional neuroanatomy of the motor system
Ion-Florin Talos, Daniel L. Rubin, Michael Halle, Mark A. Musen, Ron Kikinis |
J. Biomed. Informatics | 2 |
| 2008 | Network Analysis of Intrinsic Functional Brain Connectivity in Alzheimer's DiseaseabstractFunctional brain networks detected in task-free ("resting-state") functional magnetic resonance imaging (fMRI) have a small-world architecture that reflects a robust functional organization of the brain. Here, we examined whether this functional organization is disrupted in Alzheimer's disease (AD). Task-free fMRI data from 21 AD subjects and 18 age-matched controls were obtained. Wavelet analysis was applied to the fMRI data to compute frequency-dependent correlation matrices. Correlation matrices were thresholded to create 90-node undirected-graphs of functional brain networks. Small-world metrics (characteristic path length and clustering coefficient) were computed using graph analytical methods. In the low frequency interval 0.01 to 0.05 Hz, functional brain networks in controls showed small-world organization of brain activity, characterized by a high clustering coefficient and a low characteristic path length. In contrast, functional brain networks in AD showed loss of small-world properties, characterized by a significantly lower clustering coefficient (p<0.01), indicative of disrupted local connectivity. Clustering coefficients for the left and right hippocampus were significantly lower (p<0.01) in the AD group compared to the control group. Furthermore, the clustering coefficient distinguished AD participants from the controls with a sensitivity of 72% and specificity of 78%. Our study provides new evidence that there is disrupted organization of functional brain networks in AD. Small-world metrics can characterize the functional organization of the brain in AD, and our findings further suggest that these network measures may be useful as an imaging-based biomarker to distinguish AD from healthy aging. Kaustubh Supekar, Vinod Menon, Daniel L. Rubin, Mark A. Musen, Michael D. Greicius |
PLoS Comput. Biol. | 3 |
| 2008 | Translating the Foundational Model of Anatomy into OWL
Natasha F. Noy, Daniel L. Rubin |
J. Web Semant. | 2 |
| 2007 | LesionViewer: A Tool for Tracking Cancer Lesions Over Time
Mia A. Levy, Aaron Tam, Yael Garten, Daniel L. Rubin |
AMIA | 5 |
| 2007 | Annotation and query of tissue microarray data using the NCI ThesaurusabstractBACKGROUND: The Stanford Tissue Microarray Database (TMAD) is a repository of data serving a consortium of pathologists and biomedical researchers. The tissue samples in TMAD are annotated with multiple free-text fields, specifying the pathological diagnoses for each sample. These text annotations are not structured according to any ontology, making future integration of this resource with other biological and clinical data difficult. RESULTS: We developed methods to map these annotations to the NCI thesaurus. Using the NCI-T we can effectively represent annotations for about 86% of the samples. We demonstrate how this mapping enables ontology driven integration and querying of tissue microarray data. We have deployed the mapping and ontology driven querying tools at the TMAD site for general use. CONCLUSION: We have demonstrated that we can effectively map the diagnosis-related terms describing a sample in TMAD to the NCI-T. The NCI thesaurus terms have a wide coverage and provide terms for about 86% of the samples. In our opinion the NCI thesaurus can facilitate integration of this resource with other biological data. Nigam H. Shah, Daniel L. Rubin, Inigo Espinosa, Kelli Montgomery, Mark A. Musen |
BMC Bioinform. | 2 |
| 2006 | Ontology-Based Representation of Simulation Models of Physiology
Daniel L. Rubin, David Grossman, Maxwell Lewis Neal, Daniel L. Cook, James B. Bassingthwaighte, Mark A. Musen |
AMIA | 1 |
| 2006 | Ontology-based Annotation and Query of Tissue Microarray Data
Nigam H. Shah, Daniel L. Rubin, Kaustubh Supekar, Mark A. Musen |
AMIA | 2 |
| 2006 | Using ontologies linked with geometric models to reason about penetrating injuries
Daniel L. Rubin, Olivier Dameron, Yasser Bashir, David Grossman, Parvati Dev, Mark A. Musen |
Artif. Intell. Medicine | 1 |
| 2005 | Challenges in Converting Frame-Based Ontology into OWL: the Foundational Model of Anatomy Case-Study
Olivier Dameron, Daniel L. Rubin, Mark A. Musen |
AMIA | 2 |
| 2005 | Use of Description Logic Classification to Reason about Consequences of Penetrating Injuries
Daniel L. Rubin, Olivier Dameron, Mark A. Musen |
AMIA | 1 |
| 2005 | Protégé-OWL: Creating Ontology-Driven Reasoning Applications with the Web Ontology Language
Daniel L. Rubin, Holger Knublauch, Ray W. Fergerson, Olivier Dameron, Mark A. Musen |
AMIA | 1 |
| 2005 | Research Paper: Using Petri Net Tools to Study Properties and Dynamics of Biological SystemsabstractPetri Nets (PNs) and their extensions are promising methods for modeling and simulating biological systems. We surveyed PN formalisms and tools and compared them based on their mathematical capabilities as well as by their appropriateness to represent typical biological processes. We measured the ability of these tools to model specific features of biological systems and answer a set of biological questions that we defined. We found that different tools are required to provide all capabilities that we assessed. We created software to translate a generic PN model into most of the formalisms and tools discussed. We have also made available three models and suggest that a library of such models would catalyze progress in qualitative modeling via PNs. Development and wide adoption of common formats would enable researchers to share models and use different tools to analyze them without the need to convert to proprietary formats. Mor Peleg, Daniel L. Rubin, Russ B. Altman |
J. Am. Medical Informatics Assoc. | 2 |
| 2005 | Application of Information Technology: A Statistical Approach to Scanning the Biomedical Literature for Pharmacogenetics KnowledgeabstractOBJECTIVE: Biomedical databases summarize current scientific knowledge, but they generally require years of laborious curation effort to build, focusing on identifying pertinent literature and data in the voluminous biomedical literature. It is difficult to manually extract useful information embedded in the large volumes of literature, and automated intelligent text analysis tools are becoming increasingly essential to assist in these curation activities. The goal of the authors was to develop an automated method to identify articles in Medline citations that contain pharmacogenetics data pertaining to gene-drug relationships. DESIGN: The authors built and evaluated several candidate statistical models that characterize pharmacogenetics articles in terms of word usage and the profile of Medical Subject Headings (MeSH) used in those articles. The best-performing model was used to scan the entire Medline article database (11 million articles) to identify candidate pharmacogenetics articles. RESULTS: A sampling of the articles identified from scanning Medline was reviewed by a pharmacologist to assess the precision of the method. The authors' approach identified 4,892 pharmacogenetics articles in the literature with 92% precision. Their automated method took a fraction of the time to acquire these articles compared with the time expected to be taken to accumulate them manually. The authors have built a Web resource (http://pharmdemo.stanford.edu/pharmdb/main.spy) to provide access to their results. CONCLUSION: A statistical classification approach can screen the primary literature to pharmacogenetics articles with high precision. Such methods may assist curators in acquiring pertinent literature in building biomedical databases. Daniel L. Rubin, Caroline F. Thorn, Teri E. Klein, Russ B. Altman |
J. Am. Medical Informatics Assoc. | 1 |
| 2002 | Representing genetic sequence data for pharmacogenomics: an evolutionary approach using ontological and relational modelsabstractAbstract Motivation: The information model chosen to store biological data affects the types of queries possible, database performance, and difficulty in updating that information model. Genetic sequence data for pharmacogenetics studies can be complex, and the best information model to use may change over time. As experimental and analytical methods change, and as biological knowledge advances, the data storage requirements and types of queries needed may also change. Results: We developed a model for genetic sequence and polymorphism data, and used XML Schema to specify the elements and attributes required for this model. We implemented this model as an ontology in a frame-based representation and as a relational model in a database system. We collected genetic data from two pharmacogenetics resequencing studies, and formulated queries useful for analysing these data. We compared the ontology and relational models in terms of query complexity, performance, and difficulty in changing the information model. Our results demonstrate benefits of evolving the schema for storing pharmacogenetics data: ontologies perform well in early design stages as the information model changes rapidly and simplify query formulation, while relational models offer improved query speed once the information model and types of queries needed stabilize. Availability: Our ontology and relational models are available at http://smi-web.stanford.edu/projects/helix/pubs/ismb02/. Contact: [email protected]@[email protected] Keywords: ontologies; relational databases; schema; data models; pharmacogenomics. Daniel L. Rubin, Farhad Shafa, Diane E. Oliver, Micheal Hewett, Russ B. Altman |
ISMB | 1 |
| 2000 | A Bayesian network for mammography
Elizabeth S. Burnside, Daniel L. Rubin, Ross D. Shachter |
AMIA | 2 |
| 2000 | Knowledge representation and tool support for critiquing clinical trial protocols
Daniel L. Rubin, John H. Gennari, Mark A. Musen |
AMIA | 1 |
| 1999 | Tool support for authoring eligibility criteria for cancer trials
Daniel L. Rubin, John H. Gennari, Sandra Srinivas, Allen Yuen, Herbert Kaizer, Mark A. Musen, John S. Silva |
AMIA | 1 |