Jacob D. Furst

dblp:08/1640-1 · also Jacob David Furst, Jacob Furst 0001 · DBLP profile ↗
← Back
36ranked-venue papers
2as first author
10since 2021 · last 2025
0000-0002-1015-6290ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 21 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 11 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 9 · 2 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2
YearPublicationVenuePosition
2025 Enhancing Pulmonary Nodule Localization Based on Latent Representations
abstract
Accurately localizing pulmonary nodules relative to other anatomical structures is crucial for disease management, guiding biopsies, and formulating effective treatment strategies. This study introduces a fully automated approach for classifying nodules detected in computed tomography (CT) images as pleural (near the pleura as$N_{p}$) or non-pleural (distant from the pleura as$D_{p}$). We propose a combination of Principal Component Analysis (PCA) and deep learning approach to determine a threshold for the optimal correlation between the original image and the latent PCA-based image reconstruction used to classify lung nodule as$N_{p}$or$D_{p}$. Applying our methodology to the Lung Image Database Consortium and Image Database Resource Initiative (LIDC-IDRI) dataset, we found that the PCA-based approach demonstrated substantial agreement with human evaluations of the proximity of the lung nodule to the pleura, achieving a Cohen's Kappa value of 0.76, outperforming an intensitybased baseline model with a Cohen's Kappa value of 0.55. Further analysis revealed significant variability in radiologists' semantic ratings for$N_{p}$versus$D_{p}$, with the highest variability observed in the texture feature. These findings demonstrate that our approach has the potential to enhance the classification of lung nodule localization when integrated into computer-aided diagnosis (CAD) systems.
Charmi Patel, Yiyang Wang 0003, Roselyne Tchoua, Thiruvarangan Ramaraj, Jacob D. Furst, Daniela Raicu
CBMS5
2025 Leveraging Hidden Patterns in Open-Ended Community Health Workers' Notes to Improve Prediction of Patient Readmission
abstract
Patient readmissions to Emergency Departments (EDs) pose significant challenges to healthcare systems, often indicating suboptimal care transitions, inadequate patient education, or insufficient post-discharge support. These unplanned readmissions not only compromise patient outcomes but also contribute to escalating healthcare costs. In response, healthcare providers are increasingly seeking predictive models to identify at-risk patients and implement preventive strategies. While advanced deep learning models have shown promise in predicting readmission risks, their "black-box" nature and substantial data requirements often limit clinical applicability due to a lack of interpretability. This study investigates the integration of Machine Learning (ML) and natural language processing (NLP) techniques to enhance the accuracy and explainability of patient readmission risk predictions. Utilizing data from the Sinai Urban Health Institute (SUHI), we analysed both structured data including demographics, interaction logs, social determinants of health (SDoH) survey answers and unstructured data, specifically notes capturing conversations between patients and Community Health Workers (CHWs). Our findings indicate that incorporating unstructured textual data improved model performance, with the area under the receiver operating characteristic curve (AUC) increasing from 0.68 to 0.74. This enhancement suggests that patient-CHW conversations capture critical, non-medical factors influencing readmissions, such as personal needs and social support deficits, which are not typically recorded in standard medical records. The study underscores the value of integrating patient-CHW contact notes into predictive modelling, helping to highlight the important role of community health in informing targeted interventions to reduce preventable readmissions.
Naveen Kumar Reddy Veeramreddy, Ankita Mishra, Navika Maglani, Sameer Shaik, Kelly McCabe, Jacob D. Furst, Daniela Raicu, Roselyne Tchoua, Jamshid Sourati
eScience6
2025 Video Label Refinement for Temporal Localization
abstract
When performing temporal localization, achieving high precision results or high accuracy at the highest temporal Intersection over Union (tIoU) requires training a model with class labels that have precise boundaries. Obtaining high-precision class labels solely from human annotators presents challenges, such as the difficulty of accurately discerning when a movement begins and ends by watching a video. This work expands on an approach improving the temporal boundaries of class labels for activity recognition in video to handle complex movement patterns in human interactions. This method applies signal processing methods to motion features to identify shapes in signals associated with activities to be classified and adjust the boundaries accurately. When these methods are applied to existing human-provided activity annotations, they can improve the accuracy of class label boundaries. This improves temporal localization model performance at the high tIoUs.
Jennifer Piane, Thiruvarangan Ramaraj, Jacob D. Furst, Daniela Raicu
ICME3
2024 Clustering density based gene expression data in the mouse brainstem and comparison to actual neuroanatomy
abstract
The medullary reticular formation is a network of brainstem nuclei and neurons in the brain that is composed of circuit pathways from the brain to the spinal cord. It carries out vital roles and orofacial behaviors such as vocalization, swallowing, and breathing. A disruption in these neuronal circuits can be threatening for persons with neurodevelopmental and neurodegenerative disorders, leading to deficits in breathing, swallowing, and communication. However, the neurons associated with orofacial motor controls are not clearly defined or localized within the reticular formation due to the lack of molecular markers and distinct brain tissue cells. This project serves to understand the organization of the reticular formation and define where orofacial motor behaviors are localized. We hypothesize that differences in gene expression and three-dimensional proximities can identify functional subpopulations of the reticular formation that control orofacial motor behaviors. Using gene density measures of the coronal plane in the adult male mouse brain that are attributed to the reticular formation, we computed 13 clusters using K-Means clustering with Euclidean distance measure based on brainstem nuclei that are biologically significant to the control of orofacial behaviors. This clustering pattern has implications for understanding the organization of the reticular formation. By understanding the organization of these brainstem nuclei, neuroscientists can better understand how neurodegenerative disorders work and potentially lead to effective treatments for persons suffering from them.
Colin Buenvenida, Brandon Kong, Shane Jung, Kaiwen Kam, Jacob D. Furst, Thiruvarangan Ramaraj
CIBCB5
2023 Curriculum gDRO: Improving Lung Malignancy Classification through Robust Curriculum Task Learning
abstract
Deep learning models used in Computer-Aided Diagnosis (CAD) systems are often trained with Empirical Risk Minimization (ERM) loss. These models often achieve high overall classification accuracy but with lower classification accuracy on certain subgroups. In the context of lung nodule malignancy classification task, these atypical subgroups exist due to the lung cancer heterogeneity. In this study, we characterize lung nodule malignancy subgroups using the malignancy likelihood ratings given by radiologists and improve the worst subgroup performance by utilizing group Distributionally Robust Optimization (gDRO). However, we noticed that gDRO improves on worst subgroup performance from the benign category, which has less clinical importance than improving classification accuracy for a malignant subgroup. Therefore, we propose a novel curriculum gDRO training scheme that trains for an “easy” task (nodule malignancy is determinate or indeterminate for radiologists) first, then for a “hard” task (malignant, benign, or indeterminate nodule). Our results indicate that our approach boosts the worst group subclass accuracy from the malignant category, by up to 6 percentage points compared to standard methods that address and improve worst group classification performance.
Arun Sivakumar, Yiyang Wang 0003, Roselyne Tchoua, Thiruvarangan Ramaraj, Jacob D. Furst, Daniela Raicu
CBMS5
2022 Text Summarization towards Scientific Information Extraction
abstract
Despite the exponential growth in scientific textual content, publications remain the primary means of disseminating vital research to experts within their respective fields. These texts are predominantly written for human consumption, resulting in two fundamental challenges; experts cannot efficiently remain well-informed to leverage the latest discoveries, and applications which rely on valuable insights buried in these texts cannot effectively build upon published results. Consequently, scientific progress stalls. Automatic Text Summarization (ATS) and Information Extraction (IE) are two essential fields which address this problem. While the two research topics are often studied independently, this work proposes to look at ATS in the context of IE, specifically as it relates to Scientific IE. However, Scientific Information Extraction faces several challenges; chiefly, the scarcity of relevant entities and insufficient training data. In this paper, we focus on extractive ATS, which identifies the most valuable sentences from textual content for the purpose of ultimately extracting scientific relations. We account for the associated challenges by means of an ensemble method through the integration of three weakly supervised learning models, one for each entity of the target relation. Notably, while the relation is well defined, we do not require previously annotated data for the entities composing the relation. The central objective is to generate balanced training data, which many advanced natural language processing models require. We apply this idea in the domain of materials science, extracting the polymer-glass transition temperature relation and achieve 94.7% recall (i.e., sentences which contain relations annotated by humans), while reducing the text by 99.3% of the original document.
Abigail Keller, Jacob D. Furst, Daniela Raicu, Peter M. Hastings, Roselyne Tchoua
e-Science2
2022 Predicting ME/CFS After Infectious Mononucleosis Using Cytokine Network Correlations
abstract
We investigated if a predictive modeling strategy based on the interdependence of the cytokine network could accurately predict if a patient would develop Myalgic Encephalomyelitis/Chronic Fatigue Syndrome (ME/CFS) after contracting infectious mononucleosis (IM). We analyzed previously collected data from Northwestern University (NU) students in a three-stage experiment, following them from the start of the school year (Stage 1), to development of IM (Stage 2), to six months post development of IM (Stage 3). At all three stages, blood was stored from participants for cytokine measurement and analysis. Additionally, eight psychological and behavioral scales were used to identify participants as healthy controls or as ME/CFS. Using participants’ measured cytokine expression levels, we built a predictive model based on the inherent correlations within the cytokine network. We found that we could predict ME/CFS in patients 6 months after IM with 86.84% accuracy using correlation matrices made from cytokines taken during IM infection. These results suggest that there may be potential in using an approach that is based on the interdependence of the cytokine network to predict ME/CFS post IM. Future work may explore the validity of these findings and if such an approach could have applications in other diseases.
Jennifer Schwabe, Chelsea Hua, Emma M. Allen, Leonard A. Jason, Jacob D. Furst, Daniela Raicu
ICMLA5
2022 Lung Nodule Malignancy Subtype Discovery with Semantic Learning
abstract
Computer-aided diagnosis (CAD) systems have been widely used as second readers in lung cancer diagnosis. However, lung cancer heterogeneity and lack of using human annotated semantic characteristics are significant obstacles to an accurate and explainable CAD outcome. We propose a novel CAD scheme that characterizes lung nodule malignancy subtypes semantically and classifies nodule malignancy through a semantic learning process. We built and evaluated our method on a publicly available dataset, Lung Image Database Consortium (LIDC). We discovered and characterized two malignant nodule subtypes and two benign subtypes. In addition, we achieved a significantly improved malignancy classification result by incorporating the semantic learning process when compared with a result without using the semantic learning. This study suggests that future CAD systems should include disease heterogeneity as a model input to refine the model explainability; our classification result provides case-base evidence that learning image semantic explanations is valuable for improving CAD accuracy. While we tested our approach on one dataset, the proposed approach is applicable to other CAD tasks if semantic annotations are available.
Yiyang Wang 0003, Bowen Qiu, Thiruvarangan Ramaraj, Ilyas Ustun, Jacob D. Furst, Daniela Raicu
ICPR5
2021 Approaches to Evaluating Eye Gaze Patterns between Physician-Patient Interaction in Primary Care Clinic
abstract
Assessing nonverbal communication in primary care settings may provide insights into health disparities and inform the design of health information technologies in those settings. The aim of this study was to develop approaches to measure and compare a single nonverbal communication behavior (eye gaze) between physicians and their patients in primary care clinics, and to assess whether physician gaze is consistent between patients and with other physicians. This analysis method could lead to design guidelines for technologies and more effective assessments of interventions. Data came from a study that included 11 physicians and 77 patients from Federally Qualified Health Centers (FQHCs) in Chicago, Illinois. We propose two approaches, K-means Clustering and Dynamic Time Warping (DTW), to evaluate eye gaze behavior patterns consistency between physicians and patients. Results from K-means clustering showed some consistency of gaze patterns across patients within the same physician in some visits. Furthermore, some similarities of gaze patterns across patients of different physicians were found. The t-test results from DTW showed some consistency of gaze patterns across physicians at p-value of 0.05 and 0.01. DTW results showed the viability of the approach of evaluating the consistency in physician-patient interaction. Additional approaches to evaluate eye gaze patterns, such as gaze patterns with technology, may also be considered to better understand the interaction between physicians and patients.
Amal N. Almansour, Jacob D. Furst, Daniela Raicu, Enid Montague
BIBM2
2021 Ensemble Labeling towards Scientific Information Extraction (ELSIE) - Blob Extraction
abstract
Scientific publications constitute an extremely valuable repository of knowledge and collection of facts crucial to the advancement of science and development of applications, which grows as researchers learn from previous works and scientists use results in the literature to design and create. With the exponential growth of available publications, reading and extracting this wealth of information has become impractical for humans. Despite great progress in natural language processing, machine-learned solutions require large amounts of carefully annotated data for good performance. This is especially true in the context of accurately labeling and extracting complex scientific data. Towards our ultimate goal of extracting scientific facts from the literature, we first aim to identify blobs of text that contain all of the facts in a publication to be later automatically extracted or scrutinized by experts. Our previous work identified some facts missed by experts yet missed others due to the assumption that the target relation—here, a polymer and its glass transition temperature—would be contained within the same sentence. We set out to enhance our approximate labeling system to look back and ahead for missing information and successfully achieved 100% recall of scientific facts while reducing the full-text publication to 6% of its original size. Moreover, we assign confidence scores to sentences to further assist expert curators in identifying important sentences and facts locked in unstructured text.
Erin Murphy, Alexander Rasin, Jacob D. Furst, Daniela Raicu, Roselyne Tchoua
e-Science3
2020 Heatmap Template Generation for COVID-19 Biomarker Detection in Chest X-rays
abstract
Detecting and identifying patterns in chest X-ray images of Covid-19 patients are important tasks for understanding the disease and for making differential diagnosis. Given the relatively small number of available Covid-19 X-ray images and the need to make progress in understanding the disease, we propose a transfer learning technique applied to a pretrained VGG19 neural network to build a deep convolutional model capable of detecting four possible conditions: normal (healthy), bacteria, virus (not Covid-19), and Covid-19. The transformation of the multi-class deep learning output into binary outputs and the detection of Covid-19 image patterns using Grad-CAM technique show promising results. The discovered patterns are consistent across images from a given class of disease and constitute explanations of how the deep learning model makes classification decisions. In the long run, the identified patterns can serve as biomarkers for a given disease in chest X-ray images.
Mirtha Lucas, Miguel Lerma, Jacob D. Furst, Daniela Raicu
BIBE3
2020 Robust Physician Gaze Prediction Using a Deep Learning Approach
abstract
The patient-physician relationship is an integral part of primary care visits. To build a better relationship, understanding the communication between patient and physician is the key. This study focused on analyzing the gaze, one of the most important non-verbal behaviors found to influence patient outcomes. Gaze analysis often needs a manual rating process which might be time-consuming, costly, and unreliable. This research aimed to support automated analysis of physician-patient interaction using a deep convolutional neural network with transfer learning to a build robust model for physician gaze prediction. Utilizing only 3 minutes of 15 videos capturing 3 physicians interacting with different patients in a clinical setting, the model achieved over 98% accuracy for train, test, and validation sets. By visualizing the convolutional layers and comparing sample frames from different interactions, results highlighted several patterns shared across frames predicted correctly from both seen and unseen video sequences. The proposed work has the potential to informed the future design of technologies used to capture the clinical interaction and provide real-time feedback for physicians, which will contribute to the improvement of care quality.
Tianyi Tan, Enid Montague, Jacob D. Furst, Daniela Raicu
BIBE3
2020 Explainable Deep Learning for Biomarker Classification of OCT Images
abstract
Advanced form of age-related macular degeneration (AMD) is a major health burden that can lead to irreversible vision loss in the elderly population. The early signs of AMD are drusen, which appear as yellowish deposits under the retina. The end-stages of AMD include two forms: wet AMD (neovascular) and geographic atrophy (GA, non-neovascular). We propose a deep learning approach using a pre-trained VGG-19 neural network to classify the optical coherence tomography (OCT) images into wet AMD, GA, drusen, and healthy images. To explain the results, we quantify and present the prediction confidence and reliability as well as the regions of interest deemed to be important by the deep learning model when classifying OCT images. The sensitivity of classification for wet AMD, GA, drusen, and healthy images was 91.67%, 88.00%, 96.30% and 100% respectively, with the most confident predictions being for the drusen and healthy images. Visual inspection of the misclassified images using heatmaps using the Gradient-weighted Class Activation Mapping (Grad-CAM) algorithm revealed that, for most of the misclassified images that had high prediction confidence, the algorithm identified correctly more than one region of interest, each belonging to a different AMD category. We concluded that, rather than assigning just one label to an AMD image, algorithms for AMD classification should allow multi-labels as images generally show evidence of more than one stage of AMD simultaneously (e.g., drusen as the predominant region and GA as a small region of interest).
Yiyang Wang 0003, Mirtha Lucas, Jacob D. Furst, Amani A. Fawzi, Daniela Raicu
BIBE3
2020 Enhancing Recall Using Data Cleaning for Biomedical Big Data
abstract
In clinical practice, large amounts of heterogeneous medical data are generated on a daily basis. This data has the potential to be used for biomedical research and as a diagnostic reference for physicians. However, leveraging heterogeneous data for analysis requires integrating it first. Integration process includes a pre-processing data cleaning phase that eliminates inconsistencies and errors originating from each data source. In this paper, we describe a workflow for cleaning heterogeneous biomedical data sources. Our novel data cleaning approach can be applied for replacement of missing text and to improve the number of relevant cases retrieved by search queries. When the threshold for missing category replacement is met, our results show that our method achieves a missing content replacement precision of 85%, which represents an improvement of 18% over the baseline state of our datasets.
Priya Deshpande, Alexander Rasin, Roselyne Tchoua, Jacob D. Furst, Daniela Raicu, Sameer K. Antani
CBMS4
2020 Assessment of Medical Reports Uncertainty through Topic Modeling and Machine Learning
abstract
Medical uncertainty is identified as one of the most important factors leading to miscommunication between health care providers and patients. Along with the rapid growth of Natural Language Processing and Machine Learning techniques, more opportunities to understand medical uncertainty became available, including quantifying and modeling medical uncertainty. In this work, we obtained more than 20,000 radiology teaching files as medical reports from multiple resources. Uncertainty ontologies were also obtained from two resources. We expanded the uncertainty term list by applying topic model Word2Vec to identify similar terms to original uncertainty ontologies. The teaching files were quantified into 5 uncertainty level classes by the sum of the Term Frequency - Inverse Document Frequency (TF-IDF) of the expanded uncertainty terms list. Results from topic modelling were used to produce features. The product of TF-IDF results of uncertainty terms in teaching files and the topic model results were then used to train classifiers to predict the uncertainty level in medical reports. Our exhaustive experimental analysis showed that Decision Tree can classify uncertainty level of medical reports at overall accuracy 82%, which is higher than K-nearest Neighbor (80%) and Naive Bayes (75%). The model can be used to identify medical reports' uncertainty level and limit miscommunication between parties and reduce diagnostic errors.
Mengyuan Shang, Jacob D. Furst, Daniela Raicu
CBMS2
2019 Hand-Eye Coordination: Automating the Annotation of Physician-Patient Interactions
abstract
The widespread adoption of electronic health records within clinical settings has renewed interest in understanding physician-patient interactions. Previous work analyzing clinical interactions has mostly coupled patient surveys with manually annotated video interactions provided by human coders. Physician gaze is among the components of the non-verbal interaction which has been found to impact patient outcomes. The work described in this paper illustrates an automated system for video labeling of patient-physician interactions and shows that image features (such as areas and positioning of physicians' hands) can provide important visual aids for learning physician gaze with over 90% accuracy. While our approach focuses on physician gaze, it can be extended to capture other clinical human-human and human-technology interactions as well as connect these interactions to patient ratings of clinical interactions.
Daniel Gutstein, Enid Montague, Jacob D. Furst, Daniela Raicu
BIBE3
2019 Optical Flow, Positioning, and Eye Coordination: Automating the Annotation of Physician-Patient Interactions
abstract
The widespread adoption of electronic health records within clinical settings has renewed interest in understanding physician-patient interactions. Previous work analyzing clinical interactions has mostly coupled patient surveys with manually annotated video interactions provided by human coders. Physician gaze is among the components of the non-verbal interaction which has been found to impact patient outcomes. The work described in this paper illustrates an automated system for multi-video labeling of patient-physician interactions and shows that image features (in the form of body positioning coordinates and optical flow) can provide important visual aids for learning physician gaze with over 90% accuracy. While our approach focuses on physician gaze, it can be extended to capture other clinical human-human and human-technology interactions as well as connect these interactions to patient ratings of clinical interactions.
Daniel Gutstein, Enid Montague, Jacob D. Furst, Daniela Raicu
BIBM3
2018 Detecting Database File Tampering through Page Carving
abstract
Database Management Systems (DBMSes) secure data against regular users through defensive mechanisms such as access control, and against privileged users with detection mechanisms such as audit logging. Interestingly, these security mechanisms are built into the DBMS and are thus only useful for monitoring or stopping operations that are executed through the DBMS API. Any access that involves directly modifying database files (at file system level) would, by definition, bypass any and all security layers built into the DBMS itself. In this paper, we propose and evaluate an approach that detects direct modifications to database files that have already bypassed the DBMS and its internal security mechanisms. Our approach applies forensic analysis to first validate database indexes and then compares index state with data in the DBMS tables. We show that indexes are much more difficult to modify and can be further fortified with hashing. Our approach supports most relational DBMSes by leveraging index structures that are already built into the system to detect database storage tampering that would currently remain undetectable.
James Wagner, Alexander Rasin, Tanu Malik, Karen Heart, Jacob D. Furst, Jonathan Grier
EDBT5
2016 Compression-based distance methods as an alternative to statistical methods for constructing phylogenetic trees
abstract
Distance based methods for constructing phylogenetic trees have long been considered inconsistent and inferior to the more dominant statistical methods. However, use of compression methods specific to DNA could prove valuable in improving the effectiveness of distance based methods. To demonstrate the validity of distance-based methods when utilizing current DNA compression algorithms, such as MFCompress, we have applied such a method to datasets of closely related species of fish from the suborder Labroidei and to strains of Ebola. In both cases, we have managed to produce trees that are either very similar or identical to published trees produced using statistically based methods. This suggests that distance based methods can perform comparably to statistically based methods without requiring as much pre-processing of original DNA sequences or system resources. Additionally, the results also stress the importance of using accurate methods of calculating species distance due to the way that one specific DNA compression algorithm, MFCompress, consistently and convincingly managed to outperform other popular, general use compression algorithms.
Mohamed El-Dirany, Forrest Wang, Jacob D. Furst, Daniela Raicu
BIBM3
2016 C. elegans search behavior analysis using Multivariate Dynamic Time Warping
abstract
The nematode Caenorhabditis elegans is widely used as a genetic model organism because of its relevance to human biology and disease. Quantifying C. elegans movement behavior is important for understanding the differences between genotypes under various food conditions. One of the most common ways to quantify C. elegans movement is through path analysis. However, comparing the movement paths between different C. elegans genotypes is challenging because of the variations in the trajectories' time lengths. We propose a novel approach that combines Multivariate Dynamic Time Warping and step-length path encoding to cluster C. elegans movement path data and further show that the proposed combination avoids the need for extracting complex path features.
Yiyang Wang 0003, Carleton Smith, Mingfei Shao, Jacob D. Furst, Daniela Raicu, Hongkyun Kim
BIBM5
2014 Machine-Sourced Segmentations vs. Expert-Sourced Segmentations for the Classification of Lung Nodules with Outlier Removal
abstract
Computer-aided diagnosis systems can provide additional opinions that serve as an aid to radiologists in the early detection of lung nodules. Previous CAD models have relied on radiologist-delineated contours to extract image features and classify lung nodules into semantic ratings. Manually creating these contours can be time-consuming and expensive. This paper proposes a different CAD system based on multiple machine-sourced segmentations that can provide semantic ratings at least as accurate as a panel of experts in order to aid in the diagnostic process. However, the mass production of machine-sourced segmentations may sometimes produce unwanted noise. Therefore, we propose to filter out the bad segmentations by applying an outlier detection algorithm that identifies segmentations that are far away from the majority of the segmentations. Our results are compared to a CAD system based on expert-sourced contours and a reference truth generated by radiologists' semantic ratings. Using the Lung Image Database Consortium dataset, we show that machine-sourced segmentations provide predictions at least as good as expert-sourced segmentations and how outlier removal affects mostly shape-dependent semantic ratings.
Mayra Alejandra Molina Puentes, Jacob D. Furst, Daniela Raicu
ICMLA2
2013 Weak Segmentations and Ensemble Learning to Predict Semantic Ratings of Lung Nodules
abstract
Computer-aided diagnosis (CAD) can be used as "second readers" in the imaging diagnostic process. Typically to create a CAD system, the region of interest (ROI) has to be first detected and then delineated. This can be done either manually or automatically. Given that manually delineating ROIs is a time consuming and costly process, we propose a CAD system based on multiple computer-derived weak segmentations (WSCAD) and show that its diagnosis performance is at least as good as the predictions developed using manual radiologist segmentations. The proposed CAD system extracts a set of image features from the weak segmentations and uses them in an ensemble of classification algorithms to predict semantic ratings such as malignancy. These automated results are compared against a reference truth based on ratings and segmentations provided by radiologists to determine if it is necessary to obtain manual radiologist segmentations in order to develop a CAD. By developing a pair of CADs using the Lung Image Database Consortium (LIDC) data, we show that WSCADs are at least as accurate in predicting semantic ratings as CADs based on radiologist segmentation.
Ethan Smith, Patrick Stein, Jacob D. Furst, Daniela Raicu
ICMLA (2)3
2012 Expanding diagnostically labeled datasets using content-based image retrieval
abstract
In computer-aided diagnosis (CAD), having an accurate ground truth is critical. However, the number of databases containing medical images with diagnostic information is limited. Using pulmonary computed tomography (CT) scans, we develop a content-based image retrieval (CBIR) approach to exploit the limited images with diagnostically labeled data in order to annotate unlabeled images with diagnoses. By applying this CBIR method iteratively, we expand the set of diagnosed data available for CAD systems. We evaluate the method by implementing a CAD system that uses undiagnosed lung nodules as queries and retrieves similar nodules from the diagnostically labeled dataset. In calculating the precision of this system, radiologist- and computer-predicted malignancy data are used as ground truth for the undiagnosed query nodules. Our results indicate that CBIR expansion is an effective method for labeling undiagnosed images in order to improve the performance of CAD systems.
Anne-Marie Giuca, Kerry A. Seitz Jr., Jacob D. Furst, Daniela Raicu
ICIP3
2011 Top issues in providing successful undergraduate research experiences
abstract
Undergraduate research is becoming increasingly common in colleges and universities, and, to support this, there is a need to have best practices and forums for promoting exchange of ideas. In particular, a working group at a recent National Science Foundation (NSF) Computer and Information Science and Engineering (CISE) Research Experiences for Undergraduates (REU) sites PI's meeting identified four important issues in undergraduate research: 1) how to design a good research project, 2) how to prepare students for research, 3) how to measure outcomes of undergraduate research and 4) incentives for undergraduates to publish as result of their participation in research. The panelists have all served as PIs or Co-PIs on NSF REU projects in computing and have mentored many undergraduates in a large variety of research projects both in REU settings as well as during the regular academic year. They will each address one of the issues identified above, and share their expertise in addressing the issue, providing solid guidance to anyone interested in promoting undergraduate research. A significant amount of time will be set aside for audience participation and discussion.
Hans-Peter Bischof, Jacob D. Furst, Daniela Raicu, Susan Darling Urban
SIGCSE2
2009 A statistical analysis of the effects of CT acquisition parameters on low-level features extracted from CT images of the lung
abstract
We propose a solution for automatic classification of lung nodules in an environment with heterogeneous computed tomography (CT) acquisition parameters. Such a classification system needs to take into account the differences in CT acquisition parameters used when obtaining and processing each medical image. Using analysis of variance (ANOVA), our current research proposes to better understand the effects of CT acquisition parameters on predicting various semantic characteristics (such as spiculation, subtlety, and margin) used in the diagnosis interpretation process. All of the parameters were found to affect the low-level image features used in the classification models of these semantic characteristics. When this knowledge is used to normalize those parameters, the final semantic model will become unaffected by the CT acquisition parameters.
Joseph S. Wantroba, Daniela Raicu, Jacob D. Furst
ICIP3
2009 Enhancing undergraduate education: a REU model for interdisciplinary research
abstract
This paper presents a successful model for undergraduate research where student participants work on interdisciplinary research projects; in our case, at the frontier between computer science and medicine. Students are part of research teams comprised of other undergraduates, graduate students, faculty and medical experts, participate in professional development and training activities within the larger group, and disseminate their results at the host institutions or conferences specific to the interdisciplinary focus. The model outcomes at the end of the first three years (2005-2007) indicate that the interdisciplinary model successfully 1) expanded the student participation in research by recruiting students who might not otherwise have research opportunities, 2) attracted a diversified pool of talented students into science, 3) promoted interdisciplinary undergraduate studies in computer science and medical informatics as well as in future graduate studies; and 4) trained students in all phases of research, including writing and presenting research papers at conferences.
Daniela Raicu, Jacob D. Furst
SIGCSE2
2009 Comparison and Evaluation of Methods for Liver Segmentation From CT Datasets
abstract
This paper presents a comparison study between 10 automatic and six interactive methods for liver segmentation from contrast-enhanced CT images. It is based on results from the "MICCAI 2007 Grand Challenge" workshop, where 16 teams evaluated their algorithms on a common database. A collection of 20 clinical images with reference segmentations was provided to train and tune algorithms in advance. Participants were also allowed to use additional proprietary training data for that purpose. All teams then had to apply their methods to 10 test datasets and submit the obtained results. Employed algorithms include statistical shape models, atlas registration, level-sets, graph-cuts and rule-based systems. All results were compared to reference segmentations five error measures that highlight different aspects of segmentation accuracy. All measures were combined according to a specific scoring system relating the obtained values to human expert variability. In general, interactive methods reached higher average scores than automatic approaches and featured a better consistency of segmentation quality. However, the best automatic methods (mainly based on statistical shape models with some additional free deformation) could compete well on the majority of test images. The study provides an insight in performance of different segmentation approaches under real-world conditions and highlights achievements and limitations of current image analysis techniques.
Tobias Heimann, Bram van Ginneken, Martin Styner, Yulia Arzhaeva, Volker Aurich, Christian Bauer 0001, Andreas Beck 0001, Christoph Becker 0002, Reinhard Beichel, György Bekes, Fernando Bello, Gerd Karl Binnig, Horst Bischof, Alexander Bornik, Peter Cashman, Ying Chi, Andrés Cordova, Benoit M. Dawant, Márta Fidrich, Jacob D. Furst, Daisuke Furukawa, Lars Grenacher, Joachim Hornegger, Dagmar Kainmüller, Richard Kitney, Hidefumi Kobatake, Hans Lamecker, Thomas Lange, Brian Lennon, Rui Li 0012, Senhu Li, Hans-Peter Meinzer, Gábor Németh, Daniela Raicu, Anne-Mareike Rau, Eva M. van Rikxoort, Mikaël Rousson, László Ruskó, Kinda Anna Saddi, Günter Schmidt 0001, Dieter Seghers, Akinobu Shimizu, Pieter Slagmolen, Erich Sorantin, Grzegorz Soza, Ruchaneewan Susomboon, Jonathan M. Waite, Andreas Wimmer, Ivo Wolf
IEEE Trans. Medical Imaging20
2007 Oligonucleotide microarray identification of Bacillus anthracis strains using support vector machines
abstract
The capability of a custom microarray to discriminate between closely related DNA samples is demonstrated using a set of Bacillus anthracis strains. The microarray was developed as a universal fingerprint device consisting of 390 genome-independent 9mer probes. The genomes of B. anthracis strains are monomorphic and therefore, typically difficult to distinguish using conventional molecular biology tools or microarray data clustering techniques. Using support vector machines (SVMs) as a supervised learning technique, we show that a low-density fingerprint microarray contains enough information to discriminate between B. anthracis strains with 90% sensitivity using a reference library constructed from six replicate arrays and three replicates for new isolates.
Michael V. Doran, Daniela Raicu, Jacob D. Furst, Raffaella Settimi, Matthew Schipma, Darrell P. Chandler
Bioinform.3
2006 Automatic Single-Organ Segmentation in Computed Tomography Images
abstract
In this paper, we propose a hybrid approach for automatic single-organ segmentation in computed tomography (CT) data. The approach consists of three stages: first, a probability image of the organ of interest is obtained by applying a binary classification model obtained using pixel-based texture features; second, an adaptive split-and-merge segmentation algorithm is applied on the organ probability image to remove the noise introduced by the misclassified pixels; and third, the segmented organ's boundaries from the previous stage are iteratively refined using a region growing algorithm. While we applied our approach for liver segmentation in 2-D CT images, a challenging and important task in many medical applications, the proposed approach can be applied for the segmentation of any other organ in CT images. Moreover, the proposed approach can be extended to perform automatic multiple organ segmentation and to build context-sensitive reporting tools for computer-aided diagnosis applications.
Ruchaneewan Susomboon, Daniela Raicu, Jacob D. Furst, David S. Channin
ICDM3
2005 Wavelet-Based Texture Classification of Tissues in Computed Tomography
abstract
The research presented in this article is aimed at developing an automated imaging system for classification of tissues in medical images. The article focuses on using texture analysis for the classification of tissues from CT scans. The approach consists of two steps: automatic extraction of the most discriminative texture features of regions of interest in the CT medical images and creation of a classifier that will automatically identify the various tissues. A comparative study of wavelets-based texture descriptors from three families of wavelets (Haar, Daubechies, Coiflets), coupled with the implementation of a decision tree classifier based on the Classification and Regression Tree (C&RT) approach is carried on. Preliminary results for a 3D data set from normal chest and abdomen CT scans are presented.
Lindsay Semler, Lucia Dettori, Jacob D. Furst
CBMS3
2005 Texture-Based Image Retrieval for Computerized Tomography Databases
abstract
In this paper we propose a content-based image retrieval (CBIR) system for retrieval of normal anatomical regions present in computed tomography (CT) studies of the chest and abdomen. We implement and compare eight similarity measures using local and global cooccurrence texture descriptors. The preliminary results are obtained using a CT database consisting of 344 CT images representing the segmented heart and great vessels, liver, renal and splenic parenchyma, and backbone from two different patients. We evaluate the results with respect to the retrieval precision metric for each of the similarity measures when calculated per organ and overall.
Winnie Tsang, Andrew Corboy, Ken Lee, Daniela Raicu, Jacob D. Furst
CBMS5
2005 A classification approach for anatomical regions segmentation
abstract
In this paper, a supervised pixel-based classifier approach for segmenting different anatomical regions in abdominal computed tomography (CT) studies is presented. The approach consists of three steps: texture extraction, classifier creation, and anatomical regions identification. First, a set of co-occurrence texture descriptors is calculated for each pixel from the image data sample; second, a decision tree classifier is built using the texture descriptors and the names of the tissues as class labels. At the conclusion of the classification process, a set of decision rules is generated to be used for classification of new pixels and identification of different anatomical regions by joining adjacent pixels with similar classifications. It is expected that the proposed approach will also help automate different semi-automatic segmentation techniques by providing initial boundary points for deformable models or seed points for split and merge segmentation algorithms. Preliminary results obtained for normal CT studies are presented.
Mikhail Kalinin, Daniela Raicu, Jacob D. Furst, David S. Channin
ICIP (2)3
2002 A Direct Method for Positioning the Arms of a Human Model
John McDonald 0004, Karen Alkoby, Roymieco Carter, Juliet Christopher, Mary Jo Davidson, Dan Ethridge, Jacob D. Furst, Damien Hinkle, Glenn Lancaster, Lori Smallwood, Nedjla Ougouag, Jorge Toro, Rosalee J. Wolfe
Graphics Interface7
2002 Optimal Parameter Height Ridges
Jacob D. Furst, Stephen M. Pizer
J. Vis. Commun. Image Represent.1
2001 An improved articulated model of the human hand
John McDonald 0004, Jorge Toro, Karen Alkoby, André Berthiaume, Roymieco Carter, Pattaraporn Chomwong, Juliet Christopher, Mary Jo Davidson, Jacob D. Furst, Brian Konie, Glenn Lancaster, Lopa Roychoudhuri, Eric Sedgwick, Noriko Tomuro, Rosalee J. Wolfe
Vis. Comput.9
1998 Marching Optimal-Parameter Ridges: An Algorithm to Extract Shape Loci in 3D Images
Jacob D. Furst, Stephen M. Pizer
MICCAI1