EDBT 2026 Demo / reviewers in the wild / expert
Daniela Raicu
dblp:32/1675 · also Daniela Stan, Daniela Stan Raicu
· DBLP profile ↗
32ranked-venue papers
2as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 7 since 2021Artificial intelligence and machine learning · 10 · 4 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing Pulmonary Nodule Localization Based on Latent RepresentationsabstractAccurately localizing pulmonary nodules relative to other anatomical structures is crucial for disease management, guiding biopsies, and formulating effective treatment strategies. This study introduces a fully automated approach for classifying nodules detected in computed tomography (CT) images as pleural (near the pleura as$N_{p}$) or non-pleural (distant from the pleura as$D_{p}$). We propose a combination of Principal Component Analysis (PCA) and deep learning approach to determine a threshold for the optimal correlation between the original image and the latent PCA-based image reconstruction used to classify lung nodule as$N_{p}$or$D_{p}$. Applying our methodology to the Lung Image Database Consortium and Image Database Resource Initiative (LIDC-IDRI) dataset, we found that the PCA-based approach demonstrated substantial agreement with human evaluations of the proximity of the lung nodule to the pleura, achieving a Cohen's Kappa value of 0.76, outperforming an intensitybased baseline model with a Cohen's Kappa value of 0.55. Further analysis revealed significant variability in radiologists' semantic ratings for$N_{p}$versus$D_{p}$, with the highest variability observed in the texture feature. These findings demonstrate that our approach has the potential to enhance the classification of lung nodule localization when integrated into computer-aided diagnosis (CAD) systems. Charmi Patel, Yiyang Wang 0003, Roselyne Tchoua, Thiruvarangan Ramaraj, Jacob D. Furst, Daniela Raicu |
CBMS | 6 |
| 2025 | Novel Perspective on Ensemble Clustering for Persistent Cluster Patterns: A Case Study in Disease Cluster DiscoveryabstractClustering is the process of finding natural groups within a data set such that patterns within a group are more similar than patterns belonging to different groups. It has been used in a wide range of scientific and engineering disciplines. Yet, clustering is also a difficult unsupervised problem without an absolute ground truth. In practice, the "best" clustering method is the one that produces the most interpretable results. There is no universal optimum way to select the number of clusters. Our ultimate goal is to enable domain experts in high-stakes fields to make informed decisions when choosing a final, insightful, and actionable clustering solution for complex problems. In such challenging scenarios, ensemble clustering is a technique where multiple clustering results are combined to produce a more robust and stable final clustering. Here, inspired by a medical application, we take a novel approach to ensemble clustering; rather than focusing on an optimum global clustering, we identify persistent cluster(s) by multiple techniques across multiple clustering experiments among patients from Rush University Medical Center in Chicago. The key novelty resides in reducing the uncertainty in clustering, in the absence of ground truth, by looking at persistent clusters across the ensemble clustering. In healthcare, physicians aim to treat a patient based on a series of symptoms that contribute to a single disease. Because most clinical guidelines focus on individual diseases, managing patients with multimorbidity—when a person has multiple (chronic) conditions—can be challenging. We aim to discover clusters of diseases with some level of confidence to assist physicians in diagnosing and treating patients with multimorbidity. In this case study, we identify and thoroughly characterize a cluster of diseases that affects primarily women and persists through clustering the data from k=16 to k=20 clusters using three clustering techniques. Additionally, our first proposed multimorbidity cluster identified a network connecting asthma, breast cancer, hip/pelvic fracture, endometrial cancer, and non-Alzheimer’s dementia. We use multiple clustering techniques, consensus metrics, as well as graph analysis and visualization to provide confidence and foster trust from domain experts in our results. We are working with physicians to provide a medical explanation for this data-driven disease multimorbidity discovery. Fernanda Cabral, Maria Alexandra Hubbard, Rudhvish Patel, Adeola Badmos, Daniela Raicu, Raj Shah, Roselyne Tchoua |
eScience | 5 |
| 2025 | Leveraging Hidden Patterns in Open-Ended Community Health Workers' Notes to Improve Prediction of Patient ReadmissionabstractPatient readmissions to Emergency Departments (EDs) pose significant challenges to healthcare systems, often indicating suboptimal care transitions, inadequate patient education, or insufficient post-discharge support. These unplanned readmissions not only compromise patient outcomes but also contribute to escalating healthcare costs. In response, healthcare providers are increasingly seeking predictive models to identify at-risk patients and implement preventive strategies. While advanced deep learning models have shown promise in predicting readmission risks, their "black-box" nature and substantial data requirements often limit clinical applicability due to a lack of interpretability. This study investigates the integration of Machine Learning (ML) and natural language processing (NLP) techniques to enhance the accuracy and explainability of patient readmission risk predictions. Utilizing data from the Sinai Urban Health Institute (SUHI), we analysed both structured data including demographics, interaction logs, social determinants of health (SDoH) survey answers and unstructured data, specifically notes capturing conversations between patients and Community Health Workers (CHWs). Our findings indicate that incorporating unstructured textual data improved model performance, with the area under the receiver operating characteristic curve (AUC) increasing from 0.68 to 0.74. This enhancement suggests that patient-CHW conversations capture critical, non-medical factors influencing readmissions, such as personal needs and social support deficits, which are not typically recorded in standard medical records. The study underscores the value of integrating patient-CHW contact notes into predictive modelling, helping to highlight the important role of community health in informing targeted interventions to reduce preventable readmissions. Naveen Kumar Reddy Veeramreddy, Ankita Mishra, Navika Maglani, Sameer Shaik, Kelly McCabe, Jacob D. Furst, Daniela Raicu, Roselyne Tchoua, Jamshid Sourati |
eScience | 7 |
| 2025 | Video Label Refinement for Temporal LocalizationabstractWhen performing temporal localization, achieving high precision results or high accuracy at the highest temporal Intersection over Union (tIoU) requires training a model with class labels that have precise boundaries. Obtaining high-precision class labels solely from human annotators presents challenges, such as the difficulty of accurately discerning when a movement begins and ends by watching a video. This work expands on an approach improving the temporal boundaries of class labels for activity recognition in video to handle complex movement patterns in human interactions. This method applies signal processing methods to motion features to identify shapes in signals associated with activities to be classified and adjust the boundaries accurately. When these methods are applied to existing human-provided activity annotations, they can improve the accuracy of class label boundaries. This improves temporal localization model performance at the high tIoUs. Jennifer Piane, Thiruvarangan Ramaraj, Jacob D. Furst, Daniela Raicu |
ICME | 4 |
| 2023 | Curriculum gDRO: Improving Lung Malignancy Classification through Robust Curriculum Task LearningabstractDeep learning models used in Computer-Aided Diagnosis (CAD) systems are often trained with Empirical Risk Minimization (ERM) loss. These models often achieve high overall classification accuracy but with lower classification accuracy on certain subgroups. In the context of lung nodule malignancy classification task, these atypical subgroups exist due to the lung cancer heterogeneity. In this study, we characterize lung nodule malignancy subgroups using the malignancy likelihood ratings given by radiologists and improve the worst subgroup performance by utilizing group Distributionally Robust Optimization (gDRO). However, we noticed that gDRO improves on worst subgroup performance from the benign category, which has less clinical importance than improving classification accuracy for a malignant subgroup. Therefore, we propose a novel curriculum gDRO training scheme that trains for an “easy” task (nodule malignancy is determinate or indeterminate for radiologists) first, then for a “hard” task (malignant, benign, or indeterminate nodule). Our results indicate that our approach boosts the worst group subclass accuracy from the malignant category, by up to 6 percentage points compared to standard methods that address and improve worst group classification performance. Arun Sivakumar, Yiyang Wang 0003, Roselyne Tchoua, Thiruvarangan Ramaraj, Jacob D. Furst, Daniela Raicu |
CBMS | 6 |
| 2022 | Text Summarization towards Scientific Information ExtractionabstractDespite the exponential growth in scientific textual content, publications remain the primary means of disseminating vital research to experts within their respective fields. These texts are predominantly written for human consumption, resulting in two fundamental challenges; experts cannot efficiently remain well-informed to leverage the latest discoveries, and applications which rely on valuable insights buried in these texts cannot effectively build upon published results. Consequently, scientific progress stalls. Automatic Text Summarization (ATS) and Information Extraction (IE) are two essential fields which address this problem. While the two research topics are often studied independently, this work proposes to look at ATS in the context of IE, specifically as it relates to Scientific IE. However, Scientific Information Extraction faces several challenges; chiefly, the scarcity of relevant entities and insufficient training data. In this paper, we focus on extractive ATS, which identifies the most valuable sentences from textual content for the purpose of ultimately extracting scientific relations. We account for the associated challenges by means of an ensemble method through the integration of three weakly supervised learning models, one for each entity of the target relation. Notably, while the relation is well defined, we do not require previously annotated data for the entities composing the relation. The central objective is to generate balanced training data, which many advanced natural language processing models require. We apply this idea in the domain of materials science, extracting the polymer-glass transition temperature relation and achieve 94.7% recall (i.e., sentences which contain relations annotated by humans), while reducing the text by 99.3% of the original document. Abigail Keller, Jacob D. Furst, Daniela Raicu, Peter M. Hastings, Roselyne Tchoua |
e-Science | 3 |
| 2022 | Predicting ME/CFS After Infectious Mononucleosis Using Cytokine Network CorrelationsabstractWe investigated if a predictive modeling strategy based on the interdependence of the cytokine network could accurately predict if a patient would develop Myalgic Encephalomyelitis/Chronic Fatigue Syndrome (ME/CFS) after contracting infectious mononucleosis (IM). We analyzed previously collected data from Northwestern University (NU) students in a three-stage experiment, following them from the start of the school year (Stage 1), to development of IM (Stage 2), to six months post development of IM (Stage 3). At all three stages, blood was stored from participants for cytokine measurement and analysis. Additionally, eight psychological and behavioral scales were used to identify participants as healthy controls or as ME/CFS. Using participants’ measured cytokine expression levels, we built a predictive model based on the inherent correlations within the cytokine network. We found that we could predict ME/CFS in patients 6 months after IM with 86.84% accuracy using correlation matrices made from cytokines taken during IM infection. These results suggest that there may be potential in using an approach that is based on the interdependence of the cytokine network to predict ME/CFS post IM. Future work may explore the validity of these findings and if such an approach could have applications in other diseases. Jennifer Schwabe, Chelsea Hua, Emma M. Allen, Leonard A. Jason, Jacob D. Furst, Daniela Raicu |
ICMLA | 6 |
| 2022 | Lung Nodule Malignancy Subtype Discovery with Semantic LearningabstractComputer-aided diagnosis (CAD) systems have been widely used as second readers in lung cancer diagnosis. However, lung cancer heterogeneity and lack of using human annotated semantic characteristics are significant obstacles to an accurate and explainable CAD outcome. We propose a novel CAD scheme that characterizes lung nodule malignancy subtypes semantically and classifies nodule malignancy through a semantic learning process. We built and evaluated our method on a publicly available dataset, Lung Image Database Consortium (LIDC). We discovered and characterized two malignant nodule subtypes and two benign subtypes. In addition, we achieved a significantly improved malignancy classification result by incorporating the semantic learning process when compared with a result without using the semantic learning. This study suggests that future CAD systems should include disease heterogeneity as a model input to refine the model explainability; our classification result provides case-base evidence that learning image semantic explanations is valuable for improving CAD accuracy. While we tested our approach on one dataset, the proposed approach is applicable to other CAD tasks if semantic annotations are available. Yiyang Wang 0003, Bowen Qiu, Thiruvarangan Ramaraj, Ilyas Ustun, Jacob D. Furst, Daniela Raicu |
ICPR | 6 |
| 2021 | Approaches to Evaluating Eye Gaze Patterns between Physician-Patient Interaction in Primary Care ClinicabstractAssessing nonverbal communication in primary care settings may provide insights into health disparities and inform the design of health information technologies in those settings. The aim of this study was to develop approaches to measure and compare a single nonverbal communication behavior (eye gaze) between physicians and their patients in primary care clinics, and to assess whether physician gaze is consistent between patients and with other physicians. This analysis method could lead to design guidelines for technologies and more effective assessments of interventions. Data came from a study that included 11 physicians and 77 patients from Federally Qualified Health Centers (FQHCs) in Chicago, Illinois. We propose two approaches, K-means Clustering and Dynamic Time Warping (DTW), to evaluate eye gaze behavior patterns consistency between physicians and patients. Results from K-means clustering showed some consistency of gaze patterns across patients within the same physician in some visits. Furthermore, some similarities of gaze patterns across patients of different physicians were found. The t-test results from DTW showed some consistency of gaze patterns across physicians at p-value of 0.05 and 0.01. DTW results showed the viability of the approach of evaluating the consistency in physician-patient interaction. Additional approaches to evaluate eye gaze patterns, such as gaze patterns with technology, may also be considered to better understand the interaction between physicians and patients. Amal N. Almansour, Jacob D. Furst, Daniela Raicu, Enid Montague |
BIBM | 3 |
| 2021 | Ensemble Labeling towards Scientific Information Extraction (ELSIE) - Blob ExtractionabstractScientific publications constitute an extremely valuable repository of knowledge and collection of facts crucial to the advancement of science and development of applications, which grows as researchers learn from previous works and scientists use results in the literature to design and create. With the exponential growth of available publications, reading and extracting this wealth of information has become impractical for humans. Despite great progress in natural language processing, machine-learned solutions require large amounts of carefully annotated data for good performance. This is especially true in the context of accurately labeling and extracting complex scientific data. Towards our ultimate goal of extracting scientific facts from the literature, we first aim to identify blobs of text that contain all of the facts in a publication to be later automatically extracted or scrutinized by experts. Our previous work identified some facts missed by experts yet missed others due to the assumption that the target relation—here, a polymer and its glass transition temperature—would be contained within the same sentence. We set out to enhance our approximate labeling system to look back and ahead for missing information and successfully achieved 100% recall of scientific facts while reducing the full-text publication to 6% of its original size. Moreover, we assign confidence scores to sentences to further assist expert curators in identifying important sentences and facts locked in unstructured text. Erin Murphy, Alexander Rasin, Jacob D. Furst, Daniela Raicu, Roselyne Tchoua |
e-Science | 4 |
| 2020 | Heatmap Template Generation for COVID-19 Biomarker Detection in Chest X-raysabstractDetecting and identifying patterns in chest X-ray images of Covid-19 patients are important tasks for understanding the disease and for making differential diagnosis. Given the relatively small number of available Covid-19 X-ray images and the need to make progress in understanding the disease, we propose a transfer learning technique applied to a pretrained VGG19 neural network to build a deep convolutional model capable of detecting four possible conditions: normal (healthy), bacteria, virus (not Covid-19), and Covid-19. The transformation of the multi-class deep learning output into binary outputs and the detection of Covid-19 image patterns using Grad-CAM technique show promising results. The discovered patterns are consistent across images from a given class of disease and constitute explanations of how the deep learning model makes classification decisions. In the long run, the identified patterns can serve as biomarkers for a given disease in chest X-ray images. Mirtha Lucas, Miguel Lerma, Jacob D. Furst, Daniela Raicu |
BIBE | 4 |
| 2020 | Robust Physician Gaze Prediction Using a Deep Learning ApproachabstractThe patient-physician relationship is an integral part of primary care visits. To build a better relationship, understanding the communication between patient and physician is the key. This study focused on analyzing the gaze, one of the most important non-verbal behaviors found to influence patient outcomes. Gaze analysis often needs a manual rating process which might be time-consuming, costly, and unreliable. This research aimed to support automated analysis of physician-patient interaction using a deep convolutional neural network with transfer learning to a build robust model for physician gaze prediction. Utilizing only 3 minutes of 15 videos capturing 3 physicians interacting with different patients in a clinical setting, the model achieved over 98% accuracy for train, test, and validation sets. By visualizing the convolutional layers and comparing sample frames from different interactions, results highlighted several patterns shared across frames predicted correctly from both seen and unseen video sequences. The proposed work has the potential to informed the future design of technologies used to capture the clinical interaction and provide real-time feedback for physicians, which will contribute to the improvement of care quality. Tianyi Tan, Enid Montague, Jacob D. Furst, Daniela Raicu |
BIBE | 4 |
| 2020 | Explainable Deep Learning for Biomarker Classification of OCT ImagesabstractAdvanced form of age-related macular degeneration (AMD) is a major health burden that can lead to irreversible vision loss in the elderly population. The early signs of AMD are drusen, which appear as yellowish deposits under the retina. The end-stages of AMD include two forms: wet AMD (neovascular) and geographic atrophy (GA, non-neovascular). We propose a deep learning approach using a pre-trained VGG-19 neural network to classify the optical coherence tomography (OCT) images into wet AMD, GA, drusen, and healthy images. To explain the results, we quantify and present the prediction confidence and reliability as well as the regions of interest deemed to be important by the deep learning model when classifying OCT images. The sensitivity of classification for wet AMD, GA, drusen, and healthy images was 91.67%, 88.00%, 96.30% and 100% respectively, with the most confident predictions being for the drusen and healthy images. Visual inspection of the misclassified images using heatmaps using the Gradient-weighted Class Activation Mapping (Grad-CAM) algorithm revealed that, for most of the misclassified images that had high prediction confidence, the algorithm identified correctly more than one region of interest, each belonging to a different AMD category. We concluded that, rather than assigning just one label to an AMD image, algorithms for AMD classification should allow multi-labels as images generally show evidence of more than one stage of AMD simultaneously (e.g., drusen as the predominant region and GA as a small region of interest). Yiyang Wang 0003, Mirtha Lucas, Jacob D. Furst, Amani A. Fawzi, Daniela Raicu |
BIBE | 5 |
| 2020 | Enhancing Recall Using Data Cleaning for Biomedical Big DataabstractIn clinical practice, large amounts of heterogeneous medical data are generated on a daily basis. This data has the potential to be used for biomedical research and as a diagnostic reference for physicians. However, leveraging heterogeneous data for analysis requires integrating it first. Integration process includes a pre-processing data cleaning phase that eliminates inconsistencies and errors originating from each data source. In this paper, we describe a workflow for cleaning heterogeneous biomedical data sources. Our novel data cleaning approach can be applied for replacement of missing text and to improve the number of relevant cases retrieved by search queries. When the threshold for missing category replacement is met, our results show that our method achieves a missing content replacement precision of 85%, which represents an improvement of 18% over the baseline state of our datasets. Priya Deshpande, Alexander Rasin, Roselyne Tchoua, Jacob D. Furst, Daniela Raicu, Sameer K. Antani |
CBMS | 5 |
| 2020 | Assessment of Medical Reports Uncertainty through Topic Modeling and Machine LearningabstractMedical uncertainty is identified as one of the most important factors leading to miscommunication between health care providers and patients. Along with the rapid growth of Natural Language Processing and Machine Learning techniques, more opportunities to understand medical uncertainty became available, including quantifying and modeling medical uncertainty. In this work, we obtained more than 20,000 radiology teaching files as medical reports from multiple resources. Uncertainty ontologies were also obtained from two resources. We expanded the uncertainty term list by applying topic model Word2Vec to identify similar terms to original uncertainty ontologies. The teaching files were quantified into 5 uncertainty level classes by the sum of the Term Frequency - Inverse Document Frequency (TF-IDF) of the expanded uncertainty terms list. Results from topic modelling were used to produce features. The product of TF-IDF results of uncertainty terms in teaching files and the topic model results were then used to train classifiers to predict the uncertainty level in medical reports. Our exhaustive experimental analysis showed that Decision Tree can classify uncertainty level of medical reports at overall accuracy 82%, which is higher than K-nearest Neighbor (80%) and Naive Bayes (75%). The model can be used to identify medical reports' uncertainty level and limit miscommunication between parties and reduce diagnostic errors. Mengyuan Shang, Jacob D. Furst, Daniela Raicu |
CBMS | 3 |
| 2019 | Hand-Eye Coordination: Automating the Annotation of Physician-Patient InteractionsabstractThe widespread adoption of electronic health records within clinical settings has renewed interest in understanding physician-patient interactions. Previous work analyzing clinical interactions has mostly coupled patient surveys with manually annotated video interactions provided by human coders. Physician gaze is among the components of the non-verbal interaction which has been found to impact patient outcomes. The work described in this paper illustrates an automated system for video labeling of patient-physician interactions and shows that image features (such as areas and positioning of physicians' hands) can provide important visual aids for learning physician gaze with over 90% accuracy. While our approach focuses on physician gaze, it can be extended to capture other clinical human-human and human-technology interactions as well as connect these interactions to patient ratings of clinical interactions. Daniel Gutstein, Enid Montague, Jacob D. Furst, Daniela Raicu |
BIBE | 4 |
| 2019 | Optical Flow, Positioning, and Eye Coordination: Automating the Annotation of Physician-Patient InteractionsabstractThe widespread adoption of electronic health records within clinical settings has renewed interest in understanding physician-patient interactions. Previous work analyzing clinical interactions has mostly coupled patient surveys with manually annotated video interactions provided by human coders. Physician gaze is among the components of the non-verbal interaction which has been found to impact patient outcomes. The work described in this paper illustrates an automated system for multi-video labeling of patient-physician interactions and shows that image features (in the form of body positioning coordinates and optical flow) can provide important visual aids for learning physician gaze with over 90% accuracy. While our approach focuses on physician gaze, it can be extended to capture other clinical human-human and human-technology interactions as well as connect these interactions to patient ratings of clinical interactions. Daniel Gutstein, Enid Montague, Jacob D. Furst, Daniela Raicu |
BIBM | 4 |
| 2018 | Automatic extraction of informal topics from online suicidal ideationabstractBACKGROUND: Suicide is an alarming public health problem accounting for a considerable number of deaths each year worldwide. Many more individuals contemplate suicide. Understanding the attributes, characteristics, and exposures correlated with suicide remains an urgent and significant problem. As social networking sites have become more common, users have adopted these sites to talk about intensely personal topics, among them their thoughts about suicide. Such data has previously been evaluated by analyzing the language features of social media posts and using factors derived by domain experts to identify at-risk users. RESULTS: In this work, we automatically extract informal latent recurring topics of suicidal ideation found in social media posts. Our evaluation demonstrates that we are able to automatically reproduce many of the expertly determined risk factors for suicide. Moreover, we identify many informal latent topics related to suicide ideation such as concerns over health, work, self-image, and financial issues. CONCLUSIONS: These informal topics topics can be more specific or more general. Some of our topics express meaningful ideas not contained in the risk factors and some risk factors do not have complimentary latent topics. In short, our analysis of the latent topics extracted from social media containing suicidal ideations suggests that users of these systems express ideas that are complementary to the topics defined by experts but differ in their scope, focus, and precision of language. Reilly Grant, David Kucher, Ana M. Leon, Jonathan Gemmell, Daniela Raicu, Samah Jamal Fodeh |
BMC Bioinform. | 5 |
| 2016 | Compression-based distance methods as an alternative to statistical methods for constructing phylogenetic treesabstractDistance based methods for constructing phylogenetic trees have long been considered inconsistent and inferior to the more dominant statistical methods. However, use of compression methods specific to DNA could prove valuable in improving the effectiveness of distance based methods. To demonstrate the validity of distance-based methods when utilizing current DNA compression algorithms, such as MFCompress, we have applied such a method to datasets of closely related species of fish from the suborder Labroidei and to strains of Ebola. In both cases, we have managed to produce trees that are either very similar or identical to published trees produced using statistically based methods. This suggests that distance based methods can perform comparably to statistically based methods without requiring as much pre-processing of original DNA sequences or system resources. Additionally, the results also stress the importance of using accurate methods of calculating species distance due to the way that one specific DNA compression algorithm, MFCompress, consistently and convincingly managed to outperform other popular, general use compression algorithms. Mohamed El-Dirany, Forrest Wang, Jacob D. Furst, Daniela Raicu |
BIBM | 5 |
| 2016 | C. elegans search behavior analysis using Multivariate Dynamic Time WarpingabstractThe nematode Caenorhabditis elegans is widely used as a genetic model organism because of its relevance to human biology and disease. Quantifying C. elegans movement behavior is important for understanding the differences between genotypes under various food conditions. One of the most common ways to quantify C. elegans movement is through path analysis. However, comparing the movement paths between different C. elegans genotypes is challenging because of the variations in the trajectories' time lengths. We propose a novel approach that combines Multivariate Dynamic Time Warping and step-length path encoding to cluster C. elegans movement path data and further show that the proposed combination avoids the need for extracting complex path features. Yiyang Wang 0003, Carleton Smith, Mingfei Shao, Jacob D. Furst, Daniela Raicu, Hongkyun Kim |
BIBM | 6 |
| 2014 | Machine-Sourced Segmentations vs. Expert-Sourced Segmentations for the Classification of Lung Nodules with Outlier RemovalabstractComputer-aided diagnosis systems can provide additional opinions that serve as an aid to radiologists in the early detection of lung nodules. Previous CAD models have relied on radiologist-delineated contours to extract image features and classify lung nodules into semantic ratings. Manually creating these contours can be time-consuming and expensive. This paper proposes a different CAD system based on multiple machine-sourced segmentations that can provide semantic ratings at least as accurate as a panel of experts in order to aid in the diagnostic process. However, the mass production of machine-sourced segmentations may sometimes produce unwanted noise. Therefore, we propose to filter out the bad segmentations by applying an outlier detection algorithm that identifies segmentations that are far away from the majority of the segmentations. Our results are compared to a CAD system based on expert-sourced contours and a reference truth generated by radiologists' semantic ratings. Using the Lung Image Database Consortium dataset, we show that machine-sourced segmentations provide predictions at least as good as expert-sourced segmentations and how outlier removal affects mostly shape-dependent semantic ratings. Mayra Alejandra Molina Puentes, Jacob D. Furst, Daniela Raicu |
ICMLA | 3 |
| 2013 | Weak Segmentations and Ensemble Learning to Predict Semantic Ratings of Lung NodulesabstractComputer-aided diagnosis (CAD) can be used as "second readers" in the imaging diagnostic process. Typically to create a CAD system, the region of interest (ROI) has to be first detected and then delineated. This can be done either manually or automatically. Given that manually delineating ROIs is a time consuming and costly process, we propose a CAD system based on multiple computer-derived weak segmentations (WSCAD) and show that its diagnosis performance is at least as good as the predictions developed using manual radiologist segmentations. The proposed CAD system extracts a set of image features from the weak segmentations and uses them in an ensemble of classification algorithms to predict semantic ratings such as malignancy. These automated results are compared against a reference truth based on ratings and segmentations provided by radiologists to determine if it is necessary to obtain manual radiologist segmentations in order to develop a CAD. By developing a pair of CADs using the Lung Image Database Consortium (LIDC) data, we show that WSCADs are at least as accurate in predicting semantic ratings as CADs based on radiologist segmentation. Ethan Smith, Patrick Stein, Jacob D. Furst, Daniela Raicu |
ICMLA (2) | 4 |
| 2012 | Expanding diagnostically labeled datasets using content-based image retrievalabstractIn computer-aided diagnosis (CAD), having an accurate ground truth is critical. However, the number of databases containing medical images with diagnostic information is limited. Using pulmonary computed tomography (CT) scans, we develop a content-based image retrieval (CBIR) approach to exploit the limited images with diagnostically labeled data in order to annotate unlabeled images with diagnoses. By applying this CBIR method iteratively, we expand the set of diagnosed data available for CAD systems. We evaluate the method by implementing a CAD system that uses undiagnosed lung nodules as queries and retrieves similar nodules from the diagnostically labeled dataset. In calculating the precision of this system, radiologist- and computer-predicted malignancy data are used as ground truth for the undiagnosed query nodules. Our results indicate that CBIR expansion is an effective method for labeling undiagnosed images in order to improve the performance of CAD systems. Anne-Marie Giuca, Kerry A. Seitz Jr., Jacob D. Furst, Daniela Raicu |
ICIP | 4 |
| 2011 | Top issues in providing successful undergraduate research experiencesabstractUndergraduate research is becoming increasingly common in colleges and universities, and, to support this, there is a need to have best practices and forums for promoting exchange of ideas. In particular, a working group at a recent National Science Foundation (NSF) Computer and Information Science and Engineering (CISE) Research Experiences for Undergraduates (REU) sites PI's meeting identified four important issues in undergraduate research: 1) how to design a good research project, 2) how to prepare students for research, 3) how to measure outcomes of undergraduate research and 4) incentives for undergraduates to publish as result of their participation in research. The panelists have all served as PIs or Co-PIs on NSF REU projects in computing and have mentored many undergraduates in a large variety of research projects both in REU settings as well as during the regular academic year. They will each address one of the issues identified above, and share their expertise in addressing the issue, providing solid guidance to anyone interested in promoting undergraduate research. A significant amount of time will be set aside for audience participation and discussion. Hans-Peter Bischof, Jacob D. Furst, Daniela Raicu, Susan Darling Urban |
SIGCSE | 3 |
| 2009 | A statistical analysis of the effects of CT acquisition parameters on low-level features extracted from CT images of the lungabstractWe propose a solution for automatic classification of lung nodules in an environment with heterogeneous computed tomography (CT) acquisition parameters. Such a classification system needs to take into account the differences in CT acquisition parameters used when obtaining and processing each medical image. Using analysis of variance (ANOVA), our current research proposes to better understand the effects of CT acquisition parameters on predicting various semantic characteristics (such as spiculation, subtlety, and margin) used in the diagnosis interpretation process. All of the parameters were found to affect the low-level image features used in the classification models of these semantic characteristics. When this knowledge is used to normalize those parameters, the final semantic model will become unaffected by the CT acquisition parameters. Joseph S. Wantroba, Daniela Raicu, Jacob D. Furst |
ICIP | 2 |
| 2009 | Enhancing undergraduate education: a REU model for interdisciplinary researchabstractThis paper presents a successful model for undergraduate research where student participants work on interdisciplinary research projects; in our case, at the frontier between computer science and medicine. Students are part of research teams comprised of other undergraduates, graduate students, faculty and medical experts, participate in professional development and training activities within the larger group, and disseminate their results at the host institutions or conferences specific to the interdisciplinary focus. The model outcomes at the end of the first three years (2005-2007) indicate that the interdisciplinary model successfully 1) expanded the student participation in research by recruiting students who might not otherwise have research opportunities, 2) attracted a diversified pool of talented students into science, 3) promoted interdisciplinary undergraduate studies in computer science and medical informatics as well as in future graduate studies; and 4) trained students in all phases of research, including writing and presenting research papers at conferences. Daniela Raicu, Jacob D. Furst |
SIGCSE | 1 |
| 2009 | Comparison and Evaluation of Methods for Liver Segmentation From CT DatasetsabstractThis paper presents a comparison study between 10 automatic and six interactive methods for liver segmentation from contrast-enhanced CT images. It is based on results from the "MICCAI 2007 Grand Challenge" workshop, where 16 teams evaluated their algorithms on a common database. A collection of 20 clinical images with reference segmentations was provided to train and tune algorithms in advance. Participants were also allowed to use additional proprietary training data for that purpose. All teams then had to apply their methods to 10 test datasets and submit the obtained results. Employed algorithms include statistical shape models, atlas registration, level-sets, graph-cuts and rule-based systems. All results were compared to reference segmentations five error measures that highlight different aspects of segmentation accuracy. All measures were combined according to a specific scoring system relating the obtained values to human expert variability. In general, interactive methods reached higher average scores than automatic approaches and featured a better consistency of segmentation quality. However, the best automatic methods (mainly based on statistical shape models with some additional free deformation) could compete well on the majority of test images. The study provides an insight in performance of different segmentation approaches under real-world conditions and highlights achievements and limitations of current image analysis techniques. Tobias Heimann, Bram van Ginneken, Martin Styner, Yulia Arzhaeva, Volker Aurich, Christian Bauer 0001, Andreas Beck 0001, Christoph Becker 0002, Reinhard Beichel, György Bekes, Fernando Bello, Gerd Karl Binnig, Horst Bischof, Alexander Bornik, Peter Cashman, Ying Chi, Andrés Cordova, Benoit M. Dawant, Márta Fidrich, Jacob D. Furst, Daisuke Furukawa, Lars Grenacher, Joachim Hornegger, Dagmar Kainmüller, Richard Kitney, Hidefumi Kobatake, Hans Lamecker, Thomas Lange, Brian Lennon, Rui Li 0012, Senhu Li, Hans-Peter Meinzer, Gábor Németh, Daniela Raicu, Anne-Mareike Rau, Eva M. van Rikxoort, Mikaël Rousson, László Ruskó, Kinda Anna Saddi, Günter Schmidt 0001, Dieter Seghers, Akinobu Shimizu, Pieter Slagmolen, Erich Sorantin, Grzegorz Soza, Ruchaneewan Susomboon, Jonathan M. Waite, Andreas Wimmer, Ivo Wolf |
IEEE Trans. Medical Imaging | 35 |
| 2007 | Oligonucleotide microarray identification of Bacillus anthracis strains using support vector machinesabstractThe capability of a custom microarray to discriminate between closely related DNA samples is demonstrated using a set of Bacillus anthracis strains. The microarray was developed as a universal fingerprint device consisting of 390 genome-independent 9mer probes. The genomes of B. anthracis strains are monomorphic and therefore, typically difficult to distinguish using conventional molecular biology tools or microarray data clustering techniques. Using support vector machines (SVMs) as a supervised learning technique, we show that a low-density fingerprint microarray contains enough information to discriminate between B. anthracis strains with 90% sensitivity using a reference library constructed from six replicate arrays and three replicates for new isolates. Michael V. Doran, Daniela Raicu, Jacob D. Furst, Raffaella Settimi, Matthew Schipma, Darrell P. Chandler |
Bioinform. | 2 |
| 2006 | Automatic Single-Organ Segmentation in Computed Tomography ImagesabstractIn this paper, we propose a hybrid approach for automatic single-organ segmentation in computed tomography (CT) data. The approach consists of three stages: first, a probability image of the organ of interest is obtained by applying a binary classification model obtained using pixel-based texture features; second, an adaptive split-and-merge segmentation algorithm is applied on the organ probability image to remove the noise introduced by the misclassified pixels; and third, the segmented organ's boundaries from the previous stage are iteratively refined using a region growing algorithm. While we applied our approach for liver segmentation in 2-D CT images, a challenging and important task in many medical applications, the proposed approach can be applied for the segmentation of any other organ in CT images. Moreover, the proposed approach can be extended to perform automatic multiple organ segmentation and to build context-sensitive reporting tools for computer-aided diagnosis applications. Ruchaneewan Susomboon, Daniela Raicu, Jacob D. Furst, David S. Channin |
ICDM | 2 |
| 2005 | Texture-Based Image Retrieval for Computerized Tomography DatabasesabstractIn this paper we propose a content-based image retrieval (CBIR) system for retrieval of normal anatomical regions present in computed tomography (CT) studies of the chest and abdomen. We implement and compare eight similarity measures using local and global cooccurrence texture descriptors. The preliminary results are obtained using a CT database consisting of 344 CT images representing the segmented heart and great vessels, liver, renal and splenic parenchyma, and backbone from two different patients. We evaluate the results with respect to the retrieval precision metric for each of the similarity measures when calculated per organ and overall. Winnie Tsang, Andrew Corboy, Ken Lee, Daniela Raicu, Jacob D. Furst |
CBMS | 4 |
| 2005 | A classification approach for anatomical regions segmentationabstractIn this paper, a supervised pixel-based classifier approach for segmenting different anatomical regions in abdominal computed tomography (CT) studies is presented. The approach consists of three steps: texture extraction, classifier creation, and anatomical regions identification. First, a set of co-occurrence texture descriptors is calculated for each pixel from the image data sample; second, a decision tree classifier is built using the texture descriptors and the names of the tissues as class labels. At the conclusion of the classification process, a set of decision rules is generated to be used for classification of new pixels and identification of different anatomical regions by joining adjacent pixels with similar classifications. It is expected that the proposed approach will also help automate different semi-automatic segmentation techniques by providing initial boundary points for deformable models or seed points for split and merge segmentation algorithms. Preliminary results obtained for normal CT studies are presented. Mikhail Kalinin, Daniela Raicu, Jacob D. Furst, David S. Channin |
ICIP (2) | 2 |
| 2003 | eID: a system for exploration of image databases
Daniela Raicu, Ishwar K. Sethi |
Inf. Process. Manag. | 1 |