Lawrence O. Hall

dblp:52/105 · also Larry O. Hall, Lawrence O'Higgins Hall · DBLP profile ↗
← Back
140ranked-venue papers
22as first author
8since 2021 · last 2024
0000-0002-7898-8456ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 87 · 15 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 48 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 44 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 15 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author
YearPublicationVenuePosition
2024 Active Prompting of Vision Language Models for Human-in-the-loop Classification and Explanation of Microscopy Images
abstract
Current AI-based methods for the classification of cellular features in microscopy images require time- and labor-intensive processes for training models. Specific limitations include the need for a large amount of image data and major time commitments from domain experts for accurate ground truthing. We present a solution for overcoming these limitations using a state-of-the-art vision language model. Our approach uses GPT-4, the Vision Language Model (VLM) from OpenAI, for the analysis and classification of Iba-1 immuno-stained microglia cells in tissue sections through the mouse hippocampus. We used GPT-4 to classify a dataset of low-power (20x) images of Iba-1 immuno-stained microglia cells from tissue sections treated with saline or a potent neurotoxin (tri-methyl-tin, TMT). Rather than training with images from each class, the GPT-4 input consists of minimal ground-truth prompts for visual question answering. We introduce a novel human-in-the-loop approach to automate the selection of example image-text pairs as input prompts and generate explanatory text as the basis for separating images into distinct classes. We assess test accuracy and efficiency compared to the baseline results using a convolutional neural net applied to the same dataset. Compared to the baseline, the equivalence in accuracy (91%) and substantial (86%) improvement in throughput efficiency with considerably lower needs for input data or domain experts highlight the effectiveness of our new method for automatic image classification. Unlike traditional methods focused on labeling images with one-to-two-word tags, our pipeline generates understandable ground truth by incorporating explanatory text for each image.
Abhiram Kandiyana, Peter R. Mouton, Lawrence O. Hall, Dmitry B. Goldgof
CBMS3
2024 Anonymized Identity Tracking: Privacy Preserving Facial Encoding
abstract
Background: The need for sharing large-scale datasets in training deep learning models, particularly in healthcare, raises significant data security and privacy concerns. To address these issues, methods such as data encryption or encoding are utilized. These techniques can encrypt the data and make it unreadable to humans, while still retaining its usefulness for training models.Method: In this study, we investigate various image encoding techniques designed to protect privacy by making images unrecognizable while still retaining their usefulness for model training. Our investigation utilized a publicly available facial database and focused on evaluating the trade-offs inherent in image encoding techniques, with a special emphasis on balancing privacy and model accuracy.Conclusion: This study navigate the balance between protecting sensitive data and meeting the data demands necessary for effective model training. It sheds light on the intricate trade-offs among different image encoding techniques and offers insights into finding an optimal balance between privacy protection and model performance.
Manas Sanjay Pakalapati, Dmitry B. Goldgof, Lawrence O. Hall, Ghada Zamzmi
CBMS3
2024 A Review of Nuclei Detection and Segmentation on Microscopy Images Using Deep Learning With Applications to Unbiased Stereology Counting
abstract
The detection and segmentation of stained cells and nuclei are essential prerequisites for subsequent quantitative research for many diseases. Recently, deep learning has shown strong performance in many computer vision problems, including solutions for medical image analysis. Furthermore, accurate stereological quantification of microscopic structures in stained tissue sections plays a critical role in understanding human diseases and developing safe and effective treatments. In this article, we review the most recent deep learning approaches for cell (nuclei) detection and segmentation in cancer and Alzheimer's disease with an emphasis on deep learning approaches combined with unbiased stereology. Major challenges include accurate and reproducible cell detection and segmentation of microscopic images from stained sections. Finally, we discuss potential improvements and future trends in deep learning applied to cell detection and segmentation.
Saeed S. Alahmari, Dmitry B. Goldgof, Lawrence O. Hall, Peter R. Mouton
IEEE Trans. Neural Networks Learn. Syst.3
2023 Unsupervised Prostate Cancer Histopathology Image Segmentation via Meta-Learning
abstract
We propose a novel unsupervised meta-learning based segmentation algorithm for histopathology images. The proposed algorithm does not require any kind of patch-level annotations and relies solely on image labels, corresponding to any classification task, and direct feedback from a classifier. Furthermore, instead of simply segmenting histopathology images into different types of tissue, our algorithm determines the relative importance of each tissue region. After thresholding, the produced segmentations can also be used as regions of interest for various machine learning based diagnosis systems. We have tested our approach on Prostate cANcer graDe Assessment (PANDA) dataset and obtained 0.79 AUC, when testing the segmentation performance at patch-level, and 0.432 Dice coefficient, when testing precise segmentation, which is comparable to 0.446, described in related work which performed a supervised segmentation with U-Net. Note that no pixel level annotations were used.
Nikolai Fetisov, Lawrence O. Hall, Dmitry B. Goldgof, Matthew B. Schabath
CBMS2
2023 MIMO YOLO - A Multiple Input Multiple Output Model for Automatic Cell Counting
abstract
Across basic research studies, cell counting requires significant human time and expertise. Trained experts use thin focal plane scanning to count (click) cells in stained biological tissue. This computer-assisted process (optical disector) requires a well-trained human to select a unique best z-plane of focus for counting cells of interest. Though accurate, this approach typically requires an hour per case and is prone to inter-and intra-rater errors. Our group has previously proposed deep learning (DL)-based methods to automate these counts using cell segmentation at high magnification. Here we propose a novel You Only Look Once (YOLO) model that performs cell detection on multi-channel z-plane images (disector stack). This automated Multiple Input Multiple Output (MIMO) version of the optical disector method uses an entire z-stack of microscopy images as its input, and outputs cell detections (counts) with a bounding box of each cell and class corresponding to the z-plane where the cell appears in best focus. Compared to the previous segmentation methods, the proposed method does not require time-and labor-intensive ground truth segmentation masks for training, while producing comparable accuracy to current segmentation-based automatic counts. The MIMO-YOLO method was evaluated on systematic-random samples of NeuN-stained tissue sections through the neocortex of mouse brains (n=7). Using a cross validation scheme, this method showed the ability to correctly count total neuron numbers with accuracy close to human experts and with 100% repeatability (Test-Retest).
Hunter Morera, Palak Dave, Saeed S. Alahmari, Yaroslav Kolinko, Lawrence O. Hall, Dmitry B. Goldgof, Peter R. Mouton
CBMS5
2023 Using SMOTE-based Data Augmentation for Social Media Time Series Prediction
abstract
In the context of predicting activity on a social network, data for any individual will be limited. Also, low levels of activity for a topic of interest may make it difficult to build a strong predictive model. This work examines how augmentation by oversampling all activity data used to build a predictive model of user activity on Twitter can be used to improve the fidelity of predictions. Our features are counts of activity by deidentified users at the granularity of hours. It is shown that for some topics oversampling the data by creating new synthetic examples provides an effective way to increase the accuracy of predictions of future activity.
Frederick Mubang, Lawrence O. Hall
SMC2
2023 VAM: An End-to-End Simulator for Time Series Regression and Temporal Link Prediction in Social Media Networks
abstract
We present a machine-learning-driven end-to-end simulator, called theVolume-Audience-Match(VAM) simulator. VAM’s purpose is to simulate future phenomena related to various topics of discussion in social media networks. We focus our attention on the social media platform, Twitter, due to its abundant use in today’s world. VAM was applied to do time series forecasting to predict the future: 1) number of total activities; 2) number of active old users; and 3) number of newly active users over the span of 24 hours from the start time of prediction. VAM then used these macroscopic volume predictions (VPs) to perform user link predictions. A user–user edge was assigned to each of the activities in the 24 future time steps. We report that VAM outperformed multiple baseline models in the time series task, which were the auto-regressive integrated moving average (ARIMA), auto regressive moving average (ARMA), auto regressive (AR), moving average (MA), Persistence Baseline, and state-of-the-art tNodeEmbed models. Furthermore, we show that VAM outperformed the Persistence Baseline and tNodeEmbed models used for the user-assignment tasks. Finally, it is also shown that using Reddit activity data improves prediction accuracy.
Frederick Mubang, Lawrence O. Hall
IEEE Trans. Comput. Soc. Syst.2
2022 Simulating New and Old Twitter User Activity with XGBoost and Probabilistic Hybrid Models
abstract
The Volume Audience Match Simulator is an end-to-end approach for predicting user-to-user interactions on a given social media platform. It is comprised of 2 components: firstly, an XGBoost-driven volume prediction module that predicts the number of: (1) total activities, (2) active old users, and (3) newly active users over the span of 24 hours from the start time of prediction. Secondly, VAM contains a User-Assignment Module that takes as input the volume predictions and predicts the user-to-user interactions of the old and new users.In previous work, VAM has been used to predict Twitter discussions related to political crises. In this work, VAM was used to predict future activity on Twitter related to international economic affairs. We include more experiments and analyses than previous work performed with VAM. In this work, VAM is used to predict all types of retweets, including quotes and replies, unlike previous work, which only focused on regular retweets. Furthermore, we show that YouTube features, in addition to Reddit features can improve prediction performance. We examine the importance of the time series features used in VAM’s Volume Prediction module. Lastly, we show that VAM’s performance is significantly more accurate than other approaches when predicting highly-skewed, lowly-skewed, highly-sparse, and lowly-sparse time series.
Frederick Mubang, Lawrence O. Hall
ICMLA2
2020 Fuzzy Set Similarity for Feature Selection in Classification
abstract
A problem for machine learning research occurs when many possible features exist but the training data examples are very few. For example, microarray data typically have a much larger number of features, the genes, as compared to the number of training data examples, the patients. One approach is to first determine the best features for prediction and then to group features based on a measure of their relatedness. The concordance correlation coefficient has been used to place somewhat correlated features into disjoint groups of similar features. Multiple base classifiers are created by randomly picking one feature from each of the feature groups and then the collection of base classifiers is used in an ensemble classifier. Each classifier in the ensemble provides a vote. The majority vote is used to produce the final class prediction. This paper investigates grouping features using fuzzy set similarity measures as well as the concordance correlation coefficient as a relatedness measure. The performance of these different measures is compared in terms of accuracy, sensitivity, specificity, and F-measure using the ensemble classifiers created with the different relatedness measures. Four microarray gene expression data sets are used in the experiments to determine the usefulness of fuzzy set similarity measures and how they compare with the concordance correlation coefficient. Using the concordance correlation coefficient to guide clustering is not superior to fuzzy set similarity measures. Depending on the particular data set and performance measure being used, different fuzzy set similarity measures perform better than or just as well as the concordance correlation coefficient.
Valerie V. Cross, Michael Zmuda, Rahul Paul, Lawrence O. Hall
FUZZ-IEEE4
2020 Simulating Temporal User Activity on Social Networks with Sequence to Sequence Neural Models
abstract
The prediction of long-term activities of groups of users and clusters of activities around a subject in social networks is a very challenging task. In this paper, we propose a novel temporal neural network framework that tracks user engagement and activity associated with particular subjects (e.g. CVE IDs) across online platforms. The framework is able to simulate which user will do what activity and at what time. Furthermore, this framework captures groups of users reacting to an event. It also captures responses to an event on a platform and the influence of the event on activity on other platforms over time. The proposed framework aims to predict future user activity related to specific subjects across platforms. The framework also illustrates the importance influence of activities that occur on other platforms when predicting user activity for particular events on a different platform. The learned model can do simulations in a timely manner. We evaluated our user group activity prediction method on the CVE (Common Vulnerabilities and Exposures) related user groups (software vulnerability) using 3 public online social network datasets: Github, Reddit, and Twitter. Groups of users who work on a particular CVE ID are identified. Each user group has information on all users' activities related to a CVE ID. The 3 datasets from Github, Reddit, and Twitter contain more than 490,000 cross platform activities related to over 20,000 user groups (CVE IDs) from more than 50,000 users. Compared to the proposed baseline, our simulation method is better in both predictions of total activity volume over time and activity associated with an individual CVE ID.
Renhao Liu, Frederick Mubang, Lawrence O. Hall
SMC3
2019 Neuroimaging Based Survival Time Prediction of GBM Patients Using CNNs from Small Data
abstract
Here we investigate the application of convolutional neural networks (CNNs) to predict the survival time of patients with Glioblastoma Multiforme (GBM) brain tumor. Our dataset consists of T1-weighted high-resolution MRI images of just 68 GBM patients. We compare two analytic methods for predicting survival time. The first consists of training a small convolutional neural network (CNN) and the second uses extracted deep features from a pre-trained CNN. Our method is completely automated, except for tumor region segmentation. In addition, we utilize a snapshot ensemble approach to boost test accuracy when dealing with limited availability of medical images for CNN training purposes. Our approach achieves an accuracy of 72.06% using a trained small network and 66.18% using a pre-trained deep CNN. Our results compare favorably with the accuracy of 54.41% using histogram of oriented gradients (HOG) features and a non-neural network classifier.
Kaoutar Ben Ahmed, Lawrence O. Hall, Renhao Liu, Robert A. Gatenby, Dmitry B. Goldgof
SMC2
2019 Automatic Cell Counting using Active Deep Learning and Unbiased Stereology
abstract
Training deep learning models for unbiased stereology requires a large data set with associated ground truth. However manual ground truth annotation is tedious, time-consuming, and expert dependent. We propose an active deep learning method for automatic stereology counts using a snapshot ensemble approach. The method provides a confidence score for each mask in an unlabeled pool that reduces user verification to only images with high information content for training the deep learning model. The proposed method reduces the error rate to less than 1% for unbiased stereology cell counts on immunostained brain cells compared to manual stereology and requires ~25% less expert verification time compared to a previously proposed iterative deep learning approach.
Saeed S. Alahmari, Dmitry B. Goldgof, Lawrence O. Hall, Peter R. Mouton
SMC3
2019 Predicting Longitudinal User Activity at Fine Time Granularity in Online Collaborative Platforms
abstract
This paper introduces a decomposition approach to address the problem of predicting different user activities at hour granularity over a long period of time. Our approach involves two steps. First, we used a temporal neural network ensemble to predict the number of each type of activity that occurred in a day. Second, we used a set of neural networks to assign the events to a user-repository pair in a particular hour. We focused this work on a subset of the public GitHub dataset that records the activities of over 2 million users on over 400,000 software repositories. Our experiments show we were able to predict hourly user-repo activity with reasonably low error. Our simulations are accurate for 1-3 weeks (168-504 hours) after inception, with accuracy gradually falling off. It was shown that activity on Twitter and Reddit increases the accuracy of activity prediction on GitHub for most events.
Renhao Liu, Frederick Mubang, Lawrence O. Hall, Sameera Horawalavithana, Adriana Iamnitchi, John Skvoretz
SMC3
2019 Mentions of Security Vulnerabilities on Reddit, Twitter and GitHub
abstract
Activity on social media is seen as a relevant sensor for different aspects of the society. In a heavily digitized society, security vulnerabilities pose a significant threat that is publicly discussed on social media. This study presents a comparison of user-generated content related to security vulnerabilities on three digital platforms: two social media conversation channels (Reddit and Twitter) and a collaborative software development platform (GitHub). Our data analysis shows that while more security vulnerabilities are discussed on Twitter, relevant conversations go viral earlier on Reddit. We show that the two social media platforms can be used to accurately predict activity on GitHub.
Sameera Horawalavithana, Abhishek Bhattacharjee, Renhao Liu, Nazim Choudhury, Lawrence O. Hall, Adriana Iamnitchi
WI5
2018 Iterative Deep Learning Based Unbiased Stereology with Human-in-the-Loop
abstract
Lack of enough labeled data is a major problem in building machine learning based models when the manual annotation (labeling) is error-prone, expensive, tedious, and time-consuming. In this paper, we introduce an iterative deep learning based method to improve segmentation and counting of cells based on unbiased stereology applied to regions of interest of extended depth of field (EDF) images. This method uses an existing machine learning algorithm called the adaptive segmentation algorithm (ASA) to generate masks (verified by a user) for EDF images to train deep learning models. Then an iterative deep learning approach is used to feed newly predicted and accepted deep learning masks/images (verified by a user) to the training set of the deep learning model. The error rate in unbiased stereology count of cells on an unseen test set reduced from about 3 % to less than 1 % after 5 iterations of the iterative deep learning based unbiased stereology process.
Saeed S. Alahmari, Dmitry B. Goldgof, Lawrence O. Hall, Palak Dave, Hady Ahmady Phoulady, Peter R. Mouton
ICMLA3
2018 Predicting Nodule Malignancy using a CNN Ensemble Approach
abstract
Lung cancer is the leading cause of cancer-related deaths globally, which makes early detection and diagnosis a high priority. Computed tomography (CT) is the method of choice for early detection and diagnosis of lung cancer. Radiomics features extracted from CT-detected lung nodules provide a good platform for early detection, diagnosis, and prognosis. In particular when using low dose CT for lung cancer screening, effective use of radiomics can yield a precise non-invasive approach to nodule tracking. Lately, with the advancement of deep learning, convolutional neural networks (CNN) are also being used to analyze lung nodules. In this study, our own trained CNNs, a pre-trained CNN and radiomics features were used for predictive analysis. Using subsets of participants from the National Lung Screening Trial, we investigated if the prediction of nodule malignancy could be further enhanced by an ensemble of classifiers using different feature sets and learning approaches. We extracted probability predictions from our different models on an unseen test set and combined them to generate better predictions. Ensembles were able to yield increased accuracy and area under the receiver operating characteristic curve (AUC). The best-known AUC of 0.96 and accuracy of 89.45% were obtained, which are significant improvements over the previous best AUC of 0.87 and accuracy of 76.79%.
Rahul Paul, Lawrence O. Hall, Dmitry B. Goldgof, Matthew B. Schabath, Robert J. Gillies
IJCNN2
2018 Representation of Deep Features using Radiologist defined Semantic Features
abstract
Semantic features are common radiological traits used to characterize a lesion by a trained radiologist. These features have been recently formulated, quantified on a point scale in the context of lung nodules by our group. Certain radiological semantic traits have been shown to extremely predictive of malignancy [26]. Semantic traits observed by a radiologist at examination describe the nodules and the morphology of the lung nodule shape, size, border, attachment to vessel or pleural wall, location and texture etc. Deep features are numeric descriptors often obtained from a convolutional neural network (CNN) which are widely used for classification and recognition. Deep features may contain information about texture and shape, primarily. Lately, with the advancement of deep learning, convolutional neural networks (CNN) are also being used to analyze lung nodules. In this study, we relate deep features to semantic features by looking for similarity in ability to classify. Deep features were obtained using a transfer learning approach from both an ImageNet pre-trained CNN and our trained CNN architecture. We found that some of the semantic features can be represented by one or more deep features. In this process, we can infer that some deep feature(s) have similar discriminatory ability as semantic features.
Rahul Paul, Lawrence O. Hall, Dmitry B. Goldgof, Yoganand Balagurunathan, Matthew B. Schabath, Robert J. Gillies
IJCNN4
2017 Synthetic minority image over-sampling technique: How to improve AUC for glioblastoma patient survival prediction
abstract
Real-world datasets are often imbalanced, with an important class having many fewer examples than other classes. In medical data, normal examples typically greatly outnumber disease examples. A classifier learned from imbalanced data, will tend to be very good at the predicting examples in the larger (normal) class, yet the smaller (disease) class is typically of more interest. Imbalance is dealt with at the feature vector level (create synthetic feature vectors or discard some examples from the larger class) or by assigning differential costs to errors. Here, we introduce a novel method for over-sampling minority class examples at the image level, rather than the feature vector level. Our method was applied to the problem of Glioblastoma patient survival group prediction. Synthetic minority class examples were created by adding Gaussian noise to original medical images from the minority class. Uniform local binary patterns (LBP) histogram features were then extracted from the original and synthetic image examples with a random forests classifier. Experimental results show the new method (Image SMOTE) increased minority class predictive accuracy and also the AUC (area under the receiver operating characteristic curve), compared to using the imbalanced dataset directly or to creating synthetic feature vectors.
Renhao Liu, Lawrence O. Hall, Kevin W. Bowyer, Dmitry B. Goldgof, Robert A. Gatenby, Kaoutar Ben Ahmed
SMC2
2017 Finding label noise examples in large scale datasets
abstract
Mislabeled examples are difficult to avoid while building large scale datasets. In this paper we discuss an efficient approach for finding those mislabeled examples. Our approach involves selecting a small number of potentially mislabeled examples for review by an expert. We demonstrate the utility of our method by finding some mislabeled examples in one large scale dataset (ImageNet). We found 92 errors by automatically selecting 3607 examples to review out of 22951 images from 18 classes. This requires reviewing 9 times fewer examples than the random sampling method to find an equivalent number of mislabels.
Ekambaram Rajmadhan, Dmitry B. Goldgof, Lawrence O. Hall
SMC3
2017 Active Multitask Learning With Trace Norm Regularization Based on Excess Risk
abstract
This paper addresses the problem of active learning on multiple tasks, where labeled data are expensive to obtain for each individual task but the learning problems share some commonalities across multiple related tasks. To leverage the benefits of jointly learning from multiple related tasks and making active queries, we propose a novel active multitask learning approach based on trace norm regularized least squares. The basic idea is to induce an optimal classifier which has the lowest risk and at the same time which is closest to the true hypothesis. Toward this aim, we devise a new active selection criterion that takes into account not only the risk but also the excess risk, which measures the distance to the true hypothesis. Based on this criterion, our proposed algorithm actively selects the instance to query for its label based on the combination of the two risks. Experiments on both synthetic and real-world datasets show that our proposed algorithm provides superior performance as compared to other state-of-the-art active learning methods.
Jie Yin 0001, Lawrence O. Hall, Dacheng Tao
IEEE Trans. Cybern.3
2016 Automatic quantification and classification of cervical cancer via Adaptive Nucleus Shape Modeling
abstract
Decisions about cervical cancer diagnosis and classification currently require microscopic examination of cervical tissue by an expert pathologist. In the present study, which focused on full automation of this approach, we solely use nucleus-level features to classify tissues as normal or cancer. We propose Adaptive Nucleus Shape Modeling (ANSM) algorithm for nucleus-level analysis which consists of two steps to capture the nucleus-level information: adaptive multilevel thresholding segmentation; and shape approximation by ellipse fitting. After applying the proposed algorithm, the features are extracted for tissue classification. Experiments show that ANSM can achieve an accuracy of 93.33% with a false negative rate of zero in classifying cancer and healthy cervical tissues using nucleus texture features. This provides evidence that nucleus-level analysis is valuable in cervical histology image analysis.
Hady Ahmady Phoulady, Mu Zhou, Dmitry B. Goldgof, Lawrence O. Hall, Peter R. Mouton
ICIP4
2016 Spectral sparsification in spectral clustering
abstract
Graph spectral clustering algorithms have been shown to be effective in finding clusters and generally outperform traditional clustering algorithms, such as k-means. However, they have scalibility issues in both memory usage and computational time. To overcome these limitations, the common approaches sparsify the similarity matrix by zeroing out some of its elements. They generally consider local neighborhood relationships between the data instances such as the k-nearest neighbor method. Although, they eventually work with the Laplacian matrix, there is no evidence about preservation of its spectral properties with approximation guarantees. As a result, in this paper, we adopt the idea of spectral sparsification to sparsify the Laplacian matrix. A spectral sparsification algorithm takes a dense graph G with n vertices and m edges (that is usually O(n2)), and returns a new graph H with the same set of vertices and many fewer edges, on the order of O(n log n), that approximately preserves the spectral properties of the input graph. We study the effect of the spectral sparsification method based on sampling by effective resistance on the spectral clustering outputs. Through experiments, we show that the clustering results obtained from sparsified graphs are very similar to the results of the original non-sparsified graphs.
Alireza Chakeri, Hamidreza Farhidzadeh, Lawrence O. Hall
ICPR3
2016 Exploring deep features from brain tumor magnetic resonance images via transfer learning
abstract
Finding appropriate feature representations from radiological images is a vital task for prediction and diagnosis. Deep convolutional neural networks have recently achieved state-of-the-art performance in classification problems from several different domains. Research has also shown the feasibility of using a pre-trained deep neural network as a feature extractor when only a small dataset is available. This paper proposes a novel image feature extraction method for predicting survival time from brain tumor magnetic resonance images using pretrained deep neural networks. Since all tumors are different sizes, we also explore different image resizing methods in the paper. We demonstrate that deep features can result in better survival time prediction with the highest accuracy of 95.45% versus conventional feature extraction methods from magnetic resonance images of the brain.
Renhao Liu, Lawrence O. Hall, Dmitry B. Goldgof, Mu Zhou, Robert A. Gatenby, Kaoutar Ben Ahmed
IJCNN2
2016 Improving malignancy prediction through feature selection informed by nodule size ranges in NLST
abstract
Computed tomography (CT) is widely used during diagnosis and treatment of Non-Small Cell Lung Cancer (NSCLC). Current computer-aided diagnosis (CAD) models, designed for the classification of malignant and benign nodules, use image features, selected by feature selectors, for making a decision. In this paper, we investigate automated selection of different image features informed by different nodule size ranges to increase the overall accuracy of the classification. The NLST dataset is one of the largest available datasets on CT screening for NSCLC. We used 261 cases as a training dataset and 237 cases as a test dataset. The nodule size, which may indicate biological variability, can vary substantially. For example, in the training set, there are nodules with a diameter of a couple millimeters up to a couple dozen millimeters. The premise is that benign and malignant nodules have different radiomic quantitative descriptors related to size. After splitting training and testing datasets into three subsets based on the longest nodule diameter (LD) parameter accuracy was improved from 74.68% to 81.01% and the AUC improved from 0.69 to 0.79. We show that if AUC is the main factor in choosing parameters then accuracy improved from 72.57% to 77.5% and AUC improved from 0.78 to 0.82. Additionally, we show the impact of an oversampling technique for the minority cancer class. In some particular cases from 0.82 to 0.87.
Dmitry Cherezov, Samuel H. Hawkins, Dmitry B. Goldgof, Lawrence O. Hall, Yoganand Balagurunathan, Robert J. Gillies, Matthew B. Schabath
SMC4
2016 A quantitative histogram-based approach to predict treatment outcome for Soft Tissue Sarcomas using pre- and post-treatment MRIs
abstract
The goal of this paper is to show the use of data mining techniques to predict the Soft Tissue Sarcoma (STS) tumor progression. STS are cancers which occur in different parts of the body such as fat, muscle and nerves. The lack of effective treatments and the difficulty in predicting treatment response make them challenging for physicians, and has likely slowed the evolution of new therapeutic agents. To design a prediction model, we propose a novel quantitative histogram-based method to analyze the difference in histograms obtained from pre and post-treatment multi-modality magnetic resonance images. Here, we used Radiomics techniques as a non-invasive method for outcome prediction. This study could help physicians identify distinctive patterns within each tumor to find more patient-specific treatments. We demonstrated the new approach on two practical tasks: tumor recurrence prediction (metastasis) and rate of necrosis prediction. Our learned model shows 87.79% prediction accuracy for metastasis with a 0.73 AUC and 82.22% prediction accuracy for necrosis with a 0.65 AUC.
Hamidreza Farhidzadeh, Dmitry B. Goldgof, Lawrence O. Hall, Jacob G. Scott, Robert A. Gatenby, Robert J. Gillies, Meera Raghavan
SMC3
2016 Combining deep neural network and traditional image features to improve survival prediction accuracy for lung cancer patients from diagnostic CT
abstract
Lung cancer is caused by abnormal and uncontrolled growth of cells in the lungs and the mortality rate of lung cancer is the highest among all types of cancer. It can be identified and treated with the help of computed tomography (CT) images. For an automated classifier, identifying good features from an image is a key concern. Deep feature extraction using pre-trained convolutional neural networks has been successful for some image domains recently. In our study, we apply a pre-trained convolutional neural network (CNN) to extract deep features from lung cancer CT images and then train classifiers to predict short and long term survivors. The best accuracy of 77.5% was with a cropping approach using a decision tree classifier in a leave one out cross validation with ten features chosen using symmetric uncertainty feature ranking. We mixed extracted deep neural network features along with quantitative (traditional image) features and obtained the best accuracy of 82.5% with a nearest neighbor classifier in a leave one out cross validation using the symmetric uncertainty feature ranking algorithm.
Rahul Paul, Samuel H. Hawkins, Lawrence O. Hall, Dmitry B. Goldgof, Robert J. Gillies
SMC3
2016 Active cleaning of label noise
Ekambaram Rajmadhan, Sergiy Fefilatyev, Matthew Shreve, Kurt Kramer, Lawrence O. Hall, Dmitry B. Goldgof, Rangachar Kasturi
Pattern Recognit.5
2015 Correlation Based Random Subspace Ensembles for Predicting Number of Axillary Lymph Node Metastases in Breast DCE-MRI Tumors
abstract
An important problem in quantitative medical image analysis is a large number of features (often highly correlated) to instance ratio. To handle this, we developed a feature selector and an ensemble classifier based on a modified version of random subspace method. We propose using a fusion of feature selection concepts: ranking based, correlation based and random subspaces, to develop a concordance correlation coefficient based random subspace method (CCC RSM) feature selector. It forms random feature subsets with weakly correlated yet relevant features while the ensemble classification is achieved by training the base classifier with these feature subsets. Axillary lymph node (ALNs) metastases is one of the most important prognostic factors in breast cancer. We applied CCC RSM for four binary class classifications based on the number of metastatic ALNs: (i) = 1 vs 0 (ii) 1-3 vs 0 (iii) 1-3 vs = 4, and (iv) 4 vs 0. We extracted textural kinetics from habitats of fifty eight dynamic contrast enhanced magnetic resonance imaging breast tumors. We used three classifiers to compare the accuracies achieved by CCC RSM with random subspaces (RS), wrappers and correlation based feature selector (CFS). For each binary classification we achieved the best accuracy (= 78%) using CCC RSM.
Baishali Chaudhury, Dmitry B. Goldgof, Lawrence O. Hall, Robert A. Gatenby, Robert J. Gillies, Jennifer S. Drukteinis
SMC3
2015 Texture Feature Analysis to Predict Metastatic and Necrotic Soft Tissue Sarcomas
abstract
Soft Tissue Sarcomas (STS) are malignant tumors which emanate from soft tissues of the body. They are challenging for physicians because of the infrequency of their occurrence and non-predictable outcomes. In this paper, we propose a novel framework to classify STS which focuses on radio logically defined sub-regions, so-called 'habitats'. The distinctive habitats are regions where tumor evolution may be observed. We assess T1 post- and pre-contrast gadolinium and T2 non-contrast Magnetic Resonance Images (MRIs) of 36 patients prior to treatment. This paper considers spatially distinct habitats, which may be helpful in clinical treatment, especially chemotherapy and radiation. Our approach contains three main steps: (1) intra-tumor segmentation into habitats based on pixel intensity, (2) texture analysis within each distinctive habitat to capture heterogeneity, and (3) prediction of metastatic and necrotic tumor. The experimental results show the individual cases were correctly classified as metastatic or non-metastatic disease with 86.11% accuracy based on 5 features and for necrosis =90% or necrosis <; 90% with 81.81% accuracy based on 4 features by using several meta-classifiers.
Hamidreza Farhidzadeh, Dmitry B. Goldgof, Lawrence O. Hall, Robert A. Gatenby, Robert J. Gillies, Meera Raghavan
SMC3
2015 A Robust Approach for Automated Lung Segmentation in Thoracic CT
abstract
Lung segmentation in thoracic computed tomography (CT) scans is an important preprocessing step for computer-aided diagnosis (CAD) of lung diseases. This paper focuses on the segmentation of the lung field in thoracic CT images. Traditional lung segmentation is based on Gray level thresholding techniques, which often requires setting a threshold and is sensitive to image contrasts. In this paper, we present a fully automated method for robust and accurate lung segmentation, which includes a enhanced thresholding algorithm and a refinement scheme based on a texture-aware active contour model. In our thresholding algorithm, a histogram based image stretch technique is performed in advance to uniformly increase contrasts between areas with low Hounsfield unit (HU) values and areas with high HU in all CT images. This stretch step enables the following threshold-free segmentation, which is the Otsu algorithm with contour analysis. However, as a threshold based segmentation, it has common issues such as holes, noises and inaccurate segmentation boundaries that will cause problems in future CAD for lung disease detection. To solve these problems, a refinement technique is proposed that captures vessel structures and lung boundaries and then smooths variations via texture-aware active contour model. Experiments on 2,342 diagnosis CT images demonstrate the effectiveness of the proposed method. Performance comparison with existing methods shows the advantages of our method.
Hailing Zhou, Dmitry B. Goldgof, Samuel H. Hawkins, Lei Wei 0002, Douglas C. Creighton, Robert J. Gillies, Lawrence O. Hall, Saeid Nahavandi
SMC8
2014 Relational data partitioning using evolutionary game theory
abstract
This paper presents a new approach for relational data partitioning using the notion of dominant sets. A dominant set is a subset of data points satisfying the constraints of internal homogeneity and external in-homogeneity, i.e. a cluster. However, since any subset of a dominant set cannot be a dominant set itself, dominant sets tend to be compact sets. Hence, in this paper, we present a novel approach to enumerate well distributed clusters where the number of clusters need not be known. When the number of clusters is known, in order to search the solution space appropriately, after finding each dominant set, data points are partitioned into two disjoint subsets of data points using spectral graph image segmentation methods to enumerate the other well distributed dominant sets. For the latter case, we introduce a new hierarchical approach for relational data partitioning using a new class of evolutionary game theory dynamics called InImDynamics which is very fast and linear, in computational time, with the number of data points. In this regard, at each level of the proposed hierarchy, Dunn's index is used to find the appropriate number of clusters. Then the objects are partitioned based on the projected number of clusters using game theoretic relations. The same method is applied to each partition to extract its underlying structure. Although the resulting clusters exist in their equivalent partitions, they may not be clusters of the entire data. Hence, they are checked for being an actual cluster and if they are not, they are extended to an existing cluster of the data. The approach can also be used to assign unseen data to existing clusters, as well.
Lawrence O. Hall, Alireza Chakeri
CIDM1
2014 Dominant Sets as a Framework for Cluster Ensembles: An Evolutionary Game Theory Approach
abstract
Ensemble clustering aggregates partitions obtained from several individual clustering algorithms. This can improve the accuracy of results from individual methods and provide robustness against variability in the methods applied. Theorems show one can find dominant sets (clusters) very efficiently by using an evolutionary game theoretic approach. Experiments on an MRI data set consisting of about 4 million data are detailed. The distributed dominant set framework generates partitions of quality slightly better than clustering all the data using fuzzy C means.
Alireza Chakeri, Lawrence O. Hall
ICPR2
2014 Exploring Brain Tumor Heterogeneity for Survival Time Prediction
abstract
Brain tumor heterogeneity is well recognized in clinical MRI imaging and it is a challenging problem to quantitatively explore the underlying variations. It is known that brain tumors in different patients can have remarkably diverse visual appearances. In this paper, we propose a novel concept to categorize brain tumors with emphasis on spatial "habitats": a tumor can be quantified into distinctive sub-regions where the potential dynamics of tumor evolution may be evident. Our work is aimed at discovering spatially distinctive habitats within the tumor region, which may be useful in clinical practice for image-guided therapy. In particular, the heterogeneity can be well captured by two main steps: (a) intra-tumor segmentation, (b) spatial mapping scheme from a multi-modality MRI imaging dataset (Tl-weighted, FLAIR and T2-weighted MRI slices). A tumor region is initially segmented into high and low signal groups and then a joint mapping scheme is used to consider the correlation between different input modalities. In addition, focusing on signal contrast, we propose a set of quantitative features to measure differences between sub-regions. We further examined the application of survival time prediction for patients with malignant Glioblastoma multiforme (GBM). Experimental results showed that these features enabled classifiers to predict survival groups.
Mu Zhou, Lawrence O. Hall, Dmitry B. Goldgof
ICPR2
2014 Using features from tumor subregions of breast DCE-MRI for estrogen receptor status prediction
abstract
In breast cancer, tumor heterogeneity is a reflection of differing tumor subtypes, which may display markedly different genotypes and clinical phenotypes. Although pathological and qualitative (based on contrast enhancement patterns) studies suggest the presence of clinical and molecular predictive tumor subregions, this has not been fully investigated. Our goal is to develop a novel algorithm to utilize the potential information available in different tumor subregions (periphery and core) by extracting textural kinetic features, for the purpose of estrogen receptor (ER) classification. We show that features from different tumor subregions, at appropriate scales and quantization levels, can be used to better classify ER subtypes than features averaged from the whole tumor. We analyzed representative two dimensional (2D) slices from twenty breast tumors with volumetric dynamic contrast enhanced magnetic resonance imaging (DCE-MRI) available by extracting multi-parametric textural kinetic features from the periphery, core and whole tumor. The utility of the features from different subregions are evaluated using six meta-classifiers (feature selector and classifier pairs), formed from two feature selectors and three classifiers. Classification accuracy approached 94%.
Baishali Chaudhury, Mu Zhou, Dmitry B. Goldgof, Lawrence O. Hall, Robert A. Gatenby, Robert J. Gillies, Jennifer S. Drukteinis
SMC4
2014 Experiments with large ensembles for segmentation and classification of cervical cancer biopsy images
abstract
To classify cervical cells as normal or cancer, the histological image must be segmented. After segmentation mean nuclear volume can be used to distinguish between normal and cancer cells. Due to the rapid reproduction of cancer cells, they have higher mean nuclear volume than typical normal cells. We propose a large ensemble of segmentations which separate normal and cancer cases based on the single feature of mean nuclear volume. Four basic segmentors with different parameters generate the segmentations. The mean nuclear volume is extracted from the segmentations. The dataset used for this paper contained multiple images from 30 normal and 32 cancer patients. Hematoxylin and eosin (H&E) was used to stain archival tissue sections from the normal cervix and cervical cancers. Results show it is possible to predict class with greater than 84% accuracy.
Hady Ahmady Phoulady, Baishali Chaudhury, Dmitry B. Goldgof, Lawrence O. Hall, Peter R. Mouton, Ardeshir Hakam, Erin M. Siegel
SMC4
2014 Accelerating Fuzzy-C Means Using an Estimated Subsample Size
abstract
Many algorithms designed to accelerate the Fuzzy c-Means (FCM) clustering algorithm randomly sample the data. Typically, no statistical method is used to estimate the subsample size, despite the impact subsample sizes have on speed and quality. This paper introduces two new accelerated algorithms, GOFCM and MSERFCM, that use a statistical method to estimate the subsample size. GOFCM, a variant of SPFCM, also leverages progressive sampling. MSERFCM, a variant of rseFCM, gains a speedup from improved initialization. A general, novel stopping criterion for accelerated clustering is introduced. The new algorithms are compared to FCM and four accelerated variants of FCM. GOFCM's speedup was 4-47 times that of FCM and faster than SPFCM on each of the six datasets used in experiments. For five of the datasets, partitions were within 1% of those of FCM. MSERFCM's speedup was 5-26 times that of FCM and produced partitions within 3% of those of FCM on all datasets. A unique dataset, consisting of plankton images, exposed the strengths and weaknesses of many of the algorithms tested. It is shown that the new stopping criterion is effective in speeding up algorithms such as SPFCM and the final partitions are very close to those of FCM.
Jonathon K. Parker, Lawrence O. Hall
IEEE Trans. Fuzzy Syst.2
2013 Dempster-Shafer theory of evidence in Single Pass Fuzzy C Means
abstract
Clustering large data sets has become very importantas the amount of available unlabeled data increases. Single Pass Fuzzy C Means (SPFCM) is useful when memory is too limited to load the whole data set. The main idea is to divide dataset into several chunks and to apply FCM to each chunk. SPFCM uses the weighted cluster centers of the previous chunk in the next chunks. Although when the number of chunks is increased, the algorithm shows sensitivity to the order the data processed. Hence, we improved SPFCM by recognizing boundary and noisy data in each chunk and using it to influence clustering in the next chunks. In this regard, the proposed approach transfers the boundary and noisy data as well as the weighted cluster centers to the next chunks. We show that our proposed approach is significantly less sensitive to the order in which the data is loaded in each chunk.
Alireza Chakeri, Iman Nekooeimehr, Lawrence O. Hall
FUZZ-IEEE3
2013 Effect of Texture Features in Computer Aided Diagnosis of Pulmonary Nodules in Low-Dose Computed Tomography
abstract
Low-dose helical computed tomography (LDCT) has facilitated the early detection of lung cancer through pulmonary screening of patients. There have been a few attempts to develop a computer-aided diagnosis system for classifying pulmonary nodules using size and shape, with little attention to texture features. In this work, texture and shape features were extracted from pulmonary nodules selected from the LIDC data set. Several classifiers including Decision Trees, Nearest Neighbor, and Support Vector Machines (SVM) were used for classifying malignant and benign pulmonary nodules. An accuracy of 90.91% was achieved using a 5-nearest-neighbors algorithm and a data set containing texture features only. Laws and Wavelet features received the highest rank when using feature selection implying a larger contribution in the classification process. Considering the improvement in classification accuracy, the use of texture features appears to be a promising direction in computer-aided diagnosis of pulmonary nodules in LDCT.
Henry Krewer, Benjamin Geiger, Lawrence O. Hall, Dmitry B. Goldgof, Yuhua Gu, Melvyn S. Tockman, Robert J. Gillies
SMC3
2013 A Texture Feature Ranking Model for Predicting Survival Time of Brain Tumor Patients
abstract
Automated prediction of patient-specific disease progression can significantly contribute to clinical treatment. This paper presents a computer-assisted framework to tackle the survival time prediction problem. Inspired by the assumption that niche tumor regions may play a significant role in cancer diagnosis, we explore local visual variations from multiple MRI sequences. The research consists of three parts: 1) the extraction of multi-scale Local Binary Patterns (LBP) to describe the visual variations, 2) a supervised forward feature selection approach, called the Feature Ranking Model (FRM) which captures single feature predictive ability efficiently, and combines the top features to form a feature subset, 3) We cast the clinical survival time prediction task as a binary category classification problem. We tested the framework using a dataset of 32 cases collected from The Cancer Genome Atlas (TCGA). We obtained a 93.75% accuracy rate for the prediction of survival time.
Mu Zhou, Lawrence O. Hall, Dmitry B. Goldgof, Robert A. Gatenby, Robert J. Gillies
SMC2
2013 Automated delineation of lung tumors from CT images using a single click ensemble segmentation approach
Yuhua Gu, Lawrence O. Hall, Dmitry B. Goldgof, Ching-Yen Li, René Korn, Claus Bendtsen, Emmanuel Rios Velazquez, Andre Dekker, Hugo J. W. L. Aerts, Philippe Lambin, Xiuli Li, Robert A. Gatenby, Robert J. Gillies
Pattern Recognit.3
2012 A novel algorithm for automated counting of stained cells on thick tissue sections
abstract
Design-based (unbiased) stereology provides an accurate, precise, and efficient method to quantify morphological parameters of biological microstructures, such as the total number of three-dimensional (3D) objects (cells) in stained tissue sections. The current requirement for extensive user interaction with commercially available computerized stereology systems limits the throughput of data collection. To increase the efficiency of this process, an algorithm was developed to automate data collection from stained objects in thick, transparent tissue sections. We present a novel approach to extract, count and classify stained objects of interest in 3D by linking them through a z-stack of images. Skeletonization and erosion are used to further segment the under segmented (overlapping) cells resulting from the extraction of out of focus cells in conjunction with in focus cells. Finally, 3D shape features, computed from the re-linked cells, are used for final classification of counted objects into “cells” and “not-cells”. We achieve a classification accuracy of 85% using SVM in a leave one-out experiment. The results demonstrate the effectiveness of our algorithm to count cells in 3D from thick, transparent tissue sections.
Baishali Chaudhury, Kurt Kramer, Daniel Elozory, Gerry Hernandez, Dmitry B. Goldgof, Lawrence O. Hall, Peter R. Mouton
CBMS6
2012 Comparison of scalable fuzzy clustering methods
abstract
Fuzzy c-means (FCM) is a well-known algorithm for clustering data, but for large datasets termination takes significant time. As a result, a number of scalable algorithms based on FCM have been developed. In this paper, four scalable variants of FCM are compared to the base algorithm. Runtime and three quality metrics are calculated. Experimental results using five data sets are analyzed. We show that the scalable algorithms are consistent with regard to speedup, but less consistent when quality is considered. The three quality measures are shown to have little correlation and vary in magnitude across datasets. Selection of a scalable algorithm must consider a tradeoff between the quality of results and speed. Of the variants, single pass FCM (SPFCM) is fastest with good fidelity to FCM, and extensible fast FCM (eFFCM) is almost as fast as SPFCM (as implemented) with very good fidelity to FCM. Random FCM is the fastest overall and often close in quality to FCM. The results showed that scalable algorithms occasionally produce better optimized results than FCM.
Jonathon K. Parker, Lawrence O. Hall, James C. Bezdek
FUZZ-IEEE2
2012 Label-noise reduction with support vector machines
Sergiy Fefilatyev, Matthew Shreve, Kurt Kramer, Lawrence O. Hall, Dmitry B. Goldgof, Rangachar Kasturi, Kendra Daly, Andrew Remsen, Horst Bunke
ICPR4
2012 Iterative Feature perturbation as a gene Selector for microarray Data
abstract
Gene-expression microarray datasets often consist of a limited number of samples with a large number of gene-expression measurements, usually on the order of thousands. Therefore, dimensionality reduction is critical prior to any classification task. In this work, the iterative feature perturbation method (IFP), an embedded gene selector, is introduced and applied to four microarray cancer datasets: colon cancer, leukemia, Moffitt colon cancer, and lung cancer. We compare results obtained by IFP to those of support vector machine-recursive feature elimination (SVM-RFE) and the t-test as a feature filter using a linear support vector machine as the base classifier. Analysis of the intersection of gene sets selected by the three methods across the four datasets was done. Additional experiments included an initial pre-selection of the top 200 genes based on their p values. IFP and SVM-RFE were then applied on the reduced feature sets. These results showed up to 3.32% average performance improvement for IFP across the four datasets. A statistical analysis (using the Friedman/Holm test) for both scenarios showed the highest accuracies came from the t-test as a filter on experiments without gene pre-selection. IFP and SVM-RFE had greater classification accuracy after gene pre-selection. Analysis showed the t-test is a good gene selector for microarray data. IFP and SVM-RFE showed performance improvement on a reduced by t-test dataset. The IFP approach resulted in comparable or superior average class accuracy when compared to SVM-RFE on three of the four datasets. The same or similar accuracies can be obtained with different sets of genes.
Juana Canul-Reich, Lawrence O. Hall, Dmitry B. Goldgof, John N. Korecki, Steven Eschrich
Int. J. Pattern Recognit. Artif. Intell.2
2012 Fuzzy c-Means Algorithms for Very Large Data
abstract
Very large (VL) data or big data are any data that you cannot load into your computer's working memory. This is not an objective definition, but a definition that is easy to understand and one that is practical, because there is a dataset too big for any computer you might use; hence, this is VL data for you. Clustering is one of the primary tasks used in the pattern recognition and data mining communities to search VL databases (including VL images) in various applications, and so, clustering algorithms that scale well to VL data are important and useful. This paper compares the efficacy of three different implementations of techniques aimed to extend fuzzy c-means (FCM) clustering to VL data. Specifically, we compare methods that are based on 1) sampling followed by noniterative extension; 2) incremental techniques that make one sequential pass through subsets of the data; and 3) kernelized versions of FCM that provide approximations based on sampling, including three proposed algorithms. We use both loadable and VL datasets to conduct the numerical experiments that facilitate comparisons based on time and space complexity, speed, quality of approximations to batch FCM (for loadable data), and assessment of matches between partitions and ground truth. Empirical results show that random sampling plus extension FCM, bit-reduced FCM, and approximate kernel FCM are good choices to approximate FCM for VL data. We conclude by demonstrating the VL algorithms on a dataset with 5 billion objects and presenting a set of recommendations regarding the use of different VL FCM clustering schemes.
Timothy C. Havens, James C. Bezdek, Christopher Leckie, Lawrence O. Hall, Marimuthu Palaniswami
IEEE Trans. Fuzzy Syst.4
2011 Increased classification accuracy and speedup through pair-wise feature selection for support vector machines
abstract
Support vector machines are binary classifiers that can implement multi-class classifiers by creating a classifier for each possible combination of classes or for each class using a one class versus all strategy. Feature selection algorithms often search for a single set of features to be used by each of the binary classifiers. This ignores the fact that features that may be good discriminators for two particular classes might not do well for other class combinations. As a result, the feature selection process may not include these features in the common set to be used by all support vector machines. It is shown that by selecting features for each binary class combination, overall classification accuracy can be improved (as much as 2.1%), feature selection time can be significantly reduced (speed up of 3.2 times), and time required for training a multi-class support vector machine is reduced. Another benefit of this approach is that considerably less time is required for feature selection when additional classes are added to the training data. This is because the features selected for the existing class combinations are still valid, so that feature selection only needs to be run for the new class combinations created.
Kurt Kramer, Dmitry B. Goldgof, Lawrence O. Hall, Andrew Remsen
CIDM3
2011 Developing a classifier model for lung tumors in CT-scan images
abstract
A CT-scan is a vital tool for the diagnosis of lung cancer via tumor detection. Developing a classifier to make use of the information in CT-scan images could provide a non-invasive alternative to histopathological techniques such as needle biopsy to identify tumor types. Image features extracted from 74 lung tumor objects of CT-scan images are used in classifying tumor types. Classification is done into two major classes of non-small cell lung tumors, Adenocarcinoma and Squamous-cell Carcinoma, each constituting 30% of all lung tumors. In this first of its kind investigation, a large group of 2D and 3D image features which were hypothesized to be useful are evaluated for effectiveness in classifying the tumors. Classifiers including decision trees and support vector machines are used along with feature selection techniques (Wrappers and Relief-F) to build models for tumor classification. Results show that over the large feature space for both 2D and 3D features it is possible to recognize tumor classes with about 68% accuracy, showing new features may be of help. The accuracy achieved using 2D and 3D features is similar with 3D easier to use.
Satrajit Basu, Lawrence O. Hall, Dmitry B. Goldgof, Yuhua Gu, Jung Choi, Robert J. Gillies, Robert A. Gatenby
SMC2
2011 Procedure for stability analysis of gene selection from cross-site gene expression data
abstract
Typically, thousands of gene expression levels are recorded for a group of patients, leading to the situation where the number of features far exceeds the number of examples. To combat this, researchers would want to combine gene expression data collected at different sites into one data set to reduce the magnitude of the difference between the number of features (genes) and examples (samples). This makes gene selection a critical component of any process to build models using gene expression data. For instance, in the domain of ordering cancer patients based on survival time, one might assume that utilizing genes related to cancer development and progression will allow the best model to be built. In this paper, we explore two different gene selection techniques and examine how well the genes selected compare between methods. We also check gene set consistency between data sets collected using the same protocols at different research institutions. It is shown that gene selection can result in very different sets given different training data.
John N. Korecki, Lawrence O. Hall, Dmitry B. Goldgof, Steven Eschrich
SMC2
2011 Detecting and ordering salient regions
Larry Shoemaker, Robert E. Banfield, Lawrence O. Hall, Kevin W. Bowyer, W. Philip Kegelmeyer
Data Min. Knowl. Discov.3
2011 Convergence of the Single-Pass and Online Fuzzy C-Means Algorithms
abstract
Scalable versions of the widely used fuzzy c-means clustering algorithm called single-pass fuzzy c-means and online fuzzy c-means have been recently introduced. Both algorithms facilitate scaling to very large numbers of examples while providing partitions that very closely approximate those one would obtain using fuzzy c-means. Both algorithms have been successfully applied to a number of datasets, most notably, magnetic resonance image volumes of the human brain. In this letter, we show that weighting examples in the fuzzy c-means algorithm does not cause a violation in its convergence proof, and we provide a separate proof of convergence that holds for any dataset.
Lawrence O. Hall, Dmitry B. Goldgof
IEEE Trans. Fuzzy Syst.1
2010 On convergence properties of the singlepass and online fuzzy c-means algorithm
abstract
Single pass fuzzy c-means and Online fuzzy c-means are two scalable versions of the widely used fuzzy c-means clustering algorithm. They both facilitate scaling to very large numbers of examples while providing partitions that very closely approximate those one would obtain using fuzzy c-means. They have been successfully applied to a number of data sets, most notably magnetic resonance image volumes of the human brain. In practice, the algorithms have converged on the data sets to which they been applied. Computers are of finite precision, which will allow real values to be converted to integers with minor loss of information. In this paper, we show that they will converge to local minima or saddle points of the modified objective function for any data set when weights are integers.
Lawrence O. Hall, Dmitry B. Goldgof
FUZZ-IEEE1
2010 Scalable fuzzy neighborhood DBSCAN
abstract
The majority of data available in most disciplines is unlabeled and unclassified. The amount of data is often massive, hence scalable processing methods are required. One method of providing structure to unlabeled data is to group it by clustering. Density based methods discover the number of clusters. Additionally, the shape of such clusters can also be irregular. In this paper we examine a version of DBSCAN modified to use fuzzy membership functions (FN-DBSCAN). FN-DBSCAN was implemented using the WEKA data mining framework and a scalable technique (SFN-DBSCAN) is simulated using the framework. Experimental results show that SFN-DBSCAN can be over three times as fast as FN-DBSCAN for small to medium size data. The resulting cluster assignments match at an average rate of 90% when compared with assignments by FN-DBSCAN. SFN-DBSCAN's speed increases proportionally with respect to the number of subsets, but cluster assignment concurrence between FN-DBSCAN and SFN-DBSCAN suffers from degradation as the number of subsets increase.
Jonathon K. Parker, Lawrence O. Hall, Abraham Kandel
FUZZ-IEEE2
2010 Filtering for improved gene selection on microarray data
abstract
Many genes and a small number of samples are problematic characteristics of microarray datasets. We investigated the impact on classification accuracy of gene selection approaches on filtered-to-200-gene datasets. Four datasets were used with 3 filters: Student's t-test, information gain, and reliefF. We applied Iterative Feature Perturbation (IFP) and Recursive Feature Elimination (SVM-RFE) for further gene selection. Both approaches resulted in accuracy improvement when used with the t-test-filtered datasets, but not when information gain or reliefF were used. An AUC analysis of the IFP and SVM-RFE accuracy curves indicated that both methods reached the highest AUC values after t-test filtering. A statistical analysis of accuracy across the best 50 genes using the Friedman/Holm test showed that IFP and SVM-RFE were significantly more accurate more often when applied to the t-test-filtered gene sets. Surprisingly, the simple t-test, applied as a filter, results in the best overall SVM accuracy and is at least as accurate as the other, more complicated filter methods.
Juana Canul-Reich, Lawrence O. Hall, Dmitry B. Goldgof, Steven Eschrich
SMC2
2010 Evaluating scalable fuzzy clustering
abstract
Clustering large data has the problem of not having all the data fit in the memory at one time. It is a challenge to apply fuzzy clustering algorithms to get a partition in a timely manner. In this paper, we compare the online fuzzy clustering and single pass fuzzy clustering algorithms, which can be used to cluster very large data sets which might be treated as streaming data, with fuzzy c-means. We introduce more meaningful partition comparison measurements based on cluster center location instead of using the difference in Rmvalue. We obtained results on several large volumes of magnetic resonance images which indicate that the online FCM algorithm produces partitions which are very close to what you could get if you clustered all the data at one time. We also show online FCM outperforms single pass FCM and it can process streaming data as it comes without degradation in most cases.
Yuhua Gu, Lawrence O. Hall, Dmitry B. Goldgof
SMC2
2009 Automatic Red Tide Detection from MODIS Satellite Images
abstract
Red tides pose a significant environmental and economic threat in the Gulf of Mexico. Timely detection of red tides is important for understanding this phenomenon. In this paper, learning approaches based on k-nearest neighbors, random forests and support vector machines have been evaluated for red tide detection from MODIS satellite images. Detection results from our algorithms were compared with ground truth red tide data collected in situ. Our results show that red tide identification methods based on machine learning approaches outperform baseline algorithms based on bio-optical characterization.
Weijian Cheng, Lawrence O. Hall, Dmitry B. Goldgof, Chuanmin Hu, Inia M. Soto
SMC2
2009 A scalable framework for cluster ensembles
Prodip Hore, Lawrence O. Hall, Dmitry B. Goldgof
Pattern Recognit.2
2009 Fast Support Vector Machines for Continuous Data
abstract
Support vector machines (SVMs) can be trained to be very accurate classifiers and have been used in many applications. However, the training time and, to a lesser extent, prediction time of SVMs on very large data sets can be very long. This paper presents a fast compression method to scale up SVMs to large data sets. A simple bit-reduction method is applied to reduce the cardinality of the data by weighting representative examples. We then develop SVMs trained on the weighted data. Experiments indicate that bit-reduction SVM produces a significant reduction in the time required for both training and prediction with minimum loss in accuracy. It is also shown to typically be more accurate than random sampling when the data are not overcompressed.
Kurt Kramer, Lawrence O. Hall, Dmitry B. Goldgof, Andrew Remsen
IEEE Trans. Syst. Man Cybern. Part B2
2008 Semi-supervised learning on large complex simulations
abstract
Complex simulations can generate very large amounts of data stored disjointedly across many local disks. Learning from this data can be problematic due to the difficulty of obtaining labels for the data. We present an algorithm for the application of semi-supervised learning on disjoint data generated by complex simulations. Our semi-supervised technique shows a statistically significant accuracy improvement over supervised learning using the same underlying learning algorithm and requires less labeled data for comparable results.
John N. Korecki, Robert E. Banfield, Lawrence O. Hall, Kevin W. Bowyer, W. Philip Kegelmeyer
ICPR3
2008 Detecting and ordering salient regions for efficient browsing
abstract
We describe an ensemble approach to learning salient regions from data partitioned according to the distributed processing requirements of large-scale simulations. The volume of the data is such that classifiers can train only on data local to a given partition. Classes will likely be missing from some, or even most, partitions. We combine a fast ensemble learning algorithm with scaled probabilistic majority voting in order to learn an accurate classifier from such data. We order predicted regions to increase the likelihood that most of the initial set of presented regions are salient. Results from a simulated casing being dropped show that regions of interest are successfully identified and ordered. This approach is much faster than manually browsing and visualizing terabyte or larger simulations to find regions of interest.
Larry Shoemaker, Robert E. Banfield, Lawrence O. Hall, Kevin W. Bowyer, W. Philip Kegelmeyer
ICPR3
2008 Feature selection for microarray data by AUC analysis
abstract
Microarray datasets are often limited to a small number of samples with a large number of gene expressions. Therefore, dimensionality reduction through a feature/gene selection process is highly important for classification purposes. In this paper, a feature perturbation method we previously introduced is applied to do gene selection from microarray data. A publicly available colon cancer dataset is used in our experiments. In comparison with SVM-RFE, our method is better with feature sets of between 10 and 80, however for less than 10 features SVM-RFE results in higher accuracy. An analysis of the area under the curve of the feature perturbation method for the top 50 and 25 features is performed, aiming to determine the proper amount of noise to be applied. We show that a good set of small features/genes can be found using the feature perturbation method.
Juana Canul-Reich, Lawrence O. Hall, Dmitry B. Goldgof, Steven Eschrich
SMC2
2008 Automatically countering imbalance and its empirical relationship to cost
Nitesh V. Chawla, David A. Cieslak, Lawrence O. Hall, Ajay Joshi
Data Min. Knowl. Discov.3
2007 Multivariate Feature Selection using Random Subspace Classifiers for Gene Expression Data
abstract
Gene expression analysis techniques identify important genes that predict specified outcomes based on sample characteristics. Given the small sample sizes common to these studies and the large dimensionality of the data, feature selection methods are essential. In addition, cancer-related expression analysis often involves imbalanced datasets due to rare forms of disease. Popular methods of feature selection employ univariate techniques to identify the features most suitable for analysis. We propose a multivariate technique for selecting accurate subsets of features using an approach based on random subspaces. The random subspace method is used to explore random combinations of features and only subspaces that produce accurate classifiers are retained. The method is tested on two independent gene expression datasets and compared with a univariate approach. The multivariate feature selection method resulted in a 33% improvement in classification accuracy overall and 90% improvement in classification accuracy for the minority class.
Vidya P. Kamath, Lawrence O. Hall, Timothy Yeatman, Steven Eschrich
BIBE2
2007 Ensembles of Fuzzy Classifiers
abstract
The use of bagging is explored to create an ensemble of fuzzy classifiers. The learning algorithm used was ANFIS (adaptive neuro-fuzzy inference systems). We compare results from bagging to those of a single classifier using both crisp and fuzzy classifier combination methods. Results on 20 data sets show that bagging results in a significantly more accurate classifier.
Juana Canul-Reich, Larry Shoemaker, Lawrence O. Hall
FUZZ-IEEE3
2007 Single Pass Fuzzy C Means
abstract
Recently several algorithms for clustering large data sets or streaming data sets have been proposed. Most of them address the crisp case of clustering, which cannot be easily generalized to the fuzzy case. In this paper, we propose a simple single pass (through the data) fuzzy c means algorithm that neither uses any complicated data structure nor any complicated data compression techniques, yet produces data partitions comparable to fuzzy c means. We also show our simple single pass fuzzy c means clustering algorithm when compared to fuzzy c means produces excellent speed-ups in clustering and thus can be used even if the data can be fully loaded in memory. Experimental results using five real data sets are provided.
Prodip Hore, Lawrence O. Hall, Dmitry B. Goldgof
FUZZ-IEEE2
2007 Noise-Based Feature Perturbation as a Selection Method for Microarray Data
Dmitry B. Goldgof, Lawrence O. Hall, Steven Eschrich
ISBRA3
2007 Clinical deployment of a medical expert system to increase accruals for clinical trials: Challenges
abstract
Before new medical treatments become available to the public, clinicians must conduct extensive trials to determine the efficacy of the novel therapy. In order for the clinical trial to be successful, a significant number of patients with an appropriate set of medical conditions must be accrued. We have implemented a web-based expert system at the H. Lee Moffitt Cancer Center & Research Institute in the Gastrointestinal Tumor Clinic (GITC) to help physicians screen patients for phase II trials. Our system allows physicians to screen a patient for multiple trials simultaneously. Our experiments have shown that adaptation of the system into a clinical environment and the success of the system are related to the amount of time physicians are willing to spend entering data. We also found significant regulatory issues (HIPAA) that make implementation challenging.
Sergiy Fefilatyev, Tim V. Ivanovskiy, Lawrence O. Hall, Dmitry B. Goldgof, Shibendra S. Pobi, Halina Greenstien, Amit P. Pathak, Christopher R. Garret
SMC3
2007 A fuzzy c means variant for clustering evolving data streams
abstract
Clustering algorithms for streaming data sets are gaining importance due to the availability of large data streams from different sources. Recently a number of streaming algorithms have been proposed using crisp algorithms such as hard c means or its variants. The crisp cases may not be easily generalized to fuzzy cases as these two groups of algorithms try to optimize different objective functions. In this paper we propose a streaming variant of the fuzzy c means algorithm. At any stage during processing, a good streaming algorithm should be able to summarize data seen so far and also respond to evolving distributions. We study the tradeoff involved between summarization of data seen and response to an evolving distribution by varying the amount of history used by a streaming algorithm. Empirical evaluation of the performance of our algorithm using both artificial and real data sets under a noisy setting shows its effectiveness.
Prodip Hore, Lawrence O. Hall, Dmitry B. Goldgof
SMC2
2007 A Comparison of Decision Tree Ensemble Creation Techniques
abstract
We experimentally evaluate bagging and seven other randomization-based approaches to creating an ensemble of decision tree classifiers. Statistical tests were performed on experimental results from 57 publicly available data sets. When cross-validation comparisons were tested for statistical significance, the best method was statistically more accurate than bagging on only eight of the 57 data sets. Alternatively, examining the average ranks of the algorithms across the group of data sets, we find that boosting, random forests, and randomized trees are statistically significantly better than bagging. Because our results suggest that using an appropriate ensemble size is important, we introduce an algorithm that decides when a sufficient number of classifiers has been created for an ensemble. Our algorithm uses the out-of-bag error estimate, and is shown to result in an accurate ensemble for those methods that incorporate bagging into the construction of the ensemble.
Robert E. Banfield, Lawrence O. Hall, Kevin W. Bowyer, W. Philip Kegelmeyer
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Fuzzy Ants and Clustering
abstract
A swarm-intelligence-inspired approach to clustering data is described. The algorithm consists of two stages. In the first stage of the algorithm, ants move the cluster centers in feature space. The cluster centers found by the ants are evaluated using a reformulated fuzzy C-means (FCM) criterion. In the second stage, the best cluster centers found are used as the initial cluster centers for the FCM algorithm. Results on 18 data sets show that the partitions found using the ant initialization are better optimized than those obtained from random initializations. The use of a reformulated fuzzy partition validity metric as the optimization criterion is shown to enable determination of the number of cluster centers in the data for several data sets. Hard C-means (HCM) was also used after reformulation, and the partitions obtained from the ant-based algorithm were better optimized than those from randomly initialized HCM.
Parag M. Kanade, Lawrence O. Hall
IEEE Trans. Syst. Man Cybern. Part A2
2006 Kernel Based Fuzzy Ant Clustering with Partition Validity
abstract
We introduce a new swarm intelligence based algorithm for data clustering with a kernel-induced distance metric. Previously a swarm based approach using artificial ants to optimize the fuzzy c-means (FCM) criterion using the Euclidean distance was developed. However, FCM is not suitable for clusters which are not hyper-spherical and FCM requires the number of cluster centers be known in advance. The swarm based algorithm determines the number of cluster centers of the input data by using a modification to the fuzzy cluster validity metric proposed by Xie and Beni. The partition validity metric was developed based on the kernelized distance measure. Experiments were done with three data sets; the Iris data, an artificially generated data set and a Magnetic Resonance brain image data set. The results show how effectively the kernelized version of validity metric with a fuzzy ant method finds the number of clusters in the data and that it can be used to partition the data.
Yuhua Gu, Lawrence O. Hall
FUZZ-IEEE2
2006 Horizon Detection Using Machine Learning Techniques
abstract
Detecting a horizon in an image is an important part of many image related applications such as detecting ships on the horizon, flight control, and port security. Most of the existing solutions for the problem only use image processing methods to identify a horizon line in an image. This results in good accuracy for many cases and is fast in computation. However, for some images with difficult environmental conditions like a foggy or cloudy sky these image processing methods are inherently inaccurate in identifying the correct horizon. This paper investigates how to detect the horizon line in a set of images using a machine learning approach. The performance of the SVM, J48, and naive Bayes classifiers, used for the problem, has been compared. Accuracy of 90-99% in identifying horizon was achieved on image data set of 20 images
Sergiy Fefilatyev, Volha Smarodzinava, Lawrence O. Hall, Dmitry B. Goldgof
ICMLA3
2006 Learning to Predict Salient Regions from Disjoint and Skewed Training Sets
abstract
We present an ensemble learning approach that achieves accurate predictions from arbitrarily partitioned data. The partitions come from the distributed processing requirements of a large scale simulation where the volume of the data is such that classifiers can train only on data local to a given partition. As a result of the partition reflecting the need for efficient simulation analysis, rather than the needs of data mining, the class statistics vary across partitions; indeed some classes will likely be absent from some partitions. We combine a fast ensemble learning algorithm with majority voting to generate an accurate working model of the simulation. Results from several simulations show that regions of interest are successfully identified in spite of training set class imbalances. Accuracy is analyzed both at the level of nodes in the simulation data structure, and in terms of higher-level regions of interest. It is shown that over 98% of salient regions are found in independent test sets. Hence, this approach will be a significant time saver for simulation users and developers
Larry Shoemaker, Robert E. Banfield, Lawrence O. Hall, Kevin W. Bowyer, W. Philip Kegelmeyer
ICTAI3
2006 Predicting Juvenile Diabetes from Clinical Test Results
abstract
Two approaches to building models for prediction of the onset of Type 1 diabetes mellitus in juvenile subjects were examined. A set of tests performed immediately before diagnosis was used to build classifiers to predict whether the subject would be diagnosed with juvenile diabetes. A second training set consisting of differences between test results taken at different times was used to build classifiers to predict whether a subject would be diagnosed with juvenile diabetes. Neural networks were compared with decision trees and ensembles of both types of classifiers. The highest known predictive accuracy was obtained when the data was encoded to explicitly indicate missing attributes in both cases. In the latter case, high accuracy was achieved without test results which, by themselves, could indicate diabetes.
Shibendra S. Pobi, Lawrence O. Hall
IJCNN2
2006 A Cluster Ensemble Framework for Large Data sets
abstract
Combining multiple clustering solutions is important for obtaining a robust clustering solution, merging distributed clustering solutions, and scaling to large data sets. The combination of multiple clustering solutions within a scalable and robust framework for large data sets is discussed. A scalable framework requires both cluster ensemble creation and merging to be efficient in terms of time and memory complexity. We also introduce the concept of filtering malformed clusters from the ensemble. They result from unfortunate initialization or unbalanced data distribution or noise. Experimental results on real data sets show that this approach will scale and provide cluster partitions which are functionally better or equivalent when compared to clustering all the data at once and clustering solutions contained in the ensemble. We have also compared our algorithm with other ensemble merging and scalable algorithms to point out its strengths and limitations.
Prodip Hore, Lawrence O. Hall, Dmitry B. Goldgof
SMC2
2006 A constrained genetic approach for computing material property of elastic objects
abstract
This paper presents a constrained genetic approach for reconstructing the material properties of elastic objects. The considered reconstruction problem is ill-posed and must be constrained properly so that a unique and stable numerical solution can be obtained. Qualitative prior information is incorporated using a rank-based scheme to constrain the admissible solutions. Experiments show that the proposed approach is robust when presented with noisy data and can reconstruct the elastic property accurately and reliably. In a comparison study with the deterministic Gauss-Newton methods, the constrained genetic approach also shows very consistent performance.
Yong Zhang 0017, Lawrence O. Hall, Dmitry B. Goldgof, Sudeep Sarkar
IEEE Trans. Evol. Comput.2
2005 Swarm Based Fuzzy Clustering with Partition Validity
abstract
Swarm based approaches to clustering have been shown to be able to skip local extrema by doing a form of global search. We previously reported on the use of a swarm based approach using artificial ants to do fuzzy clustering by optimizing the fuzzy c-means (FCM) criterion. FCM requires that one choose the number of cluster centers (c). In the event that the user of the algorithm is unsure of the number of cluster centers, they can try several different choices and evaluate them with a cluster validity metric. In this work, we use a fuzzy cluster validity metric proposed by Xie and Beni as the criterion for evaluating a partition produced by swarm based clustering. Interestingly, when provided with more clusters than exist in the data our ant-based approach produces a partition with empty clusters and/or very lightly populated clusters. We used two data sets, Iris and an artificially generated data set, to show that optimizing a cluster validity metric with a swarm based approach can effectively provide an indication of how many clusters there are in the data
Lawrence O. Hall, Parag M. Kanade
FUZZ-IEEE1
2005 Bit Reduction Support Vector Machine
abstract
Support vector machines are very accurate classifiers and have been widely used in many applications. However, the training and to a lesser extent prediction time of support vector machines on very large data sets can be very long. This paper presents a fast compression method to scale up support vector machines to large data sets. A simple bit reduction method is applied to reduce the cardinality of the data by weighting representative examples. We then develop support vector machines which may be trained on weighted data. Experiments indicate that the bit reduction support vector machine produces a significant reduction in the time required for both training and prediction with minimum loss in accuracy. It is also shown to be more accurate than random sampling, when the data is not over-compressed.
Lawrence O. Hall, Dmitry B. Goldgof, Andrew Remsen
ICDM2
2005 Sequence tolerant segmentation system of brain MRI
abstract
An automatic human brain segmentation system for magnetic resonance images is presented. It has two main parts: a fuzzy clustering algorithm and a set of cluster combination rules. Images are segmented into ten classes by the unsupervised fuzzy c-means clustering algorithm. Then a knowledge-based system labels the clusters into the tissues of interest: cerebrospinal fluid, gray matter and white matter. This approach can process MRI data that comes from different scanners with different sequences and head coils, using several different spin-echo images (with different echo times) and different slice thickness. The system adapts without manual intervention. Segmented synthetic image data from the brainWeb simulated normal brain database resulted in a one voxel away accuracy of 90%. The results from real data from various magnetic resonance imagers were compared with a radiologist's segmentation and found to generally agree within 10%, the typical range of inter-rater radiologist agreement.
Yuhua Gu, Lawrence O. Hall, Dmitry B. Goldgof, Parag M. Kanade, F. Reed Murtagh
SMC2
2005 Active Learning to Recognize Multiple Types of Plankton
abstract
This paper presents an active learning method which reduces the labeling effort of domain experts in multi-class classification problems. Active learning is applied in conjunction with support vector machines to recognize underwater zooplankton from higher-resolution, new generation SIPPER II images. Most previous work on active learning with support vector machines only deals with two class problems. In this paper, we propose an active learning approach "breaking ties" for multi-class support vector machines using the one-vs-one approach with a probability approximation. Experimental results indicate that our approach often requires significantly less labeled images to reach a given accuracy than the approach of labeling the least certain test example and random sampling. It can also be applied in batch mode resulting in an accuracy comparable to labeling one image at a time and retraining.
Kurt Kramer, Dmitry B. Goldgof, Lawrence O. Hall, Scott Samson, Andrew Remsen, Thomas Hopkins
J. Mach. Learn. Res.4
2004 Using Probabilistic Methods to Optimize Data Entry in Accrual of Patients to Clinical Trials
abstract
A clinical trial is a study conducted on a group of patients to evaluate a new treatment procedure. Usually, clinicians manually select patients for a clinical trial; the choice of eligible patients is a labor-intensive process, and clinicians are often unable to identify sufficient number of patients, which delays the evaluation of new treatments. We have developed a Web-based system that helps clinicians to determine the eligibility of patients for multiple clinical trials. It uses probabilistic techniques that minimize the amount of manual data entry, by ordering the related data-entry steps. We describe the developed system and give the results of applying it to retrospective data of breast cancer patients at the Moffitt Cancer Center.
Bhavesh D. Goswami, Lawrence O. Hall, Dmitry B. Goldgof, Eugene Fink, Jeffrey P. Krischer
CBMS2
2004 Scalable clustering: a distributed approach
abstract
The ever-increasing size of data sets and poor scalability of clustering algorithms has drawn attention to distributed clustering for partitioning large data sets. In this paper we propose an algorithm to cluster large-scale data sets without clustering all the data at a time. Data is randomly divided into almost equal size disjoint subsets. We then cluster each subset using the hard-k means or fuzzy k-means algorithm. The centroids of subsets form an ensemble. A centroid correspondence algorithm transitively solves the correspondence problem among the ensemble of centroids. The centroids are combined to form a global set of centroids. Experimental results show that most of the time the pattern of clusters generated by our algorithm is similar to the pattern of clusters generated by clustering all the data at a time. We have shown that the disputed examples between the clusters generated by our algorithm and clustering all the data at a time lay on the spatial border of clusters.
Prodip Hore, Lawrence O. Hall
FUZZ-IEEE2
2004 Fuzzy ant clustering by centroid positioning
abstract
We present a swarm intelligence based algorithm for data clustering. The algorithm uses ant colony optimization principles to find good partitions of the data. In the first stage of the algorithm ants move the cluster centers in feature space. The cluster centers found by the ants are evaluated using a reformulated fuzzy c-means criterion. In the second stage the best cluster centers found are used as the initial cluster centers for the fuzzy c-means (FCM) algorithm. Results on 8 datasets show that the partitions found by FCM using the ant initialization are better optimized than those from randomly initialized FCM. Hard c-means was also used in the second stage and the partitions from the algorithm are better optimized than those from randomly initialized hard c-means.
Parag M. Kanade, Lawrence O. Hall
FUZZ-IEEE2
2004 Selection of patients for clinical trials: an interactive web-based system
Eugene Fink, Princeton K. Kokku, Savvas Nikiforou, Lawrence O. Hall, Dmitry B. Goldgof, Jeffrey P. Krischer
Artif. Intell. Medicine4
2004 Learning Ensembles from Bites: A Scalable and Accurate Approach
Nitesh V. Chawla, Lawrence O. Hall, Kevin W. Bowyer, W. Philip Kegelmeyer
J. Mach. Learn. Res.2
2004 Comments on "A Parallel Mixture of SVMs for Very Large Scale Problems"
abstract
Collobert, Bengio, and Bengio (2002) recently introduced a novel approach to using a neural network to provide a class prediction from an ensemble of support vector machines (SVMs). This approach has the advantage that the required computation scales well to very large data sets. Experiments on the Forest Cover data set show that this parallel mixture is more accurate than a single SVM, with 90.72% accuracy reported on an independent test set. Although this accuracy is impressive, their article does not consider alternative types of classifiers. We show that a simple ensemble of decision trees results in a higher accuracy, 94.75%, and is computationally efficient. This result is somewhat surprising and illustrates the general value of experimental comparisons using different types of classifiers.
Lawrence O. Hall, Kevin W. Bowyer
Neural Comput.2
2004 List of Reviewers
abstract
The publication offers a note of thanks and lists its reviewers.
Lawrence O. Hall
IEEE Trans. Syst. Man Cybern. Part B1
2004 Recognizing plankton images from the shadow image particle profiling evaluation recorder
abstract
We present a system to recognize underwater plankton images from the shadow image particle profiling evaluation recorder (SIPPER). The challenge of the SIPPER image set is that many images do not have clear contours. To address that, shape features that do not heavily depend on contour information were developed. A soft margin support vector machine (SVM) was used as the classifier. We developed a way to assign probability after multiclass SVM classification. Our approach achieved approximately 90% accuracy on a collection of plankton images. On another larger image set containing manually unidentifiable particles, it also provided 75.6% overall accuracy. The proposed approach was statistically significantly more accurate on the two data sets than a C4.5 decision tree and a cascade correlation neural network. The single SVM significantly outperformed ensembles of decision trees created by bagging and random forests on the smaller data set and was slightly better on the other data set. The 15-feature subset produced by our feature selection approach provided slightly better accuracy than using all 29 features. Our probability model gave us a reasonable rejection curve on the larger data set.
Kurt Kramer, Dmitry B. Goldgof, Lawrence O. Hall, Scott Samson, Andrew Remsen, Thomas Hopkins
IEEE Trans. Syst. Man Cybern. Part B4
2004 Errata to "Recognizing Plankton Images From the Shadow Image Particle Profiling Evaluation Recorder"
Kurt Kramer, Dmitry B. Goldgof, Lawrence O. Hall, Scott Samson, Andrew Remsen, Thomas Hopkins
IEEE Trans. Syst. Man Cybern. Part B4
2003 Learning from soft partitions of data: reducing the variance
abstract
Distributed machine learning can be realized using a divide and conquer methodology. One such divide and conquer method is learning from soft partitions of data. By examining the decomposition of classifier error into bias and variance terms, we see that learning from smaller partitions of data introduces higher variance. In this paper, we investigate the use of a particular variance reduction technique, randomized C4.5, when learning from soft partitions of data. This approach maintains the distributed nature of the learning algorithm while boosting the overall classification accuracy. Experiments on six machine learning datasets demonstrate the improved accuracy gains by reducing classifier variance. In particular, learning from soft partitions of data can produce more accurate classifiers than using an ensemble of randomized decision trees constructed from the entire dataset, which in turn results in a more accurate classifier than building a single decision tree.
Steven Eschrich, Lawrence O. Hall
FUZZ-IEEE2
2003 Comparing Pure Parallel Ensemble Creation Techniques Against Bagging
abstract
We experimentally evaluate randomization-based approaches to creating an ensemble of decision-tree classifiers. Unlike methods related to boosting, all of the eight approaches considered here create each classifier in an ensemble independently of the other classifiers. Experiments were performed on 28 publicly available datasets, using C4.5 release 8 as the base classifier. While each of the other seven approaches has some strengths, we find that none of them is consistently more accurate than standard bagging when tested for statistical significance.
Lawrence O. Hall, Kevin W. Bowyer, Robert E. Banfield, Divya Bhadoria, W. Philip Kegelmeyer, Steven Eschrich
ICDM1
2003 SMOTEBoost: Improving Prediction of the Minority Class in Boosting
Nitesh V. Chawla, Aleksandar Lazarevic, Lawrence O. Hall, Kevin W. Bowyer
PKDD3
2003 Experiments on the automated selection of patients for clinical trials
abstract
When clinicians test a new treatment procedure, they need to identify and recruit patients with appropriate medical conditions. We have developed an expert system that helps clinicians select patients for experimental treatments, and to reduce the number and overall cost of related medical tests. We describe experiments on selecting patients for new treatments at the Moffitt Cancer Center. The experiments have shown that the system can increase the number of selected patients by a factor of three, and that it can also reduce the cost of the selection process.
Eugene Fink, Lawrence O. Hall, Dmitry B. Goldgof, Bhavesh D. Goswami, Matthew Boonstra, Jeffrey P. Krischer
SMC2
2003 Why are neural networks sometimes much more accurate than decision trees: an analysis on a bio-informatics problem
abstract
Bio-informatics data sets may be large in the number of examples and/or the number of features. Predicting the secondary structure of proteins from amino acid sequences is one example of high dimensional data for which large training sets exist. The data from the KDD Cup 2001 on the binding of compounds to thrombin is another example of a very high dimensional data set. This type of data set can require significant computing resources to train a neural network. In general, decision trees will require much less training time than neural networks. There have been a number of studies on the advantages of decision trees relative to neural networks for specific data sets. There are often statistically significant, though typically not very large, differences. Here, we examine one case in which a neural network greatly outperforms a decision tree; predicting the secondary structure of proteins. The hypothesis that the neural network learns important features of the data through its hidden units is explored by a using a neural network to transform data for decision tree training. Experiments show that this explains some of the performance difference, but not all. Ensembles of decision trees are compared with a single neural network. It is our conclusion that the problem of protein secondary structure prediction exhibits some characteristics that are fundamentally better exploited by a neural network model.
Lawrence O. Hall, Kevin W. Bowyer, Robert E. Banfield
SMC1
2003 Learning to recognize plankton
abstract
We present a system to recognize underwater plankton images from the Shadow Image Particle Profiling Evaluation Recorder. As some images do not have clear contours, we developed several features that do not heavily depend on the contour information. A soft margin support vector machine (SVM) was used as the classifier. We developed a new way to assign probability after multi-class SVM classification. Our approach achieved approximately 90% accuracy on a collection of images with minimal noise. On another image set containing manually unidentifiable particles, it also provided promising results. Furthermore, our approach is more accurate on the two data sets than a C4.5 decision tree and a cascade correlation neural network at the 95% confidence level.
Kurt Kramer, Dmitry B. Goldgof, Lawrence O. Hall, Scott Samson, Andrew Remsen, Thomas Hopkins
SMC4
2003 Distributed learning with bagging-like performance
Nitesh V. Chawla, Thomas E. Moore, Lawrence O. Hall, Kevin W. Bowyer, W. Philip Kegelmeyer, Clayton Springer
Pattern Recognit. Lett.3
2003 Visualizing fuzzy points in parallel coordinates
abstract
Exploratory data analysis heavily relies on methods to visualize data and models in a user friendly and interpretable manner. We show how models consisting of a collection of fuzzy points can be visualized in parallel coordinates. In contrast to existing techniques that display only lines representing centroids or shaded areas showing the general variance of cluster centers, the proposed technique shows the spread of the fuzzy membership in each dimension in detail. This allows for a better interpretation of overlap in fuzzy rules.
Michael R. Berthold, Lawrence O. Hall
IEEE Trans. Fuzzy Syst.2
2003 Fast accurate fuzzy clustering through data reduction
abstract
Clustering is a useful approach in image segmentation, data mining, and other pattern recognition problems for which unlabeled data exist. Fuzzy clustering using fuzzy c-means or variants of it can provide a data partition that is both better and more meaningful than hard clustering approaches. The clustering process can be quite slow when there are many objects or patterns to be clustered. This paper discusses the algorithm brFCM, which is able to reduce the number of distinct patterns which must be clustered without adversely affecting the partition quality. The reduction is done by aggregating similar examples and then using a weighted exemplar in the clustering process. The reduction in the amount of clustering data allows a partition of the data to be produced faster. The algorithm is applied to the problem of segmenting 32 magnetic resonance images into different tissue types and the problem of segmenting 172 infrared images into trees, grass and target. Average speed-ups of as much as 59-290 times a traditional implementation of fuzzy c-means were obtained using brFCM, while producing partitions that are equivalent to those produced by fuzzy c-means.
Steven Eschrich, Jingwei Ke, Lawrence O. Hall, Dmitry B. Goldgof
IEEE Trans. Fuzzy Syst.3
2002 A constrained genetic approach for reconstructing Young's modulus of elastic objects from boundary displacement measurements
abstract
This paper presents a constrained genetic approach (CGA) for reconstructing the Young's modulus of elastic objects. Qualitative a priori information is incorporated using a rank based scheme to constrain the admissible solutions. Balance between the fitness function (adhesion to the measurement data) and the penalty function (fidelity to a priori knowledge) is achieved by a stochastic sort algorithm. The over-smoothing of Young's modulus discontinuity is avoided without the need of computing a deterministic weight coefficient. The experiment on synthetic data indicates that the proposed method not only reconstructed reliable Young's modulus from noisy data, but also expedited the convergence process significantly.
Yong Zhang 0017, Lawrence O. Hall, Dmitry B. Goldgof, Sudeep Sarkar
IEEE Congress on Evolutionary Computation2
2002 Text classification with enhanced semi-supervised fuzzy clustering
abstract
Given the increasing volume of information available on the Web, it is important to meaningfully organize online documents. Hence, the design of efficient and accurate text classification systems is of interest. In this paper, we explore a framework, in which we improve the performance of a base classifier, by clustering unlabeled data with labeled data using probabilistic and fuzzy approaches. We have used expectation maximization and semi-supervised fuzzy c-means for clustering the unlabeled data with labeled data. The naive Bayes classifier was the base classifier utilizing both the original labeled data and then additional data labeled through clustering. Utilizing unlabeled data from semi-supervised fuzzy clustering results in an improved classifier.
Girish Keswani, Lawrence O. Hall
FUZZ-IEEE2
2002 Error-Based Pruning of Decision Trees Grown on Very Large Data Sets Can Work!
abstract
It has been asserted that, using traditional pruning methods, growing decision trees with increasingly larger amounts of training data will result in larger tree sizes even when accuracy does not increase. With regard to error-based pruning, the experimental data used to illustrate this assertion have apparently been obtained using the default setting for pruning strength; in particular, using the default certainty factor of 25 in the C4.5 decision tree implementation. We show that, in general, an appropriate setting of the certainty factor for error-based pruning will cause decision tree size to plateau when accuracy is not increasing with more training data.
Lawrence O. Hall, Richard Collins, Kevin W. Bowyer, Robert E. Banfield
ICTAI1
2002 SMOTE: Synthetic Minority Over-sampling Technique
abstract
An approach to the construction of classifiers from imbalanced datasets is described. A dataset is imbalanced if the classification categories are not approximately equally represented. Often real-world data sets are predominately composed of ``normal'' examples with only a small percentage of ``abnormal'' or ``interesting'' examples. It is also the case that the cost of misclassifying an abnormal (interesting) example as a normal example is often much higher than the cost of the reverse error. Under-sampling of the majority (normal) class has been proposed as a good means of increasing the sensitivity of a classifier to the minority class. This paper shows that a combination of our method of over-sampling the minority (abnormal) class and under-sampling the majority (normal) class can achieve better classifier performance (in ROC space) than only under-sampling the majority class. This paper also shows that a combination of our method of over-sampling the minority class and under-sampling the majority class can achieve better classifier performance (in ROC space) than varying the loss ratios in Ripper or class priors in Naive Bayes. Our method of over-sampling the minority class involves creating synthetic minority class examples. Experiments are performed using C4.5, Ripper and a Naive Bayes classifier. The method is evaluated using the area under the Receiver Operating Characteristic curve (AUC) and the ROC convex hull strategy.
Nitesh V. Chawla, Kevin W. Bowyer, Lawrence O. Hall, W. Philip Kegelmeyer
J. Artif. Intell. Res.3
2002 A generic knowledge-guided image segmentation and labeling system using fuzzy clustering algorithms
abstract
Segmentation of an image into regions and the labeling of the regions is a challenging problem. In this paper, an approach that is applicable to any set of multifeature images of the same location is derived. Our approach applies to, for example, medical images of a region of the body; repeated camera images of the same area; and satellite images of a region. The segmentation and labeling approach described here uses a set of training images and domain knowledge to produce an image segmentation system that can be used without change on images of the same region collected over time. How to obtain training images, integrate domain knowledge, and utilize learning to segment and label images of the same region taken under any condition for which a training image exists is detailed. It is shown that clustering in conjunction with image processing techniques utilizing an iterative approach can effectively identify objects of interest in images. The segmentation and labeling approach described here is applied to color camera images and two other image domains are used to illustrate the applicability of the approach.
Lawrence O. Hall, Dmitry B. Goldgof
IEEE Trans. Syst. Man Cybern. Part B2
2001 Bagging Is a Small-Data-Set Phenomenon
abstract
Bagging forms a committee of classifiers by bootstrap aggregation of training sets from a pool of training data. A simple alternative to bagging is to partition the data into disjoint subsets. Experiments on various datasets show that, given the same size partitions and bags, disjoint partitions result in better performance than bootstrap aggregates (bags). Many applications (e.g., protein structure prediction) involve the use of datasets that are too large to handle in the memory of a typical computer. Our results indicate that, in such applications, the simple approach of creating a committee of classifiers from disjoint partitions is preferred over the more complex approach of bagging.
Nitesh V. Chawla, Thomas E. Moore, Kevin W. Bowyer, Lawrence O. Hall, Clayton Springer, W. Philip Kegelmeyer
CVPR (2)4
2001 Text Extraction from Color Documents - Clustering Approaches in Three and Four Dimensions
abstract
Colored paper documents often contain important text information. For automating the retrieval process, identification of text elements is essential. In order to reduce the number of colors in a scanned document, color clustering is usually done first. In this article two histogram-based color clustering algorithms are investigated. The first is based on the RGB color space exclusively, while the second takes spatial information into account, in addition to the colors. Experimental results have shown that the use of spatial information in the clustering algorithm has a positive impact. Thus the automatic retrieval of text information can be improved. The proposed methods for clustering are not restricted to document images. They can also be used for processing Web or video images, for example.
T. Perroud, Karin Sobottka, Horst Bunke, Lawrence O. Hall
ICDAR4
2001 Creating Ensembles of Classifiers
abstract
Ensembles of classifiers offer promise in increasing overall classification accuracy. The availability of extremely large datasets has opened avenues for application of distributed and/or parallel learning to efficiently learn models of them. In this paper, distributed learning is done by training classifiers on disjoint subsets of the data. We examine a random partitioning method to create disjoint subsets and propose a more intelligent way of partitioning into disjoint subsets using clustering. It was observed that the intelligent method of partitioning generally performs better than random partitioning for our datasets. In both methods a significant gain in accuracy may be obtained by applying bagging to each of the disjoint subsets, creating multiple diverse classifiers. The significance of our finding is that a partition strategy for even small/moderate sized datasets when combined with bagging can yield better performance than applying a single learner using the entire dataset.
Nitesh V. Chawla, Steven Eschrich, Lawrence O. Hall
ICDM3
2001 Data mining from extreme data sets: very large and/or very skewed data sets
abstract
The article describes an approach to the construction of classifiers from imbalanced data sets. A data set is imbalanced if the classification categories are not approximately equally represented. Often real-world data sets are predominately composed of normal examples with only a small percentage of abnormal or interesting examples. It is also the case that the cost of misclassifying an abnormal (interesting) example as a normal example is often much higher than the cost of the reverse error. Under-sampling of the majority (normal) class has been proposed as a good means of increasing the sensitivity of a classier to the minority class. We discuss a combination of over-sampling the minority (abnormal) class and under-sampling the majority (normal) class can achieve better classifier performance than only under-sampling the majority class. Our method of over-sampling the minority class involves creating synthetic minority class examples. Performance is measured using the area under the receiver operating characteristic curve. It is shown that generally a more diverse set of operating points can be found with the combination of over and undersampling of an imbalanced data set. Usually, the best of the true positives with minimal false negatives is found when compared with loss ratios, different classification costs, etc. Details are provided.
Lawrence O. Hall
SMC1
2001 Automatic segmentation of non-enhancing brain tumors in magnetic resonance images
Lynn M. Fletcher-Heath, Lawrence O. Hall, Dmitry B. Goldgof, F. Reed Murtagh
Artif. Intell. Medicine2
2001 Report of research activities in fuzzy AI and medicine at USF CSE
Horia-Nicolai L. Teodorescu, Abraham Kandel, Lawrence O. Hall
Artif. Intell. Medicine3
2001 Rule chaining in fuzzy expert systems
abstract
A fuzzy expert system must do rule chaining differently than a nonfuzzy expert system. In particular, any rule that can fire with a particular linguistic variable in its consequent must fire before any rule whose antecedent conditions depend upon the resultant fuzzy set value of the consequent linguistic variable is allowed to fire. The dependent rules would be considered in a chain with the fuzzy rules which generate or assert the needed fuzzy linguistic variable. A recent paper by J. Pan et al. (1998) points out that a version of the FuzzyCLIPS expert system shell does not operate with chained fuzzy rules as one would expect. They introduce FuzzyShell which is described as the only known shell to have the expected fuzzy rule chaining performance. We show several approaches to obtaining the desired behavior in FuzzyCLIPS. Further, a potential pitfall with the FuzzyShell approach to dealing with chaining is pointed out.
Lawrence O. Hall
IEEE Trans. Fuzzy Syst.1
2001 SMCB goes all electronic
K. Patipati, Lawrence O. Hall
IEEE Trans. Syst. Man Cybern. Part B2
2000 Chaining in fuzzy rule-based systems
abstract
A fuzzy expert system must do rule chaining differently than a non-fuzzy expert system. In particular, all rules that can fire with a particular linguistic variable in their consequent must fire before any rule whose antecedent conditions depend upon the resultant fuzzy set value is allowed to fire. The dependent rules would be considered in a chain with the fuzzy rules which generate or assert the needed fuzzy linguistic variable. A recent paper by Pan-DeSouza-Kak (1998) points out that the FuzzyCLIPS expert system shell does not operate with chained fuzzy rules as one would expect. They introduce FuzzyShell which is described as the only known shell to have the expected fuzzy rule chaining performance. In this paper, we show how FuzzyCLIPS can be made to have the required property. Further, a potential pitfall with the FuzzyShell approach to dealing with chaining is pointed out.
Lawrence O. Hall
FUZZ-IEEE1
2000 Finding Green River in SeaWiFS Satellite Images
abstract
Understanding oceanic primary production on a global scale can be enhanced by methods that are able to automatically track phytoplankton blooms from color satellite images. In the paper, unsupervised clustering and rule learning are combined to track green river, a plume of discolored water that forms every March-May offshore along the edge of the west Florida Shelf, from the Sea Viewing Wide Field of View Sensor which began flying in late 1997. Spatial information and sea surface temperature can be integrated into the approach to improve performance. Using cross-validation experiments over a series of 59 multi-spectral images, it is shown that the developed system is able to reliably discriminate between images with green river from those with no phytoplankton blooms or other kinds of blooms. It is also effective in identifying the region which the green river covers.
Wensheng Yao, Lawrence O. Hall, Dmitry B. Goldgof, Frank E. Müller-Karger
ICPR2
2000 A parallel decision tree builder for mining very large visualization datasets
abstract
Simulation problems in the DOE ASCI program generate visualization datasets more than a terabyte in size. The practical difficulties in visualizing such datasets motivate the desire for automatic recognition of salient events. We have developed a parallel decision tree classifier for use in this context. Comparisons to ScalParC, a previous attempt to build a fast parallelization of a decision tree classifier, are provided. Our parallel classifier executes on the "ASCI Red" supercomputer. Experiments demonstrate that datasets too large to be processed on a single processor can be efficiently handled in parallel, and suggest that there need not be any decrease in accuracy relative to a monolithic classifier constructed on a single processor.
Kevin W. Bowyer, Lawrence O. Hall, Nitesh V. Chawla, W. Philip Kegelmeyer
SMC2
2000 Knowledge-Guided Classification of Coastal Zone Color Images Off the West Florida Shelf
abstract
A knowledge-guided approach to automatic classification of Coastal Zone Color images off the West Florida Shelf is described. The approach is used to identify red tides on the West Florida Shelf, as well as areas with high concentration of dissolved organic matter such as a river plume found seasonally along the West Florida coast over the middle of the shelf. The Coastal Zone Color images are initially segmented by the unsupervised Multistage Random Sampling Fuzzy c-Means algorithm. Then, a knowledge-guided system is applied to the centroid values of resultant clusters to label case I, case II waters, a dilute river plume ("green river"), and red tide. The domain knowledge base contains information on cluster distribution in feature space, as well as spatial information such as bathymetry data. Our knowledge base consists of a rule-guided system and an embedded neural network. From 60 images, after training the system, this procedure recognizes all 15 images which contained a river plume and 45 images without. The system can correctly classify 74% of the pixels that belong to the river plume, which provides a substantial advantage to users looking for offshore extensions of riverine influence. Red tides are also successfully identified in a time series of images for which ground truth confirmed the presence of a harmful bloom.
Lawrence O. Hall, Dmitry B. Goldgof, Frank E. Müller-Karger
Int. J. Pattern Recognit. Artif. Intell.2
1999 Clustering with a genetically optimized approach
abstract
Describes a genetically guided approach to optimizing the hard (J/sub 1/) and fuzzy (J/sub m/) c-means functionals used in cluster analysis. Our experiments show that a genetic algorithm (GA) can ameliorate the difficulty of choosing an initialization for the c-means clustering algorithms. Experiments use six data sets, including the Iris data, magnetic resonance, and color images. The genetic algorithm approach is generally able to find the lowest known J/sub m/ value or a J/sub m/ associated with a partition very similar to that associated with the lowest J/sub m/ value. On data sets with several local extrema, the GA approach always avoids the less desirable solutions. Degenerate partitions are always avoided by the GA approach, which provides an effective method for optimizing clustering models whose objective function can be represented in terms of cluster centers. A series random initializations of fuzzy/hard c-means, where the partition associated with the lowest J/sub m/ value is chosen, can produce an equivalent solution to the genetic guided clustering approach given the same amount of processor time in some domains.
Lawrence O. Hall, Ibrahim Burak Özyurt, James C. Bezdek
IEEE Trans. Evol. Comput.1
1999 SQFDiag: semi-quantitative model-based fault monitoring and diagnosis via episodic fuzzy rules
abstract
A method for chemical process fault diagnosis using semi-quantitative model generated behavior envelopes is described. The method generates a sequence of rules for each fault class, with any rule in a sequence valid within the bounds of its time interval. This can be viewed as a qualitative description of the trend of numerical sensor measurements. For each variable in each fault class two sequences of episodic fuzzy rules are automatically generated, one for the lower and one for the upper numerical behavior envelope. The diagnostic system monitors a process via the measured sensors. The measurements are matched against the fuzzy rules for the current time in the rule base. In the case of an overlapping region defined by behavior envelopes, the distance introduced and time based fault belief scaling allows ranking of fault candidates. A novel abnormal situation will not pass the introduced system undetected due to a novel class detection mechanism. In two case studies, the system detected the correct fault even in cases of nearly total overlapped fault regions bounded by behavior envelopes.
Ibrahim Burak Özyurt, Lawrence O. Hall, Aydin K. Sunol
IEEE Trans. Syst. Man Cybern. Part A2
1998 Decision tree learning on very large data sets
abstract
Consider a labeled data set of 1 terabyte in size. A salient subset might depend upon the users interests. Clearly, browsing such a large data set to find interesting areas would be very time consuming. An intelligent agent which, for a given class of user, could provide hints on areas of the data that might interest the user would be very useful. Given large data sets having categories of salience for different user classes attached to the data in them, these labeled sets of data can be used to train a decision tree to label unseen data examples with a category of salience. The training set will be much larger than usual. This paper describes an approach to generating the rules for an agent from a large training set. A set of decision trees are built in parallel on tractable size training data sets which are a subset of the original data. Each learned decision tree will be reduced to a set of rules, conflicting rules resolved and the resultant rules merged into one set. Results from cross validation experiments on a data set suggest this approach may be effectively applied to large sets of data.
Lawrence O. Hall, Nitesh V. Chawla, Kevin W. Bowyer
SMC1
1998 Fast fuzzy clustering
Tai Wai Cheng, Dmitry B. Goldgof, Lawrence O. Hall
Fuzzy Sets Syst.3
1998 Automatic Tumor Segmentation Using Knowledge-Based Clustering
abstract
A system that automatically segments and labels glioblastoma-multiforme tumors in magnetic resonance images (MRI's) of the human brain is presented. The MRI's consist of T1-weighted, proton density, and T2-weighted feature images and are processed by a system which integrates knowledge-based (KB) techniques with multispectral analysis. Initial segmentation is performed by an unsupervised clustering algorithm. The segmented image, along with cluster centers for each class are provided to a rule-based expert system which extracts the intracranial region. Multispectral histogram analysis separates suspected tumor from the rest of the intracranial region, with region analysis used in performing the final tumor labeling. This system has been trained on three volume data sets and tested on thirteen unseen volume data sets acquired from a single MRI system. The KB tumor segmentation was compared with supervised, radiologist-labeled "ground truth" tumor volumes and supervised k-nearest neighbors tumor segmentations. The results of this system generally correspond well to ground truth, both on a per slice basis and more importantly in tracking total tumor volume during treatment over time.
Matthew C. Clark, Lawrence O. Hall, Dmitry B. Goldgof, Robert P. Velthuizen, F. Reed Murtagh, Martin L. Silbiger
IEEE Trans. Medical Imaging2
1997 Confirmation and denial as plausible modes of fuzzy inference
Lawrence O. Hall
Fuzzy Sets Syst.1
1997 An investigation of mountain method clustering for large data sets
Robert P. Velthuizen, Lawrence O. Hall, Laurence P. Clarke, Martin L. Silbiger
Pattern Recognit.2
1996 Knowledge-based classification of CZCS images and monitoring of red tides off the west Florida shelf
abstract
Red tides on the west Florida shelf have significant economic and public health effects. Tracking the phytoplankton bloom, known as red tide, is important to understanding the phenomena. In this paper, a knowledge-based approach to automatic classification of Coastal Zone Color Scanner satellite images is developed. The Coastal Zone Color Scanner or CZCS images are initially segmented by the unsupervised mr-FCM algorithm then an expert system utilizes rules, and an iterative clustering process, to recognize case I (deep) water, case II (shallow) water and red tide by searching for expected features. The results show that this system is effective in recognizing images with red tide and segmenting the red tide.
Lawrence O. Hall, Dmitry B. Goldgof
ICPR2
1996 Partially supervised clustering for image segmentation
Amine Bensaid, Lawrence O. Hall, James C. Bezdek, Laurence P. Clarke
Pattern Recognit.2
1996 Validity-guided (re)clustering with applications to image segmentation
abstract
When clustering algorithms are applied to image segmentation, the goal is to solve a classification problem. However, these algorithms do not directly optimize classification duality. As a result, they are susceptible to two problems: 1) the criterion they optimize may not be a good estimator of "true" classification quality, and 2) they often admit many (suboptimal) solutions. This paper introduces an algorithm that uses cluster validity to mitigate problems 1 and 2. The validity-guided (re)clustering (VGC) algorithm uses cluster-validity information to guide a fuzzy (re)clustering process toward better solutions. It starts with a partition generated by a soft or fuzzy clustering algorithm. Then it iteratively alters the partition by applying (novel) split-and-merge operations to the clusters. Partition modifications that result in improved partition validity are retained. VGC is tested on both synthetic and real-world data. For magnetic resonance image (MRI) segmentation, evaluations by radiologists show that VGC outperforms the (unsupervised) fuzzy c-means algorithm, and VGC's performance approaches that of the (supervised) k-nearest-neighbors algorithm.
Amine Bensaid, Lawrence O. Hall, James C. Bezdek, Laurence P. Clarke, Martin L. Silbiger, John A. Arrington, F. Reed Murtagh
IEEE Trans. Fuzzy Syst.2
1996 Fuzzy Systems Toolbox, Fuzzy Logic Toolbox [Software Review]
Lawrence O. Hall, Richard J. Hathaway
IEEE Trans. Fuzzy Syst.1
1995 Learning Membership Functions in a Function-Based Object Recognition System
abstract
Functionality-based recognition systems recognize objects at the category level by reasoning about how well the objects support the expected function. Such systems naturally associate a ``measure of goodness'' or ``membership value'' with a recognized object. This measure of goodness is the result of combining individual measures, or membership values, from potentially many primitive evaluations of different properties of the object's shape. A membership function is used to compute the membership value when evaluating a primitive of a particular physical property of an object. In previous versions of a recognition system known as Gruff, the membership function for each of the primitive evaluations was hand-crafted by the system designer. In this paper, we provide a learning component for the Gruff system, called Omlet, that automatically learns membership functions given a set of example objects labeled with their desired category measure. The learning algorithm is generally applicable to any problem in which low-level membership values are combined through an and-or tree structure to give a final overall membership value.
Kevin S. Woods, Diane J. Cook, Lawrence O. Hall, Kevin W. Bowyer, Louise Stark
J. Artif. Intell. Res.3
1994 Knowledge based (re-)clustering
abstract
This paper presents a novel multiparadigm segmentation method based upon knowledge based clustering with reclustering. The techniques described enhance unsupervised classification and achieve pattern labeling. First domain knowledge is utilized to decide where and how a clustering algorithm is applied, then clustering is iteratively applied to focus-of-attention patterns with the knowledge of how many expected classes there are in those patterns and the prototypical patterns of a class. Examples showing clustering improvements are given from brain MRI's and satellite images.
Matthew C. Clark, Lawrence O. Hall, Chunlin Li 0002, Dmitry B. Goldgof
ICPR (2)2
1993 Parallel search using transformation-ordering Lterative-Deepening-A
abstract
Iterative-Deepening-A (IDA*) is an optimal search technique which is useful for large search spaces, because it requires no intermediate state storage. We show how Transformation-Ordering Iterative-Deepening-A* (TOIDA*) improves the performance of IDA* by dynamically modifying the node expansion order based on results from previous cost limits. We then describe a window parallel implementation of TOIDA* on a Hypercube, and present empirical evidence that the parallel implementation dramatically reduces time spent in search. Finally, we analyze the best and worst case results of sequential and parallel TOIDA*, and compare the results with those of standard IDA* search. Empirical and analytical results show that TOIDA* can provide significant improvements in search speed over IDA* with no penalty in storage requirements, and parallel TOIDA* offers substantial cost reduction over sequential TOIDA*, though at the cost of optimality. © 1993 John Wiley & Sons, Inc.
Diane J. Cook, Lawrence O. Hall, Willard Thomas
Int. J. Intell. Syst.2
1993 Methods for Combination of Evidence in Function-Based 3-D Object Recognition
abstract
Representation schemes traditionally used in model-based vision are contrasted with the “function-based” representation scheme. A system which utilizes function-based representation has been implemented and tested, using the object category “chair” for case study. Function-based description is used to recognize classes and identify subclasses of known categories of objects, even if the specific object has never been encountered previously. Interpretation of the functionality of an object is accomplished through qualitative reasoning about its 3-D shape. During the recognition process, evidence is gathered as to how well the functional requirements are satisfied by the input shape. An investigation of different types of operators used in the combination of the functional evidence has been made. Three pairs of conjunctive and disjunctive operators have been used in the recognition process of more than 100 object shapes. The results are compared and differences are discussed.
Louise Stark, Lawrence O. Hall, Kevin W. Bowyer
Int. J. Pattern Recognit. Artif. Intell.2
1993 SC-net: A hybrid connectionist, symbolic system
Steve G. Romaniuk, Lawrence O. Hall
Inf. Sci.2
1993 Divide and Conquer Neural Networks
Steve G. Romaniuk, Lawrence O. Hall
Neural Networks2
1993 Knowledge-based classification and tissue labeling of MR images of human brain
abstract
Presents a knowledge-based approach to automatic classification and tissue labeling of 2D magnetic resonance (MR) images of the human brain. The system consists of 2 components: an unsupervised clustering algorithm and an expert system. MR brain data is initially segmented by the unsupervised algorithm, then the expert system locates a landmark tissue or cluster and analyzes it by matching it with a model or searching in it for an expected feature. The landmark tissue location and its analysis are repeated until a tumor is found or all tissues are labeled. The knowledge base contains information on cluster distribution in feature space and tissue models. Since tissue shapes are irregular, their models and matching are specially designed: 1) qualitative tissue models are defined for brain tissues such as white matter; 2) default reasoning is used to match a model with an MR image; that is, if there is no mismatch between a model and an image, they are taken as matched. The system has been tested with 53 slices of MR images acquired at different times by 2 different scanners. It accurately identifies abnormal slices and provides a partial labeling of the tissues. It provides an accurate complete labeling of all normal tissues in the absence of large amounts of data nonuniformity, as verified by radiologists. Thus the system can be used to provide automatic screening of slices for abnormality. It also provides a first step toward the complete description of abnormal images for use in automatic tumor volume determination.
Chunlin Li 0002, Dmitry B. Goldgof, Lawrence O. Hall
IEEE Trans. Medical Imaging3
1992 A Hybrid/Symbolic Connectionist Production System
abstract
The task of implementing a simple rule-based production system in terms of a hybrid symbolic/connectionist architecture is discussed. The aim is to achieve parallel execution of rules. The architecture uses a local representation and builds upon prior work, which resulted in the symbolic/connectionist expert system development tool SC-net. The Hybrid Symbolic/Connectionist Production System (HSC-PS) reported supports two types of working memory elements, the attribute/value pair and the object/attribute/value triplet. HSC-PS can provide variable binding in an individual condition and across conditions and variable value instantiation from the left-hand side to the right-hand side in a rule. The network is constructed from production rules such that only one value is bound to a variable at any one time. An implementation of the farmer's dilemma planning problem is given to illustrate the potential of this approach.>
Katsuaki Sanou, Steve G. Romaniuk, Lawrence O. Hall
ICTAI3
1992 Experimental results from parallel backward-chained expert systems
abstract
There are many applications which may be done by an expert system in real time, if the system is capable of real time response. the first Lisp- and Prolog-based expert systems have typically been too slow for real time response. This has lead to an effort to use other languages, the development of fast pattern matching techniques, and other methods of improving the speed of expert systems. Another approach to developing faster expert systems is to make use of the emerging parallel processing computer technology. A further use for parallelism is to allow reasonable response time for large knowledge bases. the size of knowledge bases may become as large as 20,000 chunks of knowledge (and more) in the near future in medical and space applications. This article describes the use of parallel processing in the EMYCIN backward chained rule-based model. Performance on two examples of shared memory multiprocessors is presented and contrasted with earlier simulations.
Lawrence O. Hall
Int. J. Intell. Syst.1
1992 Backpac: A parallel goal-driven reasoning system
Lawrence O. Hall
Inf. Sci.1
1992 A comparison of neural network and fuzzy clustering techniques in segmenting magnetic resonance images of the brain
abstract
Magnetic resonance (MR) brain section images are segmented and then synthetically colored to give visual representations of the original data with three approaches: the literal and approximate fuzzy c-means unsupervised clustering algorithms, and a supervised computational neural network. Initial clinical results are presented on normal volunteers and selected patients with brain tumors surrounded by edema. Supervised and unsupervised segmentation techniques provide broadly similar results. Unsupervised fuzzy algorithms were visually observed to show better segmentation when compared with raw image data for volunteer studies. For a more complex segmentation problem with tumor/edema or cerebrospinal fluid boundary, where the tissues have similar MR relaxation behavior, inconsistency in rating among experts was observed, with fuzz-c-means approaches being slightly preferred over feedforward cascade correlation results. Various facets of both approaches, such as supervised versus unsupervised learning, time complexity, and utility for the diagnostic process, are compared.
Lawrence O. Hall, Amine Bensaid, Laurence P. Clarke, Robert P. Velthuizen, Martin L. Silbiger, James C. Bezdek
IEEE Trans. Neural Networks1
1990 A Hybrid Connectionist, Symbolic Learning System
Lawrence O. Hall, Steve G. Romaniuk
AAAI1
1988 On the choice of ply operators for modus ponens generation in fuzzy intelligent systems
Lawrence O. Hall
Int. J. Approx. Reason.1
1988 On the validation and testing of fuzzy expert systems
abstract
A reusable expert system called Fess was designed to be easy to validate and has been used in several different domains. The validation and testing of the system from the building stage onward is discussed.>
Lawrence O. Hall, Menahem Friedman, Abraham Kandel
IEEE Trans. Syst. Man Cybern.1
1986 On the derivation of memberships for fuzzy sets in expert systems
Lawrence O. Hall, Sue Szabo, Abraham Kandel
Inf. Sci.1