Donald E. Brown

dblp:19/1149 · DBLP profile ↗
← Back
79ranked-venue papers
21as first author
16since 2021 · last 2025
0000-0002-9140-2632ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 28 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 26 · 10 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 23 · 7 first-author · 1 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 2 since 2021Security and privacy · 7 · 1 first-author · 1 since 2021Computer networks · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Theory of computation · 1
YearPublicationVenuePosition
2025 MisstepMath: A Diverse Student Mistake Dataset for AI in Mathematics Teacher Training
Shahina Mohd Azam Ansari, James P. Bywater, Sarah Lilly, Donald E. Brown, Jennifer L. Chiu
AIED (1)4
2024 Gengmm: Generalized Gaussian-Mixture-Based Domain Adaptation Model for Semantic Segmentation
abstract
Domain adaptive semantic segmentation is the task of generating precise and dense predictions for an unlabeled target domain using a model trained on a labeled source domain. While significant efforts have been devoted to improving unsupervised domain adaptation for this task, it is crucial to note that many models rely on a strong assumption that the source data is entirely and accurately labeled, while the target data is unlabeled. In real-world scenarios, however, we often encounter partially or noisy labeled data in source and target domains, referred to as Generalized Domain Adaptation (GDA). In such cases, we suggest leveraging weak or unlabeled data from both domains to narrow the gap between them, resulting in effective adaptation. We introduce the Generalized Gaussian-mixture-based (GenGMM) domain adaptation model, which harnesses the underlying data distribution in both domains to refine noisy weak and pseudo labels. The experiments demonstrate the effectiveness of our approach.
Nazanin Moradinasab, Hassan Jafarzadeh, Donald E. Brown
ICIP3
2024 ProtoGMM: Multi-prototype Gaussian-Mixture-based Domain Adaptation Model for Semantic Segmentation
abstract
Domain adaptive semantic segmentation aims to generate accurate and dense predictions for an unlabeled target domain by leveraging a supervised model trained on a labeled source domain. The prevalent self-training approach involves retraining the dense discriminative classifier of p(class/pixel feature) using the pseudo-labels from the target domain. While many methods focus on mitigating the issue of noisy pseudo-labels, they often overlook the underlying data distribution p(pixel feature/class) in both the source and target domains. To address this limitation, we propose the multi-prototype Gaussian-Mixture-based (ProtoGMM) model, which incorporates the GMM into contrastive losses to perform guided contrastive learning. Contrastive losses are commonly executed in the literature using memory banks, which can lead to class biases due to underrepresented classes. Furthermore, memory banks often have fixed capacities, potentially restricting the model's ability to capture diverse representations of the target/source domains. An alternative approach is to use global class prototypes (i.e. averaged features per category). However, the global prototypes are based on the unimodal distribution assumption per class, disregarding within-class variation. To address these challenges, we propose the ProtoGMM model. This novel approach involves estimating the underlying multi-prototype source distribution by utilizing the GMM on the feature space of the source samples. The components of the GMM model act as representative prototypes. To achieve increased intra-class semantic similarity, decreased inter-class similarity, and domain alignment between the source and target domains, we employ multi-prototype contrastive learning between source distribution and target samples. The experiments show the effectiveness of our method on UDA benchmarks.
Nazanin Moradinasab, Laura S. Shankman, Rebecca A. Deaton, Gary K. Owens, Donald E. Brown
ICMLA5
2024 Computer vision digitization of smartphone images of anesthesia paper health records from low-middle income countries
abstract
BACKGROUND: In low-middle income countries, healthcare providers primarily use paper health records for capturing data. Paper health records are utilized predominately due to the prohibitive cost of acquisition and maintenance of automated data capture devices and electronic medical records. Data recorded on paper health records is not easily accessible in a digital format to healthcare providers. The lack of real time accessible digital data limits healthcare providers, researchers, and quality improvement champions to leverage data to improve patient outcomes. In this project, we demonstrate the novel use of computer vision software to digitize handwritten intraoperative data elements from smartphone photographs of paper anesthesia charts from the University Teaching Hospital of Kigali. We specifically report our approach to digitize checkbox data, symbol-denoted systolic and diastolic blood pressure, and physiological data. METHODS: We implemented approaches for removing perspective distortions from smartphone photographs, removing shadows, and improving image readability through morphological operations. YOLOv8 models were used to deconstruct the anesthesia paper chart into specific data sections. Handwritten blood pressure symbols and physiological data were identified, and values were assigned using deep neural networks. Our work builds upon the contributions of previous research by improving upon their methods, updating the deep learning models to newer architectures, as well as consolidating them into a single piece of software. RESULTS: The model for extracting the sections of the anesthesia paper chart achieved an average box precision of 0.99, an average box recall of 0.99, and an mAP0.5-95 of 0.97. Our software digitizes checkbox data with greater than 99% accuracy and digitizes blood pressure data with a mean average error of 1.0 and 1.36 mmHg for systolic and diastolic blood pressure respectively. Overall accuracy for physiological data which includes oxygen saturation, inspired oxygen concentration and end tidal carbon dioxide concentration was 85.2%. CONCLUSIONS: We demonstrate that under normal photography conditions we can digitize checkbox, blood pressure and physiological data to within human accuracy when provided legible handwriting. Our contributions provide improved access to digital data to healthcare practitioners in low-middle income countries.
Ryan D. Folks, Bhiken I. Naik, Donald E. Brown, Marcel Durieux
BMC Bioinform.3
2024 Universal representation learning for multivariate time series using the instance-level and cluster-level supervised contrastive learning
abstract
The multivariate time series classification (MTSC) task aims to predict a class label for a given time series. Recently, modern deep learning-based approaches have achieved promising performance over traditional methods for MTSC tasks. The success of these approaches relies on access to the massive amount of labeled data (i.e., annotating or assigning tags to each sample that shows its corresponding category). However, obtaining a massive amount of labeled data is usually very time-consuming and expensive in many real-world applications such as medicine, because it requires domain experts' knowledge to annotate data. Insufficient labeled data prevents these models from learning discriminative features, resulting in poor margins that reduce generalization performance. To address this challenge, we propose a novel approach: supervised contrastive learning for time series classification (SupCon-TSC). This approach improves the classification performance by learning the discriminative low-dimensional representations of multivariate time series, and its end-to-end structure allows for interpretable outcomes. It is based on supervised contrastive (SupCon) loss to learn the inherent structure of multivariate time series. First, two separate augmentation families, including strong and weak augmentation methods, are utilized to generate augmented data for the source and target networks, respectively. Second, we propose the instance-level, and cluster-level SupCon learning approaches to capture contextual information to learn the discriminative and universal representation for multivariate time series datasets. In the instance-level SupCon learning approach, for each given anchor instance that comes from the source network, the low-variance output encodings from the target network are sampled as positive and negative instances based on their labels. However, the cluster-level approach is performed between each instance and cluster centers among batches, as opposed to the instance-level approach. The cluster-level SupCon loss attempts to maximize the similarities between each instance and cluster centers among batches. We tested this novel approach on two small cardiopulmonary exercise testing (CPET) datasets and the real-world UEA Multivariate time series archive. The results of the SupCon-TSC model on CPET datasets indicate its capability to learn more discriminative features than existing approaches in situations where the size of the dataset is small. Moreover, the results on the UEA archive show that training a classifier on top of the universal representation features learned by our proposed method outperforms the state-of-the-art approaches.
Nazanin Moradinasab, Suchetha Sharma, Ronen Bar-Yoseph, Shlomit Radom-Aizik, Kenneth C. Bilchick, Dan M. Cooper, Arthur Weltman, Donald E. Brown
Data Min. Knowl. Discov.8
2023 Global Analysis with Aggregation-based Beaconing Detection across Large Campus Networks
abstract
We present a new approach to effectively detect and prioritize malicious beaconing activities in large campus networks by profiling the server activities through aggregated signals across multiple traffic protocols and networks. Key components of our system include a novel time-series analysis algorithm that uncovers hidden periodicity in aggregated signals, and a ranking-based detection pipeline that utilizes self-training and active-learning techniques. We evaluate our detection system on 10 months of real-world traffic collected at two large campus networks, comprising over 75 billion connections. On a daily average, we detect 43% more periodic domains by aggregating signals across multiple networks compared to single-network analysis. Furthermore, our ranking pipeline successfully identifies 1,387 unique malicious domains, out of which 781 (56%) were unknown to the major online threat intelligence platform, VirusTotal, at the time of our detection.
Yizhe Zhang 0006, Hongying Dong, Alastair Nottingham, Molly Buchanan, Donald E. Brown, Yixin Sun 0004
ACSAC5
2023 A Bayesian Hierarchical Analysis on the Disparity of Emergency Department Visits for COVID-19: A Cohort Study Using National COVID Cohort Collaborative (N3C) Data
abstract
Since the COVID-19 pandemic in 2020, there are numerous studies and researches on the long term effect of COVID-19 on both patient level and social level with disparities noted in infection rates and outcomes. However, differences in the COVID related healthcare decisions after patients present for care (e.g. hospitalization after emergency department visit) has not been studied much at a national level. The National COVID Cohort Collaborative (N3C) provides researchers with abundant data collected from different clinical sites, making it suitable for Bayesian hierarchical modeling while analyzing the disparity in hospitalization after emergency department visit, where prior information or belief could be easily included in the modeling process by adjusting the prior distribution of parameters. In this analysis, we select demographic information (age, sex, race and ethnicity) and the Charlson Comorbidity Index (CCI) as features and study the relationship between these features and whether a patient would be hospitalized after having a COVID related visit to an emergency department (ED).
Johanna Loomba, Andrea Zhou, Suchetha Sharma, Saurav Sengupta, Donald E. Brown
ICMLA6
2023 Determining Risk Factors for Long COVID Using Positive Unlabeled Learning on Electronic Health Records Data from NIH N3C
abstract
Post-acute sequelae of SARS-Co V-2 infection (PASC), also known as Long COVID, is an emerging medical condition in the aftermath of the COVID-19 pandemic. Research on this disease is limited by its newness and the lack of reliable controls, which can hinder model development. The National COVID Cohort Collaborative (N3C)11https://ncats.nih.gov/n3c contains Electronic Health Record (EHR) data for 7 million COVID positive patients from 76 sites across the United States, of which there are fifty thousand Long COVID patients. For this study, we model our risk factor analysis as Positive Unlabeled (PU) problem, where we treat Long COVID patients as the positive sample and rest of the COVID positive patients as unlabeled data. We first curate reliable controls using a PU modeling technique called bagging. We then use this cohort of positive and the curated negative samples to model risk factors for Long COVID. We utilize an attention-based deep learning approach using Long Short Term Memory (LSTM) networks on historical diagnosis data prior to COVID-19 infection, to first predict for Long COVID and then extract the model attention values to score input diagnoses for each patient. Using this process, we achieve an Area Under the Receiver Operating Characteristic (AUROC) of 0.93 (0.88 F1 Score) for the prediction task, significantly outperforming the same model trained on randomly selected controls. We then use a scoring process to rank different input diagnoses for each correctly classified patient with attention values extracted from the trained model and find the temporal distribution of top diagnosis codes which, when represented graphically, becomes a helpful tool to for physicians to investigate diagnosis patterns that effect Long COVID and also evaluate model trustworthiness.
Saurav Sengupta, Johanna Loomba, Suchetha Sharma, Scott A. Chapman, Donald E. Brown
ICMLA5
2022 Vital Measurements of Hospitalized COVID-19 Patients as a Predictor of Long COVID: An EHR-based Cohort Study from the RECOVER Program in N3C
abstract
It is shown that various symptoms could remain in the stage of post-acute sequelae of SARS-CoV-2 infection (PASC), otherwise known as Long COVID. A number of COVID patients suffer from heterogeneous symptoms, which severely impact recovery from the pandemic. While scientists are trying to give an unambiguous definition of Long COVID, efforts in prediction of Long COVID could play an important role in understanding the characteristic of this new disease. Vital measurements (e.g. oxygen saturation, heart rate, blood pressure) could reflect body's most basic functions and are measured regularly during hospitalization, so among patients diagnosed COVID positive and hospitalized, we analyze the vital measurements of first 7 days since the hospitalization start date to study the pattern of the vital measurements and predict Long COVID with the information from vital measurements.
Johanna Loomba, Suchetha Sharma, Donald E. Brown
BIBM4
2022 Analyzing historical diagnosis code data from NIH N3C and RECOVER Programs using deep learning to determine risk factors for Long Covid
abstract
Post-acute sequelae of SARS-CoV-2 infection (PASC) or Long COVID is an emerging medical condition that has been observed in several patients with a positive diagnosis for COVID-19. Historical Electronic Health Records (EHR) like diagnosis codes, lab results and clinical notes have been analyzed using deep learning and have been used to predict future clinical events. In this paper, we propose an interpretable deep learning approach to analyze historical diagnosis code data from the National COVID Cohort Collective (N3C)1to find the risk factors contributing to developing Long COVID. Using our deep learning approach, we are able to predict if a patient is suffering from Long COVID from a temporally ordered list of diagnosis codes up to 45 days post the first COVID positive test or diagnosis for each patient, with an accuracy of 70.48%. We are then able to examine the trained model using Gradient-weighted Class Activation Mapping (GradCAM) to give each input diagnoses a score. The highest scored diagnosis were deemed to be the most important for making the correct prediction for a patient. We also propose a way to summarize these top diagnoses for each patient in our cohort and look at their temporal trends to determine which codes contribute towards a positive Long COVID diagnosis.
Saurav Sengupta, Johanna Loomba, Suchetha Sharma, Donald E. Brown, Lorna E. Thorpe, Melissa A. Haendel, Christopher G. Chute, Stephanie S. Hong
BIBM4
2022 MaNi: Maximizing Mutual Information for Nuclei Cross-Domain Unsupervised Segmentation
Yash Sharma 0002, Sana Syed, Donald E. Brown
MICCAI (2)3
2022 The iTHRIV Commons: a cross-institution information and health research data sharing architecture and web application
abstract
OBJECTIVE: The integrated Translational Health Research Institute of Virginia (iTHRIV) aims to develop an information architecture to support data workflows throughout the research lifecycle for cross-state teams of translational researchers. MATERIALS AND METHODS: The iTHRIV Commons is a cross-state harmonized infrastructure supporting resource discovery, targeted consultations, and research data workflows. As the front end to the iTHRIV Commons, the iTHRIV Research Concierge Portal supports federated login, personalized views, and secure interactions with objects in the ITHRIV Commons federation. The canonical use-case for the iTHRIV Commons involves an authenticated user connected to their respective high-security institutional network, accessing the iTHRIV Research Concierge Portal web application on their browser, and interfacing with multi-component iTHRIV Commons Landing Services installed behind the firewall at each participating institution. RESULTS: The iTHRIV Commons provides a technical framework, including both hardware and software resources located in the cloud and across partner institutions, that establishes standard representation of research objects, and applies local data governance rules to enable access to resources from a variety of stakeholders, both contributing and consuming. DISCUSSION: The launch of the Commons API service at partner sites and the addition of a public view of nonrestricted objects will remove barriers to data access for cross-state research teams while supporting compliance and the secure use of data. CONCLUSIONS: The secure architecture, distributed APIs, and harmonized metadata of the iTHRIV Commons provide a methodology for compliant information and data sharing that can advance research productivity at Hub sites across the CTSA network.
Johanna Loomba, Glenn S. Wasson, Ravi Kiran Reddy Chamakuri, Pabitra Kumar Dash, Stephen G. Patterson, Mary M. A. Potter, Jason Edward Krisch, Martha M. Tenzer, Karen C. Johnston, Donald E. Brown
J. Am. Medical Informatics Assoc.10
2022 Functional Data Analysis for Predicting Pediatric Failure to Complete Ten Brief Exercise Bouts
abstract
Physiological response to physical exercise through analysis of cardiopulmonary measurements has been shown to be predictive of a variety of diseases. Nonetheless, the clinical use of exercise testing remains limited because interpretation of test results requires experience and specialized training. Additionally, until this work no methods have identified which dynamic gas exchange or heart rate responses influence an individual's decision to start or stop physical activity. This research examines the use of advanced machine learning methods to predict completion of a test consisting of multiple exercise bouts by a group of healthy children and adolescents. All participants could complete the ten bouts at low or moderate-intensity work rates, however, when the bout work rates were high-intensity, 50% refused to begin the subsequent exercise bout before all ten bouts had been completed (task failure). We explored machine learning strategies to model the relationship between the physiological time series, the participant's anthropometric variables, and the binary outcome variable indicating whether the participant completed the test. The best performing model, a generalized spectral additive model with functional and scalar covariates, achieved 93.6% classification accuracy and an F1 score of 93.5%. Additionally, functional analysis of variance testing showed that participants in the 'failed' and 'success' groups have significantly different functional means in three signals: heart rate, oxygen uptake rate, and carbon dioxide uptake rate. Overall, these results show the capability of functional data analysis with generalized spectral additive models to identify key differences in the exercise-induced responses of participants in multiple bout exercise testing.
Nick Coronato, Donald E. Brown, Yash Sharma 0002, Ronen Bar-Yoseph, Shlomit Radom-Aizik, Dan M. Cooper
IEEE J. Biomed. Health Informatics2
2022 Using Machine Learning to Identify Organ System Specific Limitations to Exercise via Cardiopulmonary Exercise Testing
abstract
Cardiopulmonary Exer cise Testing (CPET) is a unique physiologic medical test used to evaluate human response to progressive maximal exercise stress. Depending on the degree and type of deviation from the normal physiologic response, CPET can help identify a patient's specific limitations to exercise to guide clinical care without the need for other expensive and invasive diagnostic tests. However, given the amount and complexity of data obtained from CPET, interpretation and visualization of test results is challenging. CPET data currently require dedicated training and significant experience for proper clinician interpretation. To make CPET more accessible to clinicians, we investigated a simplified data interpretation and visualization tool using machine learning algorithms. The visualization shows three types of limitations (cardiac, pulmonary and others); values are defined based on the results of three independent random forest classifiers. To display the models' scores and make them interpretable to the clinicians, an interactive dashboard with the scores and interpretability plots was developed. This machine learning platform has the potential to augment existing diagnostic procedures and provide a tool to make CPET more accessible to clinicians.
Julio J. Portella, Brian J. Andonian, Donald E. Brown, Joao Mansur, Derek Wales, Vivian L. West, William E. Kraus, William Edward Hammond
IEEE J. Biomed. Health Informatics3
2021 Graph Convolutional Neural Network For Weakly Supervised Abnormality Localization In Long Capsule Endoscopy Videos
abstract
Temporal abnormality localization in long Wireless Capsule Endoscopy (WCE) videos is an important problem. The cost of obtaining frame level label for long WCE videos is prohibitive. In this paper, we propose an end-to-end temporal abnormality localization for long WCE videos using only weak video level labels. Physicians use Capsule Endoscopy (CE) as a non-surgical and non-invasive method to examine the entire digestive tract in order to diagnose diseases or abnormalities. While CE has revolutionized traditional endoscopy procedures, a single CE examination could last up to 8 hours generating as much as 100,000 frames. Physicians must review the entire video, frame-by-frame, in order to identify the frames capturing relevant lesion or abnormality. This, sometimes could be as few as just a single frame. Given this very high level of redundancy, analysing long CE videos can be very tedious, time consuming and also error prone. This paper presents a novel multi-step method for an end-to-end localization of target frames capturing abnormalities of interest in the long video using only weak video labels. First we developed an automatic temporal segmentation using change point detection technique to temporally segment the video into uniform, homogeneous and identifiable segments. Then we employed Graph Convolutional Neural Network (GCNN) to learn a representation of each video segment. Using weak video segment labels, we trained our GCNN model to recognize each video segment as abnormal if it contains at least a single abnormal frame. Finally, leveraging the parameters of the trained GCNN model, we replaced the final layer of the network with a temporal pool layer to localize the relevant abnormal frames within each abnormal video segment. We experimented with multiple real patients’ endoscopy videos and achieved an accuracy of 89.9% on the graph classification task and a specificity of 97.5% on the abnormal frames localization task.
Sodiq Adewole, Philip Fernandez, James A. Jablonski, Andrew Copland, Michael D. Porter, Sana Syed, Donald E. Brown
IEEE BigData7
2021 Embeddings of genomic region sets capture rich biological associations in lower dimensions
abstract
MOTIVATION: Genomic region sets summarize functional genomics data and define locations of interest in the genome such as regulatory regions or transcription factor binding sites. The number of publicly available region sets has increased dramatically, leading to challenges in data analysis. RESULTS: We propose a new method to represent genomic region sets as vectors, or embeddings, using an adapted word2vec approach. We compared our approach to two simpler methods based on interval unions or term frequency-inverse document frequency and evaluated the methods in three ways: First, by classifying the cell line, antibody or tissue type of the region set; second, by assessing whether similarity among embeddings can reflect simulated random perturbations of genomic regions; and third, by testing robustness of the proposed representations to different signal thresholds for calling peaks. Our word2vec-based region set embeddings reduce dimensionality from more than a hundred thousand to 100 without significant loss in classification performance. The vector representation could identify cell line, antibody and tissue type with over 90% accuracy. We also found that the vectors could quantitatively summarize simulated random perturbations to region sets and are more robust to subsampling the data derived from different peak calling thresholds. Our evaluations demonstrate that the vectors retain useful biological information in relatively lower-dimensional spaces. We propose that vector representation of region sets is a promising approach for efficient analysis of genomic region data. AVAILABILITY AND IMPLEMENTATION: https://github.com/databio/regionset-embedding. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Erfaneh Gharavi, Aaron Gu, Guangtao Zheng, Jason P. Smith, Hyun Jae Cho, Aidong Zhang 0001, Donald E. Brown, Nathan C. Sheffield
Bioinform.7
2020 Semi-Supervised Classification of Noisy, Gigapixel Histology Images
abstract
One of the greatest obstacles in the adoption of deep neural networks for new medical applications is that training these models typically require a large amount of manually labeled training samples. In this body of work, we investigate the semi-supervised scenario where one has access to large amounts of unlabeled data and only a few labeled samples. We study the performance of MixMatch and FixMatch-two popular semi-supervised learning methods-on a histology dataset. More specifically, we study these models' impact under a highly noisy and imbalanced setting. The findings here motivate the development of semi-supervised methods to ameliorate problems commonly encountered in medical data applications.
J. Vince Pulido, Shan Guleria, Lubaina Ehsan, Matthew Fasullo, Robert Lippman, Pritesh Mutha, Tilak Shah, Sana Syed, Donald E. Brown
BIBE9
2020 A recurrent neural network approach to predicting hemoglobin trajectories in patients with End-Stage Renal Disease
Benjamin J. Lobo, Emaad Abdel-Rahman, Donald E. Brown, Lori Dunn, Brendan Bowman
Artif. Intell. Medicine3
2019 CeliacNet: Celiac Disease Severity Diagnosis on Duodenal Histopathological Images Using Deep Residual Networks
abstract
Celiac Disease (CD) is a chronic autoimmune disease that affects the small intestine in genetically predisposed children and adults. Gluten exposure triggers an inflammatory cascade which leads to compromised intestinal barrier function. If this enteropathy is unrecognized, this can lead to anemia, decreased bone density, and, in longstanding cases, intestinal cancer. The prevalence of the disorder is 1% in the United States. An intestinal (duodenal) biopsy is considered the "gold standard" for diagnosis. The mild CD might go unnoticed due to non-specific clinical symptoms or mild histologic features. In our current work, we trained a model based on deep residual networks to diagnose CD severity using a histological scoring system called the modified Marsh score. The proposed model was evaluated using an independent set of 120 whole slide images from 15 CD patients and achieved an AUC greater than 0.96 in all classes. These results demonstrate the diagnostic power of the proposed model for CD severity classification using histological images.
Rasoul Sali, Lubaina Ehsan, Kamran Kowsari, Marium N. Khan, Christopher A. Moskaluk, Sana Syed, Donald E. Brown
BIBM7
2019 MobiAmbulance: Optimal Scheduling of Emergency Vehicles in Catastrophic Situations
abstract
With recent experience in multiple large-scale disasters, it has been widely confirmed that the severity of a disaster is greatly dependent on the effectiveness of ambulance dispatching during disaster phase. However, previous base station (i.e., temporary or permanent hospital) based ambulance redeployment methods and dynamic ambulance scheduling methods cannot handle the ambulance dispatching problem in catastrophic situations. In this paper, we present MobiAmbulance: a human Mobility based Ambulance dispatching system that aims to maximize the total number of fulfilled patient pick-up requests, and minimize the driving delay of the fulfilled requests. We studied a state-scale human mobility dataset and found that the change of vehicle flow rate can be utilized to determine the connection status between road segments, and the distribution of people in catastrophic situations is drastically different from that in normal situations. Then, we develop a method to determine the road network connection status and the set of road segments that can still be driven through by ambulances after disaster. Based on the updated road network graph, we develop an ambulance dispatching method based on weighted driving route to maximize the total number of fulfilled patient pick-up requests, and minimize the driving delays of the fulfilled requests. Our trace-driven experiments demonstrate the superior performance of MobiAmbulance over other comparison methods.
Li Yan 0004, Shohaib Mahmud, Haiying Shen, Natasha Zhang Foutz, Donald E. Brown, Wie Yusuf, Derek Loftis, Lucas Lyons, Jonathan L. Goodall, Joshua Anton
ICCCN5
2018 Longitudinal Analysis of Linguistic flexibility of Value-motivated Groups
abstract
Increasing globalization of the world leads to an emerging need for ways to analysis and understand groups from different cultures and ideologies. Researchers have used written text as a medium to examine political discourse and analyze value-motivated groups. Previous works showed that computational linguistic analysis can be performed to infer the flexibility of value-motivated groups from their writings. The main premise of these works is that text can bring insights into individuals' and groups' way of thinking, and potentially, behaviour. While existing works provide viable solutions for characterizing groups' ideological behaviour, they perform their analyses over all text published by the groups. However, researchers have found that religious and value-motivated groups can't be analyzed collectively as they regularly evolve. To address this gap, we analyze the performance of existing methods to single documents. Experimental results show that previous features (e.g., use of pronouns and judgment statements) used to predict groups' flexibility are less predictive for single documents' flexibility. We show that a newly added feature regarding the identity of a group provides a significant contribution to the prediction process. Furthermore, due to the unbalanced nature of our data, we propose a weighting scheme for linear regression based on the inter-group variance. Results indicate that a weighted least squares significantly outperforms a traditional least squares approach. This work brings new insights into the characteristics of different linguistic and performative signals, and their relationship to the linguistic flexibility of groups. It also provides a decision making support tool for practical use by practitioners.
Mohammad Al Boni, Seth Green, Megan Stiles, Katherine Harton, Donald E. Brown
IEEE BigData5
2018 Comparing the Performance of an Immersive Virtual Reality and Traditional Desktop Cultural Game
Brian An, Forrest Matteo, Matt Epstein, Donald E. Brown
CHIRA4
2018 Analysis of Railway Accidents' Narratives Using Deep Learning
abstract
Automatic understanding of domain specific texts in order to extract useful relationships for later use is a non-trivial task. One such relationship would be between railroad accidents' causes and their correspondent descriptions in reports. From 2001 to 2016 rail accidents in the U.S. cost more than $4.6B. Railroads involved in accidents are required to submit an accident report to the Federal Railroad Administration (FRA). These reports contain a variety of fixed field entries including primary cause of the accidents (a coded variable with 389 values) as well as a narrative field which is a short text description of the accident. Although these narratives provide more information than a fixed field entry, the terminologies used in these reports are not easy to understand by a non-expert reader. Therefore, providing an assisting method to fill in the primary cause from such domain specific texts (narratives) would help to label the accidents with more accuracy. Another important question for transportation safety is whether the reported accident cause is consistent with narrative description. To address these questions, we applied deep learning methods together with powerful word embeddings such as Word2Vec and GloVe to classify accident cause values for the primary cause field using the text in the narratives. The results show that such approaches can both accurately classify accident causes based on report narratives and find important inconsistencies in accident reporting.
Mojtaba Heidarysafa, Kamran Kowsari, Laura E. Barnes, Donald E. Brown
ICMLA4
2018 Differentially Private Hypothesis Transfer Learning
Quanquan Gu, Donald E. Brown
ECML/PKDD (2)3
2017 Environmental Reservoirs of Nosocomial Infection: Imputation Methods for Linking Clinical and Environmental Microbiological Data to Understand Infection Transmission
Julia Lensing, Ketki Vilankar, Hyojung Kang, Donald E. Brown, Amy Mathers, Laura E. Barnes
AMIA4
2017 HDLTex: Hierarchical Deep Learning for Text Classification
abstract
Increasingly large document collections require improved information processing methods for searching, retrieving, and organizing text. Central to these information processing methods is document classification, which has become an important application for supervised learning. Recently the performance of traditional supervised classifiers has degraded as the number of documents has increased. This is because along with growth in the number of documents has come an increase in the number of categories. This paper approaches this problem differently from current document classification methods that view the problem as multi-class classification. Instead we perform hierarchical classification using an approach we call Hierarchical Deep Learning for Text classification (HDLTex). HDLTex employs stacks of deep learning architectures to provide specialized understanding at each level of the document hierarchy.
Kamran Kowsari, Donald E. Brown, Mojtaba Heidarysafa, Kiana Jafari Meimandi, Matthew S. Gerber, Laura E. Barnes
ICMLA2
2016 Extracting Addresses from News Reports Using Conditional Random Fields
abstract
Spatial analysis in many fields requires effective address extraction from text reports. This problem is of particular importance in social science where news reports contain information about socially relevant incidents. Previous address extraction work focuses on web pages where addresses are separated from other text, however news reports contain addresses embedded in text. Hence, the need for different methods. This paper describes and compares three supervised learning approaches and one semi-supervised learning approach to automatically extract street addresses from news reports. Experimental results with actual news reports show performance close to that achieved for web pages and some lift in accuracy from the semi-supervised approach. These results also show that different news sources produce different outcomes.
Donald E. Brown
ICMLA1
2016 Text Mining the Contributors to Rail Accidents
abstract
Rail accidents represent an important safety concern for the transportation industry in many countries. In the 11 years from 2001 to 2012, the U.S. had more than 40 000 rail accidents that cost more than $45 million. While most of the accidents during this period had very little cost, about 5200 had damages in excess of $141 500. To better understand the contributors to these extreme accidents, the Federal Railroad Administration has required the railroads involved in accidents to submit reports that contain both fixed field entries and narratives that describe the characteristics of the accident. While a number of studies have looked at the fixed fields, none have done an extensive analysis of the narratives. This paper describes the use of text mining with a combination of techniques to automatically discover accident characteristics that can inform a better understanding of the contributors to the accidents. The study evaluates the efficacy of text mining of accident narratives by assessing predictive performance for the costs of extreme accidents. The results show that predictive accuracy for accident costs significantly improves through the use of features found by text mining and predictive accuracy further improves through the use of modern ensemble methods. Importantly, this study also shows through case examples how the findings from text mining of the narratives can improve understanding of the contributors to rail accidents in ways not possible through only fixed field analysis of the accident reports.
Donald E. Brown
IEEE Trans. Intell. Transp. Syst.1
2015 Regular expression acceleration on the micron automata processor: Brill tagging as a case study
abstract
Brill tagging is a classic rule-based algorithm for part-of-speech (POS) tagging that assigns tags, such as nouns, verbs, adjectives, etc., to input tokens. Due to the the intense memory requirements of rule matching, CPU implementations of the Brill tagging algorithm have been found to be slow. We show that Micron's Automata Processor (AP) - a new computing architecture that can perform massively parallel pattern matching - can greatly accelerate the second stage of Brill tagging via rule template matching. The 218 contextual rules are first converted into regular expressions (regex). Regex is used widely in natural language processing (NLP) tasks, thus, this case study involving Brill Tagging also shows how the AP might accelerate other applications that are able to be framed as regexes. We compare single-threaded, and multithreaded versions of Regex matching on an Intel i7 CPU, an Intel XeonPhi co-processor, and the AP. The results show a 63.90X speed-up using the AP as a regex accelerator over the fastest multi-threaded CPU version. We also investigate how performance of regex matching on both CPU architectures varies depending on the complexity of the regex. Taken together, these results demonstrate the potential for significant performance improvements by using accelerators for various NLP computational tasks, particularly those that involve rule-based or pattern-matching approaches.
Keira Zhou, Jack Wadden, Jeffrey J. Fox, Ke Wang 0011, Donald E. Brown, Kevin Skadron
IEEE BigData5
2014 Measures of Entropy and Change Point Analysis as Predictors of Post-Surgical Adverse Outcomes
abstract
A variety of adverse outcomes, such as kidney injury, death, cardiac injury, and respiratory failure affect a significant number of patients after surgery. Previous research has investigated possible predictors for these outcomes including features extracted from physiologic time series. This study builds upon this previous work by exploring entropy, long-term memory, and change point analysis as different and possibly predictive measures of volatility. To do this, we use both random forest models and the robust method of L1 regularized logistic regression as modeling frameworks for the prediction. Predictive results from these models are evaluated using receiver operating characteristic (ROC) curves and their area under the curve (AUC) values. While the developed models did not show improvements in predictive accuracy, they did show that change point analysis and measures of entropy and long-term memory can be useful tools in predicting postsurgical adverse outcomes.
Zachary Terner, Timothy Carroll, Donald E. Brown
BIBE3
2013 Random Forests on Ubiquitous Data for Heart Failure 30-Day Readmissions Prediction
abstract
Heart failure is the most common reason for unplanned hospital readmissions. Typical 30 day readmission prediction models either use data that are not readily available at the majority of US hospitals or use modeling techniques that do not provide adequate prediction accuracy. Moreover, the tendency of ongoing studies is to incorporate clinical data that is only present in the most modern electronic health record systems (EHRs). This is problematic as the population most affected by heart disease, the rural poor, is also the same population whose hospitals have the slowest adoption rates of advanced EHR systems. We apply the machine learning technique random forests to administrative claims data to predict unplanned all-cause 30 day readmissions for congestive heart failure patients in a hospital system located in central Virginia, USA. We form two random forests model variants based on datasets comprised of procedure data, diagnosis data, a combination of both, and basic demographic data. Our results show significant predictive performance, yield importance rankings for candidate variables, and address heart failure readmissions in high-need areas.
Michael A. Vedomske, Donald E. Brown, James H. Harrison
ICMLA (2)2
2013 Scalable and Locally Applicable Measures of Treatment Variation That Use Hospital Billing Data
abstract
Care variation studies often use large amounts of data but approaches developed for such research are either scalable but not locally applicable or locally applicable but not scalable. We present a method that is scalable and locally applicable while being statistically significant. Using a population of patients diagnosed with both congestive heart failure and myocardial infarction, we developed and tested measures of care variation on data derived from hospital billing records. Our metrics yielded statistically significant results. Computing time for the method was found to increase linearly allowing for the desired scalability. In the future, our care variation metrics be used to gain insight into local conditions that correlate with outcomes of interest like visit charges or morbidity rates.
Michael A. Vedomske, Matthew S. Gerber, Donald E. Brown, James H. Harrison
ICMLA (2)3
2012 Spatio-temporal modeling of criminal incidents using geographic, demographic, and twitter-derived information
abstract
Personal and property crimes create large economic losses within the United States. To prevent crimes, law enforcement agencies model the spatio-temporal pattern of criminal incidents. In this paper, we present a new modeling process that combines two of our recently developed approaches for modeling criminal incidents. The first component of the process is the spatio-temporal generalized additive model (STGAM), which predicts the probability of criminal activity at a given location and time using a feature-based approach. The second component involves textual analysis. In our experiments, we automatically analyzed Twitter posts, which provide a rich, event-based context for criminal incidents. In addition, we describe a new feature selection method to identify important features. We applied our new model to actual criminal incidents in Charlottesville, Virginia. Our results indicate that the STGAM/Twitter model outperforms our previous STGAM model, which did not use Twitter information. The STGAM/Twitter model can be generalized to other applications of event modeling where unstructured text is available.
Donald E. Brown, Matthew S. Gerber
ISI2
2012 Police patrol district design using agent-based simulation and GIS
abstract
Police patrols play an important role in public safety. The patrol district design is an important factor affecting the patrol performances, such as average response time and workload variation. The redistricting procedure can be described as partitioning smaller geographical units into several larger districts with the constraints of contiguity and compactness. The size of the possible sample space is large and the corresponding graph-partitioning problem is NP-complete. In our approach, the patrol districting plans generated by a parameterized redistricting procedure are evaluated using an agent-based simulation model we implemented in Java Repast in a geographic information system (GIS) environment. The relationship between districting parameters and response variables is studied and better districting plans can be generated. After in-depth evaluations of these plans, we perform a Pareto analysis of the outputs from the simulation to find the non-dominated set of plans on each of the objectives. This paper also includes a case study for the police department of Charlottesville, VA, USA. Simulation results show that patrol performance can be improved compared with the current districting solution.
Donald E. Brown
ISI2
2011 The spatio-temporal generalized additive model for criminal incidents
abstract
Law enforcement agencies need to model spatio-temporal patterns of criminal incidents. With well developed models, they can study the causality of crimes and predict future criminal incidents, and they can use the results to help prevent crimes. In this paper, we described our newly developed spatio-temporal generalized additive model (S-T GAM) to discover underlying factors related to crimes and predict future incidents. The model can fully utilize many different types of data, such as spatial, temporal, geographic, and demographic data, to make predictions. We efficiently estimated the parameters for S-T GAM using iteratively re-weighted least squares and maximum likelihood and the resulting estimates provided for model interpretability. In this paper we showed the evaluation of S-T GAM with the actual criminal incident data from Charlottesville, Virginia. The evaluation results showed that S-T GAM outperformed the previous spatial prediction models in predicting future criminal incidents.
Donald E. Brown
ISI2
2011 Future trends in business analytics and optimization
abstract
During the last decades, the disciplines of Data Mining and Operations Research have been working mostly independent of each other. However, the increasing complexity of today's applications in areas such as business, medicine, and science requires m
Donald E. Brown, Fazel Famili, Gerhard Paass, Kate Smith-Miles, Lyn C. Thomas, Richard Weber 0002, Ricardo Baeza-Yates, Cristián Bravo, Gaston L'Huillier, Sebastián Maldonado 0001
Intell. Data Anal.1
2009 A Statistical Threat Assessment
abstract
Criminal gangs, insurgent groups, and terror networks demonstrate observable preferences in selecting the sites where they commit their crimes. Accordingly, police departments, military organizations, and intelligence agencies seek to learn these preferences and identify locations with a high probability of experiencing the particular event of interest in the near future. Often, such agencies are keen not just to predict the spatial pattern of future events but even more importantly to conduct threat assessments of particular criminal gangs or insurgent groups. These threat assessments include identifying where each of the various groups presents the greatest threat to the community, what the most likely targets are for each criminal group, what makes one location more likely to experience an attack than another, and how to most efficiently allocate resources to address the specific threats to the community. Previous research has demonstrated that applying multivariate prediction models to relate features in an area to the occurrence of crimes offers an improvement in predictive performance over traditional methods of hot-spot analysis. This paper introduces the application of multilevel modeling to these multivariate spatial choice models, demonstrating that it is possible to significantly improve the predictive performance of the spatial choice models for individual groups and leverage that information to provide improved threat assessments of the criminal elements in a given geographic area.
Samuel H. Huddleston, Donald E. Brown
IEEE Trans. Syst. Man Cybern. Part A2
2007 Global Optimization With Multivariate Adaptive Regression Splines
abstract
This paper presents a novel procedure for approximating the global optimum in structural design by combining multivariate adaptive regression splines (MARS) with a response surface methodology (RSM). MARS is a flexible regression technique that uses a modified recursive partitioning strategy to simplify high-dimensional problems into smaller yet highly accurate models. Combining MARS and RSM improves the conventional RSM by addressing highly nonlinear high-dimensional problems that can be simplified into lower dimensions, yet maintains a low computational cost and better interpretability when compared to neural networks and generalized additive models. MARS/RSM is also compared to simulated annealing and genetic algorithms in terms of computational efficiency and accuracy. The MARS/RSM procedure is applied to a set of low-dimensional test functions to demonstrate its convergence and limiting properties.
Scott T. Crino, Donald E. Brown
IEEE Trans. Syst. Man Cybern. Part B2
2006 An outlier-based data association method for linking criminal incidents
Donald E. Brown
Decis. Support Syst.2
2006 Spatial analysis with preference specification of latent decision makers for criminal event prediction
Yifei Xue, Donald E. Brown
Decis. Support Syst.2
2005 Assessing casualty densities based on sensor reports pursuant to a large-scale disaster
abstract
One of the newest innovations which are making its way more prevalently into the field of emergency response is information technology. Information technology (IT), in this sense, seeks to turn relevant data into usable information to aid in an emergency response. One of the key elements to useful beneficial IT is to quickly, accurately, and dynamically turn incoming data into usable information. This paper presents a way to statistically analyze incoming casualty reports at specific time intervals to not only estimate casualty densities, but also assess whether or not the casualty densities being observed are within some confidence interval of an expected number of casualties. Simple models of the searching process are developed and used to dynamically analyze an incoming report stream. If the numbers of casualties are sufficiently different than the expected number, then one might conclude either a secondary event has occurred or the initial estimates were simply wrong.
C. Donald Robinson, Donald E. Brown
SMC2
2005 Health-status monitoring through analysis of behavioral patterns
abstract
With the rapid growth of the elderly population, there is a need to support the ability of elders to maintain an independent and healthy lifestyle in their homes rather than through more expensive and isolated care facilities. One approach to accomplish these objectives employs the concepts of ambient intelligence to remotely monitor an elder's activities and condition. The SmartHouse project uses a system of basic sensors to monitor a person's in-home activity; a prototype of the system is being tested within a subject's home. We examined whether the system could be used to detect behavioral patterns and report the results in this paper. Mixture models were used to develop a probabilistic model of behavioral patterns. The results of the mixture-model analysis were then evaluated by using a log of events kept by the occupant.
T. S. Barger, Donald E. Brown, Majd Alwan
IEEE Trans. Syst. Man Cybern. Part A2
2004 Spatial Forecast Methods for Terrorist Events in Urban Environments
Donald E. Brown, Jason Dalton, Heidi Hoyle
ISI1
2004 A new point process transition density model for space-time event prediction
abstract
A new point process transition density model is proposed based on the theory of point patterns for predicting the likelihood of occurrence of spatial-temporal random events. The model provides a framework for discovering and incorporating event initiation preferences in terms of clusters of feature values. Components of the proposed model are specified taking into account additional behavioral assumptions such as the "journey to event" and "lingering period to resume act." Various feature selection techniques are presented in conjunction with the proposed model. Extending knowledge discovery into feature space allows for extrapolation beyond spatial or temporal continuity and is shown to be a major advantage of our model over traditional approaches. We examine the proposed model primarily in the context of predicting criminal events in space and time.
Donald E. Brown
IEEE Trans. Syst. Man Cybern. Part C2
2003 Criminal Incident Data Association Using the OLAP Technology
Donald E. Brown
ISI2
2003 Decision Based Spatial Analysis of Crime
Yifei Xue, Donald E. Brown
ISI2
2003 An Outlier-based Data Association Method for Linking Criminal Incidents
abstract
Data association is an important data mining task and it has various applications. In crime analysis, data association means linking criminal incidents committed by the same person. It helps to discover crime patterns and catch the criminals. In this paper, we present an outlier-based data association method. An outlier score function is defined to measure the extremeness of an observation, and a data association method is developed based upon the outlier score function. We applied this method to the robbery data in Richmond, Virginia, and compared the result with a similarity-based association method. The results show that the outlier-based data association method is promising.
Donald E. Brown
SDM2
2003 Geographic profiling with event prediction
abstract
Studies have shown that the target preference of a serial criminal is dependent upon the distance he or she must travel from their residence to the target. Further research has identified this relationship as the journey to crime theory. This theory states that a criminal's propensity to commit crime decreases exponential with increasing distance from their home. This paper combines the journey to crime theory along with other geographic profiling methodologies with spatial crime forecasting methodologies to produce a unique geographic profiling methodology. In this methodology, a density surface representing the likelihood of future crime occurrence at each point in the sample space is calculated using locations and location features of previous events of a serial offender. The forecasted density surface is then used to simulate a complete crime surface, where every point in the sample space is considered a crime point of varying degree. The method then models a residence likelihood surface using a density function that accounts for distance to every point in the sample space and the forecasted density score of each of these future crimes.
Justin K. Stile, Donald E. Brown
SMC2
2003 Data association methods with applications to law enforcement
Donald E. Brown, Stephen Hagen
Decis. Support Syst.1
2003 A decision model for spatial site selection by criminals: a foundation for law enforcement decision support
abstract
Crime analysis uses past crime data to predict future crime locations and times. Typically this analysis relies on hot spot models that show clusters of criminal events based on past locations of these events. It does not consider the decision making processes of criminals as human initiated events susceptible to analysis using spatial choice models. This paper analyzes criminal incidents as spatial choice processes. Spatial choice analysis can be used to discover the distribution of people's behaviors in space and time. Two adjusted spatial choice models that include models of decision making processes are presented. The comparison results show that adjusted spatial choice models provide efficient and accurate predictions of future crime patterns and can be used as the basis for a law enforcement decision support system. This paper also extends spatial choice modeling to include the class of problems where the decision makers' preferences are derived indirectly through incident reports rather than directly through survey instruments.
Yifei Xue, Donald E. Brown
IEEE Trans. Syst. Man Cybern. Part C2
2002 Design Approach for Integration of Demographic, Biologic, and Clinical Data
Jason Dalton, Steve Tropello, Sandra Pelletier, Jason A. Lyman, Bruce P. Dembling, Kenneth W. Scully, William A. Knaus, Donald E. Brown
AMIA8
2002 Mining human failure dynamics from accident data using logistic regression and decision trees
abstract
The effective operation of technology depends on the decision-making of the humans operating that technology. Of fundamental interest are the conditions that may lead to failure or accidents. The research to understand human decision-making processes that lead to failure under varying conditions has typically approached the problem either deductively or inductively through surveys or small-scale experiments. This paper describes an inductive approach based on mining multiple accident data sets for relationships between environmental factors, human factors, operational stimuli, and the probability of correct response by the human operators. We describe data mining techniques we have developed for this problem and then show their applicability to train accident data.
Donald E. Brown, Justin K. Stile, Louise F. Gunderson, Ted C. Giras
SMC1
2001 Mining Preferences from Spatial-Temporal Data
abstract
The discovery of preferences in space and time is important in a variety of applications. In this paper we first establish the correspondence between a set of preferences in space and time and density estimates obtained from observations of spatial-temporal features recorded within large databases. We perform density estimation using both kernel methods and mixture models. The density estimates constitute a probabilistic representation of preferences. We then present a point process transition density model for space-time event prediction that hinges upon the density estimates from the preference discovery process. The added dimension of preference discovery through feature space analysis enables our model to outperform traditional preference modeling approaches. We demonstrate this performance improvement using a criminal incident database from Richmond, Virginia. Criminal incidents are human-initiated events that may be governed by criminal preferences over space and time. We applied our modeling technique to breaking and entering crimes committed in both residential and commercial settings. Our approach effectively recovers the preference structure of the criminals and enables one-week ahead forecasts of threatened areas. This capability to accommodate all measurable features, identify the key features, and quantify their relationship with event occurrence over space and time makes this approach applicable to domains other than law enforcement.
Donald E. Brown, Yifei Xue
SDM1
2001 Data mining time series with applications to crime analysis
abstract
This paper is a study of methods of predicting the number of breaking and enterings (B&Es) in subcity regions of Richmond, Virginia. In this study, predictions are made for B&Es in each of four precincts as well as in regions measuring approximately 0.64 square miles. These predictions can be helpful to police efforts by helping them more effectively allocate resources. The paper includes investigation into the distribution of incidents of breaking and entering, which concludes that B&Es are not Poisson distributed. Furthermore, in the analysis of the data, incidents of B&Es also do not show evidence of seasonal patterns. The research investigates factors that many believe are related to crime, such as unemployment rates, previous incidents of crimes, and alcohol sales.
Donald E. Brown, Rosemary B. Oxford
SMC1
2001 Using cluster specific salience weighting to determine the preferences of agents for multi-agent simulations
abstract
One of the methods used to simulate human behavior is the multi-agent model. However, without survey data or the opinion of experts, the preferences of agents in these models are difficult to discover. In the domain of crime, in specific burglary, this information is not. available or is prohibitively expensive to obtain. This paper presents a new methodology for the discovery of these preferences. This methodology uses sequential clustering to discover both the preferences of the agents and the salience of these features to the agent. This result is compared with existing clustering methodologies and is demonstrated to be superior. Finally, the model in which these agents will be used is described.
Louise F. Gunderson, Donald E. Brown
SMC2
2001 Using clustering to discover the preferences of computer criminals
abstract
The ability to predict computer crimes has become increasingly important. The paper describes a method for discovering the preferences of computer criminals. This method involves sequential clustering based on the variance of clusters discovered in higher order clustering. These discovered preferences can be used for the direct protection of computer systems against ongoing attacks or for the construction of simulations of future attacks.
Donald E. Brown, Louise F. Gunderson
IEEE Trans. Syst. Man Cybern. Part A1
2000 Intelligent decision support systems
abstract
We examine characteristics common to successful intelligent decision support systems. In doing this, we attempt to bridge the gap between disparate communities engaged in building various parts of these systems. Three systems were examined in detail from widely different applications and more than 20 additional systems were considered at a lower level of detail. By examining deployed decision support systems within the context of a broad framework we hope to capture the characteristics that can guide future development efforts. We see this as a first step in developing an in-depth compendium that will help bridge the gap between important yet typically isolated fields.
Stephanie A. Guerlain, Donald E. Brown, Christina Mastrangelo
SMC2
2000 Using a multi-agent model to predict both physical and cyber criminal activity
abstract
The paper describes a multi-agent methodology for the prediction of physical crime and cyber crime. The model uses clustering algorithms to determine the number of agents in the environment. The preference of each of these agents is determined by feature selection. The agents are allowed to interact in a synthetic environment. The results of the interactions are measured and the model is updated with new information. This modeling approach holds significant promise for the simulation of human criminal behavior.
Louise F. Gunderson, Donald E. Brown
SMC2
2000 A comparison of systems engineering programs in the United States
abstract
Given the length of time systems engineering has been taught in the US, it is now appropriate to examine the types of programs being offered and to compare and contrast these programs. The paper provides this comparison with a view toward the future of systems engineering education in the US. In particular, we first examine the US undergraduate and graduate programs in systems engineering in order to understand what is taught and how it is taught. Using cluster analysis, we identify four distinct types of systems engineering undergraduate programs, and an informal analysis examines the directions in the systems engineering graduate programs. Next we look at issues in systems engineering education, which have shaped the development of the curricula over the last thirty years. These include the definition of systems engineering, associated professional societies, similar degree types, the role of an undergraduate systems engineering degree, and the role of information technology in systems engineering. We conclude with opportunities for systems engineering education within the US with regard to curricula directions and job opportunities.
Donald E. Brown, William T. Scherer
IEEE Trans. Syst. Man Cybern. Part C1
1998 The Regional Crime Analysis Program (ReCAP): a framework for mining data to catch criminals
abstract
Most law enforcement agencies today are faced with enormous quantities of data that must be processed and turned into useful information. Two technologies provide the means to turn data into information: data fusion and data mining. Data fusion organizes, combines and interprets information from multiple sources, and it overcomes confusion from conflicting reports and cluttered or noisy backgrounds. Data mining is concerned with the automatic discovery of patterns and relationships in large databases. This paper describes ReCAP (Regional Crime Analysis Program), which was built to provide crime analysts with both technologies.
Donald E. Brown
SMC1
1998 Spatial-temporal event prediction: a new model
abstract
A new model for predicting the probability of occurrence of spatial-temporal random events is proposed based on the theory of point patterns. The model allows for the incorporation of observed event characteristics or features into space-time prediction. Effective and efficient prediction also involves identifying the key characteristics or features that explain the spatial pattern as a preliminary step. A clustering-based criterion is presented to address the feature subset selection problem.
Donald E. Brown
SMC2
1998 Tracking multiple objects in terrain
abstract
The digitized battlefield of the 21st Century will revolutionize the methods used to maintain military command and control. The tremendous amount of data available will necessitate the use of intelligent automated systems that augment, and in some cases replace, the human structures currently in place. One aspect of such systems is terrain-based tracking. We discuss an intelligent terrain-based system for tracking multiple vehicles moving across terrain. Specifically, our system extracts and utilizes knowledge about groups to improve the performance of a discrete state-space motion model. Parallel programming techniques are utilized to compute probability densities for the vehicles. A learning component allows for real-time adjustment based on performance.
Edward Sobiesk, John A. Hamilton Jr., John A. Marin, Donald E. Brown, Maria L. Gini
SMC4
1998 Static data association with a terrain-based prior density
abstract
We consider the problem of estimating the states of a static set of targets, given a collection of densities, each representing the state of a single target. We assume there is no a priori knowledge of which of the given densities represent common targets, but that a prior density for the target locations is available. For a two-dimensional (2-D) location estimation problem, we construct a prior density model based on known features of the terrain. We then give a simple Gaussian association-estimation algorithm using the prior density and present some simulation results. We briefly discuss extensions to nonstatic models.
Allen L. Barker, Donald E. Brown, Worthy N. Martin
IEEE Trans. Syst. Man Cybern. Part C2
1997 A Classification Approach to Boolean Query Reformulation
abstract
One of the difficulties in using current Boolean-based information retrieval systems is that it is hard for a user, especially a novice, to formulate an effective Boolean query. Query reformulation can be even more difficult and complex than formulation since users often have difficulty incorporating the new information gained from the previous search into the next query. In this article, query reformulation is viewed as a classification problem, that is, classifying documents as either relevant or nonrelevant. A new reformulation algorithm is proposed which builds a tree-structured classifier, called a query tree, at each reformulation from a set of feedback documents retrieved from the previous search. The query tree can easily be transformed into a Boolean query. The query tree is compared to two query reformulation algorithms on benchmark test sets (CACM, CISI, and Medlars). In most experiments, the query tree showed significant improvements in precision over the two algorithms compared in this study. We attribute this improved performance to the ability of the query tree algorithm to select good search terms and to represent the relationships among search terms into a tree structure. © 1997 John Wiley & Sons, Inc.
James C. French, Donald E. Brown, Nam-Ho Kim
J. Am. Soc. Inf. Sci.2
1996 Classification trees with optimal multivariate decision nodes
Donald E. Brown, Clarence Louis Pittard, Han Park
Pattern Recognit. Lett.1
1995 Reliability Estimation During Prototyping of Knowledge-Based Systems
abstract
Many knowledge based systems are designed and built with little attention paid to the reliability of the output. In this paper, we present an approach, using partitioning of both the knowledge base and the input space, that allows for the measurement of the reliability during any program increment in a rapid prototyping development cycle. Before presenting the approach, we formalize the problem using concepts from general systems theory and then describe our three objectives: 1) measurement of the reliability of the knowledge-based system at the current program increment, 2) prediction of the reliability of the future system, and 3) support for design decisions. Finally, we apply our approach to a design-aiding knowledge-based system for the selection of materials under various climatic conditions. The design-aiding knowledge-based system is used by U.S. Army personnel in the development of equipment to be used by the U.S. Army in various regions of the world. We find that the current system, containing 40 rules, has a reliability of approximately 0.85. However, more importantly, we have discovered the rules that led to many of the failures.>
Donald E. Brown, James J. Pomykalski
IEEE Trans. Knowl. Data Eng.1
1994 A polynomial network for predicting temperature distributions
abstract
Complete temperature distributions are unavailable for many locations throughout the world. This distributional information is important for product design and operational planning. The problem of obtaining these temperature distributions is quite difficult and current techniques are limited in accuracy. This paper describes a new and effective approach to this problem that matches data-deficient locations to maximally similar locations with known distributions. A polynomial network is used to predict which of a set of sites with known distributions is most similar to a data-deficient site. Then an information theoretic criterion is optimized to find the unknown distribution that closely matches this maximally similar site. Tests with this approach demonstrate its effectiveness and its superiority to current methods.
George E. Fulcher, Donald E. Brown
IEEE Trans. Neural Networks2
1993 A comparison of decision tree classifiers with backpropagation neural networks for multimodal classification problems
Donald E. Brown, Vincent Corruble, Clarence Louis Pittard
Pattern Recognit.1
1992 Uncertainty Management with Imprecise Knowledge with Application to Design
Donald E. Brown, Wendy J. Markert
J. Autom. Reason.1
1992 A practical application of simulated annealing to clustering
Donald E. Brown, Christopher L. Huntley
Pattern Recognit.1
1992 Fast generic selection of features for neural network classifiers
abstract
The authors describe experiments using a genetic algorithm for feature selection in the context of neural network classifiers, specifically, counterpropagation networks. They present the novel techniques used in the application of genetic algorithms. First, the genetic algorithm is configured to use an approximate evaluation in order to reduce significantly the computation required. In particular, though the desired classifiers are counterpropagation networks, they use a nearest-neighbor classifier to evaluate features sets and show that the features selected by this method are effective in the context of counterpropagation networks. Second, a method called the training set sampling in which only a portion of the training set is used on any given evaluation, is proposed. Computational savings can be made using this method, i.e., evaluations can be made over an order of magnitude faster. This method selects feature sets that are as good as and occasionally better for counterpropagation than those chosen by an evaluation that uses the entire training set.
Frank Z. Brill, Donald E. Brown, Worthy N. Martin
IEEE Trans. Neural Networks2
1991 Clustering of homogeneous subsets
Donald E. Brown, Christopher L. Huntley, Paul J. Garvey
Pattern Recognit. Lett.1
1990 A Neural Network Implementation of a Data Association Algorithm
abstract
In this paper we are concerned with a time varying set of entities located in a fixed field. These entities are sensed at discrete time instances with a single sensor modality. At a given time instant a collection of bivariate Gaussian sensor reports is produced, estimating the locations of a subset of the entities present in the field. A database of reports is maintained which should ideally contain exactly one report for each entity that has been sensed. Whenever a collection of sensor reports is received the database must be updated to reflect the new information. This updating requires correspondence processing between the database reports and the new sensor reports to determine which pairs of sensor and database reports correspond to the same entity. We present an algorithm for performing this correspondence processing under the assumptions that each new collection of sensor reports contains at most one report for any entity, and that the database can be reasonably assumed to contain at most one report for any entity. This algorithm is based on pairwise distance measures between Gaussian distributions. We consider several such distance measures, and present simulation results indicating that the divergence is a reasonable choice. A formulation of the divergence between two bivariate Gaussians as a scalar product is given. We describe a neural network implementation of our algorithm, along with a proof that the network will converge and, under certain restrictions, exactly compute our correspondence algorithm. INFORMS Journal on Computing, ISSN 1091-9856, was published as ORSA Journal on Computing from 1989 to 1995 under ISSN 0899-1499.
Allen L. Barker, Donald E. Brown, Worthy N. Martin
INFORMS J. Comput.2
1989 Knowledge-based computer-aided design of materials handling systems
abstract
A knowledge-based system called MAHDE (materials handing design) is described that was devised to assist facility designers in the selection and configuration of materials handling equipment. The system implements techniques for representing and acquiring the preferences of the facility designer. This preferential knowledge is incorporated into a knowledge-based system utilizing heuristics to update the operational knowledge about the individual design. The system utilizes preference-directed search to capture improved designs by dynamically acquiring preferences throughout the design process. An example of the application of the system to a design problem is presented.>
Paula Gabbert, Donald E. Brown
IEEE Trans. Syst. Man Cybern.2
1987 An information theoretic interface for a stochastic model management system
Donald E. Brown
Decis. Support Syst.1
1987 A Decision Support System for Reliable Software Development
abstract
Reliability has become a major concern for large-scale software developers. This concern has resulted in the development of a large number of models intended to describe the failure process of software. The use of these models during development has been restricted to the later stages of integration testing. A decision support system is described here which allows for the use of existing software reliability models throughout development. The system uses information from the developer in the form of constraints on the probability distributions which describe the uncertainty associated with the parameters of the various models. Model management is data driven to allow for maximum flexibility in probability assessments which can be provided by either the user or the system.
Donald E. Brown
IEEE Trans. Syst. Man Cybern.1
1987 An Expert System Approach to Boiler Design
abstract
A process model and preliminary specifications are described for a knowledge-based system that incorporates user preference in supporting an important engineering design problem. The specific design problem of interest is the development of minimum (or near minimum) cost fossil fuel boilers. The current process of design involves multiple experts who interact to produce a design meeting customer specifications at the lowest possible cost. This iterative procedure is both tedious and frequently ineffective. Our approach integrates concepts from rule-based expert systems and the decision and cognitive sciences to produce an interactive knowledge-based system to support the design process. The central feature of this system is the inclusion of multiple sources of expertise in combination with formal reasoning and user preference to create the required design.
Donald E. Brown, Chelsea C. White III
IEEE Trans. Syst. Man Cybern.1
1986 Conflicting information integration for decision support
Donald E. Brown, Bernard G. Duren
Decis. Support Syst.1
1975 A High Throughput Packet-Switched Network Technique Without Message Reassembly
abstract
A new packet-switching network technique is described which, while utilizing certain aspects of the ARPANET technology, introduces a substantially different technique for handling traffic which is longer than a single packet in length. The technique is keyed to a common-user network environment, where a wide variety of subscriber types, ranging from computers to simple terminals, are to be serviced. Subscribers in most cases would be remotely located from the network switching nodes. By splitting the buffering between the originating and destination nodes and by essentially eliminating the segment reassembly process, substantial reductions in on-line buffering can be achieved, while still maintaining short response times for interactive messages and large bandwidths for long data exchanges. In this paper we describe the network operational concepts and traffic flow for various subscriber types, show specific examples and timing diagrams for message flows, and present a comparative analysis of the buffer sizing, throughput, and delay for this new technique compared to the well-known ARPANET technique of packet switching.
Roy Daniel Rosner, Raymond H. Bittel, Donald E. Brown
IEEE Trans. Commun.3