VLDB 2026 Research / reviewers in the wild / expert
Carla E. Brodley
dblp:24/1941
· DBLP profile ↗
78ranked-venue papers
9as first author
10since 2021 · last 2026
0009-0008-2134-6285ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 6 first-author · 1 since 2021Databases, data management, data science and information retrieval · 20 · 1 first-authorHuman-computer interaction and ubiquitous computing · 11 · 3 first-author · 9 since 2021Security and privacy · 8Applied, interdisciplinary, general and emerging computing · 8 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4Systems, architecture and hardware · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Can We Have Both? Achieving ABET Compliance with Reduced Curricular Complexity in Computer Science ProgramsabstractThe structure of a computer science (CS) degree program can create barriers that prevent students from discovering or completing the major and limit broadening participation in the field. This study examines whether ABET accreditation requirements influence the degree structure in undergraduate CS programs by analyzing 34 ABET-accredited programs and 39 non-accredited programs. Using the lens of curricular complexity, we found that ABET-accredited programs have higher curricular complexity than non-accredited programs (mean = 189.6 versus. 129.0 respectively, Cohen's d = 1.6). This result is consistent across three home college categories (Computing, Arts & Sciences, and Engineering), indicating that the increased complexity from ABET is independent of home college. A deeper analysis of degree structure reveals that the increased complexity can be attributed to the number and placement of math and CS courses and the path lengths of the prerequisite chains. Understanding these structural influences can help institutions design CS programs that maintain ABET accreditation while reducing barriers to student success. All students should be able to succeed in CS regardless of the ABET accreditation status of the program. Stefanie Colino, Albert Lionelle, Estelita Chen, Carla E. Brodley |
ITiCSE (1) | 4 |
| 2026 | Creating a Second Pathway to the Computing MajorabstractIn 2021, the Computer Science and Engineering department at UC Riverside added a second pathway to the major with a new CS1 and CS2 course sequence. Our goal was to address disparities in course outcomes between different populations, particularly between those with and without prior coding experience and majors versus non-majors. The hope was that students new to computing would have a viable path to discover whether they had interest and aptitude in computing without the added stress of being in a classroom with students who have prior coding experience. In this paper, we report the data analysis that led to our decision to add a second pathway, design choices, and challenges (particularly in university politics) in launching and adding a second pathway. We present the results over the last three years, which illustrate that the new series has leveled the playing field. Most importantly, the course outcome data and changes to computing major enrollment illustrate that we achieved our goal of attracting students new to computing and thus are broadening participation in computing. Ashley Pang, Paea LePendu, Mariam Salloum, Neftali Watkinson Medina, Carla E. Brodley |
SIGCSE (1) | 5 |
| 2025 | NNsight and NDIF: Democratizing Access to Open-Weight Foundation Model InternalsabstractWe introduce NNsight and NDIF, technologies that work in tandem to enable scientific study of the representations and computations learned by very large neural networks. NNsight is an open-source system that extends PyTorch to introduce deferred remote execution. The National Deep Inference Fabric (NDIF) is a scalable inference service that executes NNsight requests, allowing users to share GPU resources and pretrained models. These technologies are enabled by the Intervention Graph, an architecture developed to decouple experimental design from model runtime. Together, this framework provides transparent and efficient access to the internals of deep neural networks such as very large language models (LLMs) without imposing the cost or complexity of hosting customized models individually. We conduct a quantitative survey of the machine learning literature that reveals a growing gap in the study of the internals of large-scale AI. We demonstrate the design and use of our framework to address this gap by enabling a range of research methods on huge models. Finally, we conduct benchmarks to compare performance with previous approaches.
Code, documentation, and tutorials are available at https://nnsight.net/. Jaden Fiotto-Kaufman, Alexander R. Loftus, Eric Todd, Jannik Brinkmann, Koyena Pal, Dmitrii Troitskii, Michael Ripa, Adam Belfki, Can Rager, Caden Juang, Aaron Mueller, Samuel Marks, Arnab Sen Sharma, Francesca Lucchetti, Nikhil Prakash, Carla E. Brodley, Arjun Guha, Jonathan Bell 0001, Byron C. Wallace, David Bau |
ICLR | 16 |
| 2025 | Does Reducing Curricular Complexity Impact Student Success in Computer Science?abstractComputer science degree requirements often have a rigid pre- and corequisite structure, which can impede a student's progression through a degree and in particular can add one or more semesters to time to completion, particularly for those students who need to retake a course that serves as a prerequisite to other courses and for students who are not calculus-ready when they enter university. In this paper, we present the results of a comparative analysis of curricula before and after a major structural revision. The first curriculum adheres to the conventional, rigid prerequisite structure, while the second emphasizes student choice and multiple pathways through the degree. No changes were made to course content/outcomes between the two versions. Employing curricular metrics such as complexity and centrality, we examine the degree progress of 3010 students over a six-year period. Specifically, our investigation looks at the impact of reducing curricular complexity on student attrition from and attraction to the CS major. The new curriculum, with a 60% reduction in curricular structural complexity, showed both increased retention of students over the old curriculum (67% to 98%) and an increase in the number of students converting from undeclared to computer science (44% to 69%). Our findings demonstrate that reducing curricular complexity need not compromise program rigor and can benefit students by providing greater flexibility and ensuring earlier exposure to (and therefore retention in) CS. Sumukhi Ganesan, Albert Lionelle, Catherine Gill, Carla E. Brodley |
SIGCSE (1) | 4 |
| 2025 | Student Application Trends for Teaching Assistant Positions
Felix Muzny, Abdulaziz Arif Suria, Carla E. Brodley |
SIGCSE (1) | 3 |
| 2024 | Re-making CS Departments for Generation CSabstractIn response to the enrollment surge that started in many CS depart- ments around 2006, the Computing Research Association published a report on "Generation CS", which named a pervasive theme in computing education: more and a greater diversity of students are seeking computing education, even if not as traditional CS ma- jors. However, our curricula and departments have stayed much the same. We still mostly prepare students for software develop- ment jobs in the technology industry, while we rarely identify the damage that same industry has caused in our democratic societies. How do we do better? How do we change to meet the needs of a changing society? What strategies should we apply? We know that large-scale change will require structural shifts, but such shifts are likely to be slow and expensive, whereas smaller, "boots on the ground" initiatives can positively impact individuals but do little to change the systems that underlie the deeper-seated problems in computing. Navigating this paradox is imperative to our success as a field. Our panel will address the big questions about how to make structural changes in computing education in order to meet the greater needs of Generation CS. Kathleen J. Lehman, Carla E. Brodley, Mark Guzdial, Paul T. Tymann, Aman Yadav |
SIGCSE (2) | 2 |
| 2024 | Does Curricular Complexity in Computer Science Influence the Representation of Women CS Graduates?abstractNot all degree programs are created equal. Indeed, the structure, prerequisites and overall complexity of some programs create barriers that impede student success. Inspired by the methodology of previous papers investigating the inverse relationship between curricular complexity and program quality, in this paper we investigate the relationship between curricular complexity and the representation of women earning CS degrees. We created curricular maps of 60 computer science degrees and calculated measures such as program complexity, course blocking, delay factor, and total math/CS credits to understand complexity's correlation with the representation of women CS majors. Our results show that degree complexity, blocking factor, and delay factor are all inversely related to the representation of women. In addition, we present the courses that most commonly impede student progress and provide suggestions to enhance degree programs based on the insights gained. Albert Lionelle, McKenna Quam, Carla E. Brodley, Catherine Gill |
SIGCSE (1) | 3 |
| 2024 | Collecting, Analyzing, and Acting on Intersectional, Longitudinal Data and Pass/Fail/Withdraw Rates in Computing CoursesabstractWe present the Center for Inclusive Computing's data collection and visualization system, which enables computing departments to track and visualize their enrollment and course outcome data intersectionally and longitudinally. The system tracks the impact of institutional changes in how computing (particularly the introductory sequence) is discovered and experienced by undergraduates as measured by course outcome and persistence data. To date we have worked with and collected data from 52 U.S. computing departments. Collected data spans 2018-present and contains term-by- term, intersectional course enrollment and outcome data for CS 1-3, while also tracking declared majors and persistence to graduation. Drawing on our experience working with these universities we present guidelines for the analysis of intersectional, longitudinal data alongside our recommendations for actionable next steps. We present three case studies grounded in an analysis of CS1, demon- strating how an institution can understand their own computing program and develop interventions-specifically with an eye toward broadening participation in computing. Felix Muzny, Megan Giordano, Emma Sommers, Carla E. Brodley |
SIGCSE (1) | 4 |
| 2022 | Interdisciplinary Computing Majors (CS+X): Making it work at your UniversityabstractAs computing becomes increasingly relevant to all disciplines, interdisciplinary computing degrees become increasingly important. These interdisciplinary majors: 1) address the increasing need for computing knowledge across all disciplines; 2) have the potential to increase a student's employability; 3) give employers the opportunity to hire students who are trained in two fields relevant to the company; 4) by reducing the number of requirements for the computing degree, can alleviate some of the pressure faced by CS departments from booming enrollments ; and 5) broaden participation in computing - in particular, to increase the percentage of women. Despite these opportunities, as of 2022 there are only a few schools that have embraced this mission. There are significant implementation challenges to interdisciplinary majors: 1) university/college budget models often financially discourage interdisciplinary degrees; 2) difficulty determining the "unit" that will administer the combined degree; 3) already large demands on CS faculty time; 4) advisors/faculty must be trained to advise interdisciplinary majors; and 5) both disciplines have to agree to reduce their major course requirements. In this BoF session we will discuss these and other obstacles identified at the range of institutions represented at SIGCSE and how to overcome these implementation challenges. Carla E. Brodley, Valerie Barr |
SIGCSE (2) | 1 |
| 2022 | Broadening Participation in Computing via Ubiquitous Combined Majors (CS+X)abstractIn 2001, Khoury College of Computer Sciences at Northeastern University created their first combined majors with Cognitive Psychology, Mathematics and Physics. This type of degree has often been referred to as "CS+X" in the literature and is increasingly relevant as the need for interdisciplinary computer scientists grows. As of 2021, students at Northeastern can choose among three computing majors (Computer Science, Data Science or Cybersecurity) and 42 combined majors, which combine one of the three computing degrees with one of 29 distinct majors in other fields. Prior to 2014, combined majors were with the sciences, business and design. Over the last seven years, we created 29 new combined majors, explicitly creating combinations with fields where there has traditionally been greater gender diversity. The resulting increase in student interest and gender diversity over the last seven years is compelling. As of Fall 2020, 44.6% of the 2,800+ computing majors at Northeastern are pursuing combined majors, 39% of whom are women. This is substantially higher than the 21.5% reported in IPEDS for 2019 women computing graduates in the U.S. We did not observe any significant differences in racial and ethnic diversity between combined and computing only degrees. In this experience paper, we describe how we create and manage combined majors, and we present results on enrollments, admissions, graduation, internship placements, and how students discover combined majors. Carla E. Brodley, Benjamin Hescott, Jessica Biron, Ali Ressing, Melissa Peiken, Sarah Maravetz, Alan Mislove |
SIGCSE (1) | 1 |
| 2020 | A Quantitative Machine Learning Approach to Master Students Admission for Professional Institutions
Bryan Lackaye, Jennifer G. Dy, Carla E. Brodley |
EDM | 4 |
| 2020 | An MS in CS for non-CS Majors: Moving to Increase Diversity of Thought and Demographics in CSabstractWe have created, piloted and are growing the Align program, a Master of Science in Computer Science (MS in CS) for post-secondary graduates who did not major in CS. Our goal is to create a pathway to CS for all students, with particular attention to women and underrepresented minorities. Indeed, women represent 57% and underrepresented minorities represent 25% of all bachelor's recipients in the U.S., but only 19.5% and 12.6% of CS graduates, respectively. If we can fill this opportunity gap, we will satisfy a major economic need and address an issue of social equity and inclusion. In this paper, we present our "Bridge'' curriculum, which is a two-semester preparation for students to then join the traditional MS in CS students in master's-level classes. We describe co-curricular activities designed to help students succeed in the program. We present our empirical findings around enrollment, demographics, retention and job outcomes. Among our findings is that Align students outperform our traditional MS in CS students in grade point average. To date we have graduated 137 students and 827 are enrolled. Carla E. Brodley, Megan Barry, Aidan Connell, Catherine Gill, Ian Gorton, Benjamin Hescott, Bryan Lackaye, Cynthia LuBien, Leena Razzaq, Amit Shesh, Tiffani L. Williams, Andrea Pohoreckyj Danyluk |
SIGCSE | 1 |
| 2017 | Human-in-the-loop applied machine learningabstractMachine learning research in academia is often conducted in vitro, divorced from motivating practical applications. As a result researchers often lose the ability to ask the question: how can my human expert's knowledge be used to best improve the machine learning outcome? In this talk, we present three motivating applications that all benefit from human-guided machine learning: systematic reviews for evidence-based medicine, generating maps of global land cover of the Earth from remotely sensed data, and finding lesions in the MRI's of treatment resistant epilepsy patients. Our machine learning contributions span active learning, both supervised and unsupervised learning, and their combination with human input. The methods we created are applicable to a wide range of applications in science, medicine and business. Carla E. Brodley |
IEEE BigData | 1 |
| 2016 | A Non-parametric Approach to Detect Epileptogenic Lesions using Restricted Boltzmann MachinesabstractVisual detection of lesional areas on a cortical surface is critical in rendering a successful surgical operation for Treatment Resistant Epilepsy (TRE) patients. Unfortunately, 45% of Focal Cortical Dysplasia (FCD, the most common kind of TRE) patients have no visual abnormalities in their brains' 3D-MRI images. We collaborate with doctors from NYU Langone's Comprehensive Epilepsy Center and apply machine learning methodologies to identify the resective zones for these {MRI-negative} FCD patients. Our task is particularly challenging because MRI images can only provide a limited number of features. Furthermore, data from different patients often exhibit inter-patient variabilities due to age, gender, left/right handedness, etc. In this paper, we introduce a new approach which combines the restricted Boltzmann machines and a Bayesian non-parametric mixture model to address these issues. We demonstrate the efficacy of our model by applying it to a retrospective dataset of MRI-negative FCD patients who are seizure free after surgery. Thomas Thesen, Karen E. Blackmon, Jennifer G. Dy, Carla E. Brodley, Ruben Kuzniecky, Orrin Devinsky |
KDD | 6 |
| 2016 | Decrypting "Cryptogenic" Epilepsy: Semi-supervised Hierarchical Conditional Random Fields For Detecting Cortical Lesions In MRI-Negative PatientsabstractFocal cortical dysplasia (FCD) is the most common cause of pediatric epilepsy and the third most common cause in adults with treatment-resistant epilepsy. Surgical resection of the lesion is the most effective treatment to stop seizures. Technical advances in MRI have revolutionized the diagnosis of FCD, leading to high success rates for resective surgery. However, 45% of histologically confirmed FCD patients have normal MRIs (MRI-negative). Without a visible lesion, the success rate of surgery drops from 66% to 29%. In this work, we cast the problem of detecting potential FCD lesions using MRI scans of MRI-negative patients in an image segmentation framework based on hierarchical conditional random fields (HCRF). We use surface based morphometry to model the cortical surface as a two-dimensional surface which is then segmented at multiple scales to extract superpixels of different sizes. Each superpixel is assigned an outlier score by comparing it to a control population. The lesion is detected by fusing the outlier probabilities across multiple scales using a tree- structured HCRF. The proposed method achieves a higher detection rate, with superior recall and precision on a sample of twenty MRI-negative FCD patients as compared to a baseline across four morphological features and their combinations. Thomas Thesen, Karen E. Blackmon, Ruben Kuzniecky, Orrin Devinsky, Carla E. Brodley |
J. Mach. Learn. Res. | 6 |
| 2015 | Domain Induced Dirichlet Mixture of Gaussian Processes: An Application to Predicting Disease Progression in Multiple Sclerosis PatientsabstractPredicting disease course is critical in chronic progressive diseases such as multiple sclerosis (MS) for determining treatment. Forming an accurate predictive model based on clinical data is particularly challenging when data is gathered from multiple clinics/physicians as the labels vary with physicians' subjective judgment about clinical tests and further we have no a priori knowledge of the various types of physician subjectivity. At the same time, we often have some (limited) domain knowledge on how to group patients into disease progression subgroups. In this paper, we first present our rationale for choosing a Dirichlet mixture of Gaussian processes (DPMGP) model to address the subjectivity in our data. We then introduce a new approach to incorporating domain knowledge into the non-parametric mixture model. We demonstrate the efficacy of our model by applying it to two medical datasets to predict disease progression in MS patients and disability levels in early Parkinson's patients. Tanuja Chitnis, Brian C. Healy, Jennifer G. Dy, Carla E. Brodley |
ICDM | 5 |
| 2015 | Removing confounding factors via constraint-based clustering: An application to finding homogeneous groups of multiple sclerosis patients
Carla E. Brodley, Brian C. Healy, Tanuja Chitnis |
Artif. Intell. Medicine | 2 |
| 2014 | Discovering Better AAAI Keywords via Clustering with Community-Sourced ConstraintsabstractSelecting good conference keywords is important because they often determine the composition of review committees and hence which papers are reviewed by whom. But presently conference keywords are generated in an ad-hoc manner by a small set of conference organizers. This approach is plainly not ideal. There is no guarantee, for example, that the generated keyword set aligns with what the community is actually working on and submitting to the conference in a given year. This is especially true in fast moving fields such as AI. The problem is exacerbated by the tendency of organizers to draw heavily on preceding years' keyword lists when generating a new set. Rather than a select few ordaining a keyword set that that represents AI at large, it would be preferable to generate these keywords more directly from the data, with input from research community members. To this end, we solicited feedback from seven AAAI PC members regarding a previously existing keyword set and used these 'community-sourced constraints' to inform a clustering over the abstracts of all submissions to AAAI 2013. We show that the keywords discovered via this data-driven, human-in-the-loop method are at least as preferred (by AAAI PC members) as 2013's manually generated set, and that they include categories previously overlooked by organizers. Many of the discovered terms were used for this year's conference. Kelly Moran, Byron C. Wallace, Carla E. Brodley |
AAAI | 3 |
| 2014 | Hierarchical Conditional Random Fields for Outlier Detection: An Application to Detecting Epileptogenic Cortical MalformationsabstractWe cast the problem of detecting and isolating regions of abnormal cortical tissue in the MRIs of epilepsy patients in an image segmentation framework. Employing a multiscale approach we divide the surface images into segments of different sizes and then classify each segment as being an outlier, by comparing it to the same region across controls. The final classification is obtained by fusing the outlier probabilities obtained at multiple scales using a tree-structured hierarchical conditional random field (HCRF). The proposed method correctly detects abnormal regions in 90% of patients whose abnormality was detected via routine visual inspection of their clinical MRI. More importantly, it detects abnormalities in 80% of patients whose abnormality escaped visual inspection by expert radiologists. Thomas Thesen, Karen E. Blackmon, Orrin Devinsky, Ruben Kuzniecky, Carla E. Brodley |
ICML | 7 |
| 2014 | CSAX: Characterizing Systematic Anomalies in eXpression Data
Keith Noto, Carla E. Brodley, Saeed Majidi, Diana W. Bianchi, Donna K. Slonim |
RECOMB | 2 |
| 2014 | Addressing Human Subjectivity via Transfer Learning: An Application to Predicting Disease Outcome in Multiple Sclerosis PatientsabstractPredicting disease course is critical in chronic progressive diseases such as multiple sclerosis (MS). In our work we are applying machine learning methods to longitudinal records of MS patients to build a classifier that predicts whether a patient will have a significant increase in disability at the five year mark using information from the first two years of clinical visits. This prediction is key for choosing among the available treatments as some have more troubling side-effect profiles. Two challenges arise while learning with this data. First, patient data involves the physician's (possibly subjective) evaluation. Because a patient's data may come from one doctor on the first visit and different doctors on subsequent visits, it can be difficult to form an accurate predictor of disease outcome. In particular, some physicians may be biased in one direction, scoring each patient as more severe than would other physicians, while others may be biased in the opposite direction. Another challenge is that it is much easier to classify the cases with low future disability compared to cases with high future disability at early stages of the disease. This asymmetric property is due to the nature of the disease rather than class imbalance. In this paper we introduce a new transfer learning approach to handle these challenges. The algorithm builds a single SVM classifier for each doctor by dividing the entire dataset into primary (instances from the doctor) and auxiliary (instances from other doctors) sets. When applied to our dataset of MS patients our new approach is able to realize a significant increase in prediction performance over approaches that form a single SVM classifier for the entire dataset or for the physician's dataset alone. Carla E. Brodley, Tanuja Chitnis, Brian C. Healy |
SDM | 2 |
| 2012 | Active Label CorrectionabstractActive Label Correction (ALC) is an interactive method that cleans an established training set of mislabeled examples in conjunction with a domain expert. ALC presumes that the expert who conducts this review is either more accurate than the original annotator or has access to additional resources that ensure a high quality label. A high-cost re-review is possible because ALC proceeds iteratively, scoring the full training set but selecting only small batches of examples that are likely mislabeled. The expert reviews each batch and corrects any mislabeled examples, after which the classifier is retrained and the process repeats until the expert terminates it. We compare several instantiations of ALC to fully-automated methods that attempt to discard or correct label noise in a single pass. Our empirical results show that ALC outperforms single-pass methods in terms of selection efficiency and classifier accuracy. We evaluate the best ALC instantiation on our motivating task of detecting mislabeled and poorly formulated sites within a land cover classification training set from the geography domain. Umaa Rebbapragada, Carla E. Brodley, Damien Sulla-Menashe, Mark A. Friedl |
ICDM | 2 |
| 2012 | FRaC: a feature-modeling approach for semi-supervised and unsupervised anomaly detection
Keith Noto, Carla E. Brodley, Donna K. Slonim |
Data Min. Knowl. Discov. | 2 |
| 2011 | Class Imbalance, ReduxabstractClass imbalance (i.e., scenarios in which classes are unequally represented in the training data) occurs in many real-world learning tasks. Yet despite its practical importance, there is no established theory of class imbalance, and existing methods for handling it are therefore not well motivated. In this work, we approach the problem of imbalance from a probabilistic perspective, and from this vantage identify dataset characteristics (such as dimensionality, sparsity, etc.) that exacerbate the problem. Motivated by this theory, we advocate the approach of bagging an ensemble of classifiers induced over balanced bootstrap training samples, arguing that this strategy will often succeed where others fail. Thus in addition to providing a theoretical understanding of class imbalance, corroborated by our experiments on both simulated and real datasets, we provide practical guidance for the data mining practitioner working with imbalanced data. Byron C. Wallace, Kevin Small, Carla E. Brodley, Thomas A. Trikalinos |
ICDM | 3 |
| 2011 | The Constrained Weight Space SVM: Learning with Ranked Features
Kevin Small, Byron C. Wallace, Carla E. Brodley, Thomas A. Trikalinos |
ICML | 3 |
| 2011 | Who Should Label What? Instance Allocation in Multiple Expert Active LearningabstractThe active learning (AL) framework is an increasingly popular strategy for reducing the amount of human labeling effort required to induce a predictive model. Most work in AL has assumed that a single, infallible oracle provides labels requested by the learner at a fixed cost. However, real-world applications suitable for AL often include multiple domain experts who provide labels of varying cost and quality. We explore this multiple expert active learning (MEAL) scenario and develop a novel algorithm for instance allocation that exploits the meta-cognitive abilities of novice (cheap) experts in order to make the best use of the experienced (expensive) annotators. We demonstrate that this strategy outperforms strong baseline approaches to MEAL on both a sentiment analysis dataset and two datasets from our motivating application of biomedical citation screening. Furthermore, we provide evidence that novice labelers are often aware of which instances they are likely to mislabel. Byron C. Wallace, Kevin Small, Carla E. Brodley, Thomas A. Trikalinos |
SDM | 3 |
| 2010 | Anomaly Detection Using an Ensemble of Feature ModelsabstractWe present a new approach to semi-supervised anomaly detection. Given a set of training examples believed to come from the same distribution or class, the task is to learn a model that will be able to distinguish examples in the future that do not belong to the same class. Traditional approaches typically compare the position of a new data point to the set of "normal" training data points in a chosen representation of the feature space. For some data sets, the normal data may not have discernible positions in feature space, but do have consistent relationships among some features that fail to appear in the anomalous examples. Our approach learns to predict the values of training set features from the values of other features. After we have formed an ensemble of predictors, we apply this ensemble to new data points. To combine the contribution of each predictor in our ensemble, we have developed a novel, information-theoretic anomaly measure that our experimental results show selects against noisy and irrelevant features. Our results on 47 data sets show that for most data sets, this approach significantly improves performance over current state-of-the-art feature space distance and density-based approaches. Keith Noto, Carla E. Brodley, Donna K. Slonim |
ICDM | 2 |
| 2010 | Redefining class definitions using constraint-based clustering: an application to remote sensing of the earth's surfaceabstractTwo aspects are crucial when constructing any real world supervised classification task: the set of classes whose distinction might be useful for the domain expert, and the set of classifications that can actually be distinguished by the data. Often a set of labels is defined with some initial intuition but these are not the best match for the task. For example, labels have been assigned for land cover classification of the Earth but it has been suspected that these labels are not ideal and some classes may be best split into subclasses whereas others should be merged. This paper formalizes this problem using three ingredients: the existing class labels, the underlying separability in the data, and a special type of input from the domain expert. We require a domain expert to specify an L × L matrix of pairwise probabilistic constraints expressing their beliefs as to whether the L classes should be kept separate, merged, or split. This type of input is intuitive and easy for experts to supply. We then show that the problem can be solved by casting it as an instance of penalized probabilistic clustering (PPC). Our method, Class-Level PPC (CPPC) extends PPC showing how its time complexity can be reduced from O(N2) to O(NL) for the problem of class re-definition. We further extend the algorithm by presenting a heuristic to measure adherence to constraints, and providing a criterion for determining the model complexity (number of classes) for constraint-based clustering. We demonstrate and evaluate CPPC on artificial data and on our motivating domain of land cover classification. For the latter, an evaluation by domain experts shows that the algorithm discovers novel class definitions that are better suited to land cover classification than the original set of labels. Dan Preston, Carla E. Brodley, Roni Khardon, Damien Sulla-Menashe, Mark A. Friedl |
KDD | 2 |
| 2010 | Active learning for biomedical citation screeningabstractActive learning (AL) is an increasingly popular strategy for mitigating the amount of labeled data required to train classifiers, thereby reducing annotator effort. We describe a real-world, deployed application of AL to the problem of biomedical citation screening for systematic reviews at the Tufts Medical Center's Evidence-based Practice Center. We propose a novel active learning strategy that exploits a priori domain knowledge provided by the expert (specifically, labeled features)and extend this model via a Linear Programming algorithm for situations where the expert can provide ranked labeled features. Our methods outperform existing AL strategies on three real-world systematic review datasets. We argue that evaluation must be specific to the scenario under consideration. To this end, we propose a new evaluation framework for finite-pool scenarios, wherein the primary aim is to label a fixed set of examples rather than to simply induce a good predictive model. We use a method from medical decision theory for eliciting the relative costs of false positives and false negatives from the domain expert, constructing a utility measure of classification performance that integrates the expert preferences. Our findings suggest that the expert can, and should, provide more information than instance labels alone. In addition to achieving strong empirical results on the citation screening problem, this work outlines many important steps for moving away from simulated active learning and toward deploying AL for real-world applications. Byron C. Wallace, Kevin Small, Carla E. Brodley, Thomas A. Trikalinos |
KDD | 3 |
| 2010 | Semi-automated screening of biomedical citations for systematic reviewsabstractBACKGROUND: Systematic reviews address a specific clinical question by unbiasedly assessing and analyzing the pertinent literature. Citation screening is a time-consuming and critical step in systematic reviews. Typically, reviewers must evaluate thousands of citations to identify articles eligible for a given review. We explore the application of machine learning techniques to semi-automate citation screening, thereby reducing the reviewers' workload. RESULTS: We present a novel online classification strategy for citation screening to automatically discriminate "relevant" from "irrelevant" citations. We use an ensemble of Support Vector Machines (SVMs) built over different feature-spaces (e.g., abstract and title text), and trained interactively by the reviewer(s). Semi-automating the citation screening process is difficult because any such strategy must identify all citations eligible for the systematic review. This requirement is made harder still due to class imbalance; there are far fewer "relevant" than "irrelevant" citations for any given systematic review. To address these challenges we employ a custom active-learning strategy developed specifically for imbalanced datasets. Further, we introduce a novel undersampling technique. We provide experimental results over three real-world systematic review datasets, and demonstrate that our algorithm is able to reduce the number of citations that must be screened manually by nearly half in two of these, and by around 40% in the third, without excluding any of the citations eligible for the systematic review. CONCLUSIONS: We have developed a semi-automated citation screening algorithm for systematic reviews that has the potential to substantially reduce the number of citations reviewers have to manually screen, without compromising the quality and comprehensiveness of the review. Byron C. Wallace, Thomas A. Trikalinos, Joseph Lau, Carla E. Brodley, Christopher H. Schmid |
BMC Bioinform. | 4 |
| 2009 | Event Discovery in Time SeriesabstractThe discovery of events in time series can have important implications, such as identifying microlensing events in astronomical surveys, or changes in a patient's electrocardiogram. Current methods for identifying events require a sliding window of a fixed size, which is not ideal for all applications and could overlook important events. In this work, we develop probability models for calculating the significance of an arbitrary-sized sliding window and use these probabilities to find areas of significance. Because a brute force search of all sliding windows and all window sizes would be computationally intractable, we introduce a method for quickly approximating the results. We apply our method to over 100,000 astronomical time series from the MACHO survey, in which 56 different sections of the sky are considered, each with one or more known events. Our method was able to recover 100% of these events in the top 1% of the results, essentially pruning 99% of the data. Interestingly, our method was able to identify events that do not pass traditional event discovery procedures. Dan Preston, Pavlos Protopapas, Carla E. Brodley |
SDM | 3 |
| 2009 | Finding anomalous periodic time series
Umaa Rebbapragada, Pavlos Protopapas, Carla E. Brodley, Charles R. Alcock |
Mach. Learn. | 3 |
| 2009 | IP Covert Channel DetectionabstractA covert channel can occur when an attacker finds and exploits a shared resource that is not designed to be a communication mechanism. A network covert channel operates by altering the timing of otherwise legitimate network traffic so that the arrival times of packets encode confidential data that an attacker wants to exfiltrate from a secure area from which she has no other means of communication. In this article, we present the first public implementation of an IP covert channel, discuss the subtle issues that arose in its design, and present a discussion on its efficacy. We then show that an IP covert channel can be differentiated from legitimate channels and present new detection measures that provide detection rates over 95%. We next take the simple step an attacker would of adding noise to the channel to attempt to conceal the covert communication. For these noisy IP covert timing channels, we show that our online detection measures can fail to identify the covert channel for noise levels higher than 10%. We then provide effective offline search mechanisms that identify the noisy channels. Serdar Cabuk, Carla E. Brodley, Clay Shields |
ACM Trans. Inf. Syst. Secur. | 2 |
| 2008 | Generating High-Quality Training Data for Automated Land-Cover MappingabstractThis paper presents two machine learning techniques that greatly reduce the number of person-hours required to generate high-quality training data for land cover classification. The first technique uses active learning to guide the generation of training data by selecting only the most informative examples for labeling. The second technique identifies and mitigates the impact of mislabeled instances. Both techniques are tested on data from NASA's Moderate Resolution Imaging Spectroradiometer (MODIS), which has required thousands of person hours to label. Our results shows that the active learning method requires fewer labeled examples than random sampling to produce a high quality classifier. Our results on class noise mitigation show that if mislabelings occur, we can further improve classifier accuracy, and that weighting instances by their label confidence outperforms an analogous method that discards suspected mislabelings. If combined, these methods have the potential to make training data generation a more efficient and reliable process. Umaa Rebbapragada, Rachel Lomasky, Carla E. Brodley, Mark A. Friedl |
IGARSS (4) | 3 |
| 2008 | A Distance-Based Method for Detecting Horizontal Gene Transfer in Whole Genomes
Xintao Wei, Lenore Cowen, Carla E. Brodley, Arthur Brady, D. Sculley, Donna K. Slonim |
ISBRA | 3 |
| 2007 | Active Class Selection
Rachel Lomasky, Carla E. Brodley, M. Aernecke, David R. Walt, Mark A. Friedl |
ECML | 2 |
| 2007 | Class Noise Mitigation Through Instance Weighting
Umaa Rebbapragada, Carla E. Brodley |
ECML | 2 |
| 2006 | Offloading IDS Computation to the GPUabstractSignature-matching intrusion detection systems can experience significant decreases in performance when the load on the IDS-host increases. We propose a solution that off-loads some of the computation performed by the IDS to the graphics processing unit (GPU). Modern GPUs are programmable, stream-processors capable of high-performance computing that in recent years have been used in non-graphical computing tasks. The major operation in a signature-matching IDS is matching values seen operation to known black-listed values, as such, our solution implements the string-matching on the GPU. The results show that as the CPU load on the IDS host system increases, PixelSnort's performance is significantly more robust and is able to outperform conventional Snort by up to 40% Nigel Jacob, Carla E. Brodley |
ACSAC | 2 |
| 2006 | Compression and Machine Learning: A New Perspective on Feature Space VectorsabstractThe use of compression algorithms in machine learning tasks such as clustering and classification has appeared in a variety of fields, sometimes with the promise of reducing problems of explicit feature selection. The theoretical justification for such methods has been founded on an upper bound on Kolmogorov complexity and an idealized information space. An alternate view shows compression algorithms implicitly map strings into implicit feature space vectors, and compression-based similarity measures compute similarity within these feature spaces. Thus, compression-based methods are not a "parameter free" magic bullet for feature selection and data representation, but are instead concrete similarity measures within defined feature spaces, and are therefore akin to explicit feature vector models used in standard machine learning algorithms. To underscore this point, we find theoretical and empirical connections between traditional machine learning vector models and compression, encouraging cross-fertilization in future work D. Sculley, Carla E. Brodley |
DCC | 2 |
| 2006 | SmashGuard: A Hardware Solution to Prevent Security Attacks on the Function Return AddressabstractA buffer overflow attack is perhaps the most common attack used to compromise the security of a host. This attack can be used to change the function return address and redirect execution to the attacker's code. We present a hardware-based solution, called SmashGuard, to protect against all known forms of attack on the function return addresses stored on the program stack. With each function call instruction, the current return address is pushed onto a hardware stack. A return instruction compares its address to the return address from the top of the hardware stack. An exception is raised to signal the mismatch. Because the stack operations and checks are done in hardware in parallel with the usual execution of instructions, our best-performing implementation scheme has virtually no performance overhead (because we are modifying hardware, it is impossible to guarantee zero overhead without an actual hardware implementation). While previous software-based approaches' average performance degradation for the SPEC2000 benchmarks is only 2.8 percent, their worst-case degradation is up to 8.3 percent. Apart from the lack of robustness in performance, the software approaches' key disadvantages are less security coverage and the need for recompilation of applications. SmashGuard, on the other hand, is secure and does not require recompilation of applications. Hilmi Ozdoganoglu, T. N. Vijaykumar, Carla E. Brodley, Benjamin A. Kuperman, Ankit Jalote |
IEEE Trans. Computers | 3 |
| 2005 | Heat Stroke: Power-Density-Based Denial of Service in SMTabstractIn the past, there have been several denial of service (DOS) attacks which exhaust some shared resource (e.g., physical memory, process table, file descriptors, TCP connections) of the targeted machine. Though these attacks have been addressed, it is important to continue to identify and address new attacks because DOS is one of most prominent methods used to cause significant financial loss. A recent paper shows how to prevent attacks that exploit the sharing of pipeline resources (e.g., shared trace cache) in SMT to degrade the performance of normal threads. In this paper, we show that power density can be exploited in SMT to launch a novel DOS attack, called heat stroke. Heat stroke repeatedly accesses a shared resource to create a hot spot at the resource. Current solutions to hot spots inevitably involve slowing down the pipeline to let the hot spot cool down. Consequently, heat stroke slows down the entire SMT pipeline and severely degrades normal threads. We present a solution to heat stroke by identifying the thread that causes the hot spot and selectively slowing down the malicious thread while minimally affecting normal threads. Jahangir Hasan, Ankit Jalote, T. N. Vijaykumar, Carla E. Brodley |
HPCA | 4 |
| 2005 | Correlation Clustering for Learning Mixtures of Canonical Correlation ModelsabstractThis paper addresses the task of analyzing the correlation between two related domains X and Y. Our research is motivated by an Earth Science task that studies the relationship between vegetation and precipitation. A standard statistical technique for such problems is Canonical Correlation Analysis (CCA). A critical limitation of CCA is that it can only detect linear correlation between the two domains that is globally valid throughout both data sets. Our approach addresses this limitation by constructing a mixture of local linear CCA models through a process we name correlation clustering. In correlation clustering, both data sets are clustered simultaneously according to the data's correlation structure such that, within a cluster, domain X and domain Y are linearly correlated in the same way. Each cluster is then analyzed using the traditional CCA to construct local linear correlation models. We present results on both artificial data sets and Earth Science data sets to demonstrate that the proposed approach can detect useful correlation patterns, which traditional CCA fails to discover. Xiaoli Z. Fern, Carla E. Brodley, Mark A. Friedl |
SDM | 2 |
| 2004 | IP covert timing channels: design and detectionabstractA network covert channel is a mechanism that can be used to leak information across a network in violation of a security policy and in a manner that can be difficult to detect. In this paper, we describe our implementation of a covert network timing channel, discuss the subtle issues that arose in its design, and present performance data for the channel. We then use our implementation as the basis for our experiments in its detection. We show that the regularity of a timing channel can be used to differentiate it from other traffic and present two methods of doing so and measures of their efficiency. We also investigate mechanisms that attackers might use to disrupt the regularity of the timing channel, and demonstrate methods of detection that are effective against them. Serdar Cabuk, Carla E. Brodley, Clay Shields |
CCS | 2 |
| 2004 | Solving cluster ensemble problems by bipartite graph partitioningabstractA critical problem in cluster ensemble research is how to combine multiple clusterings to yield a final superior clustering result. Leveraging advanced graph partitioning techniques, we solve this problem by reducing it to a graph partitioning problem. We introduce a new reduction method that constructs a bipartite graph from a given cluster ensemble. The resulting graph models both instances and clusters of the ensemble simultaneously as vertices in the graph. Our approach retains all of the information provided by a given ensemble, allowing the similarity among instances and the similarity among clusters to be considered collectively in forming the final clustering. Further, the resulting graph partitioning problem can be solved efficiently. We empirically evaluate the proposed approach against two commonly used graph formulations and show that it is more robust and achieves comparable or better performance in comparison to its competitors. Xiaoli Z. Fern, Carla E. Brodley |
ICML | 2 |
| 2004 | User re-authentication via mouse movementsabstractWe present an approach to user re-authentication based on the data collected from the computer's mouse device. Our underlying hypothesis is that one can successfully model user behavior on the basis of user-invoked mouse movements. Our implemented system raises an alarm when the current behavior of user X, deviates sufficiently from learned "normal" behavior of user X. We apply a supervised learning method to discriminate among k users. Our empirical results for eleven users show that we can differentiate these individuals based on their mouse movement behavior with a false positive rate of 0.43% and a false negative rate of 1.75%. Nevertheless, we point out that analyzing mouse movements alone is not sufficient for a stand-alone user re-authentication system. Maja Pusara, Carla E. Brodley |
VizSEC | 2 |
| 2004 | Feature Selection for Unsupervised Learning
Jennifer G. Dy, Carla E. Brodley |
J. Mach. Learn. Res. | 2 |
| 2003 | Behavioral Authentication of Server FlowsabstractUnderstanding the nature of the information flowing into and out of a system or network is fundamental to determining if there is adherence to a usage policy. Traditional methods of determining traffic type rely on the port label carried in the packet header. This method can fail, however, in the presence of proxy servers that remap port numbers or host services that have been compromised to act as backdoors or covert channels. We present an approach to classify server traffic based on decision trees learned during a training phase. The trees are constructed from traffic described using a set of features we designed to capture stream behavior. Because our classification of the traffic type is independent of port label, it provides a more accurate classification in the presence of malicious activity. An empirical evaluation illustrates that models of both aggregate protocol behavior and host-specific protocol behavior obtain classification accuracies ranging from 82-100%. James P. Early, Carla E. Brodley, Catherine Rosenberg |
ACSAC | 2 |
| 2003 | Boosting Lazy Decision Trees
Xiaoli Z. Fern, Carla E. Brodley |
ICML | 2 |
| 2003 | Random Projection for High Dimensional Data Clustering: A Cluster Ensemble Approach
Xiaoli Z. Fern, Carla E. Brodley |
ICML | 2 |
| 2003 | An Empirical Study of Two Approaches to Sequence Learning for Anomaly Detection
Terran Lane, Carla E. Brodley |
Mach. Learn. | 2 |
| 2003 | Unsupervised Feature Selection Applied to Content-Based Retrieval of Lung ImagesabstractThis paper describes a new hierarchical approach to content-based image retrieval called the "customized-queries" approach (CQA). Contrary to the single feature vector approach which tries to classify the query and retrieve similar images in one step, CQA uses multiple feature sets and a two-step approach to retrieval. The first step classifies the query according to the class labels of the images using the features that best discriminate the classes. The second step then retrieves the most similar images within the predicted class using the features customized to distinguish "subclasses" within that class. Needing to find the customized feature subset for each class led us to investigate feature selection for unsupervised learning. As a result, we developed a new algorithm called FSSEM (feature subset selection using expectation-maximization clustering). We applied our approach to a database of high resolution computed tomography lung images and show that CQA radically improves the retrieval precision over the single feature vector approach. To determine whether our CBIR system is helpful to physicians, we conducted an evaluation trial with eight radiologists. The results show that our system using CQA retrieval doubled the doctors' diagnostic accuracy. Jennifer G. Dy, Carla E. Brodley, Avinash C. Kak, Lynn S. Broderick, Alex M. Aisen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | Integration of domain knowledge in the form of ancillary map data into supervised classification of remotely sensed dataabstractIn recent years, machine learning and data mining methods have become increasingly common in remote sensing applications. One area in which such techniques are particularly useful is classification of remotely sensed data for land cover and vegetation mapping applications. In this paper, we describe new methods to include available information (domain knowledge) in supervised classification of land cover using high dimensional remote sensing observations. specifically, land cover and vegetation classification schemes are generally designed for ecological or land use applications. As a result, the classes of interest are often poorly separable in the multi-spectral or multi-temporal feature space provided by remote sensing. In many cases, ancillary data sources can provide useful information to help distinguish between problematic classes. However, available methods for including ancillary data sources, such as the use of prior probabilities in maximum likelihood classification, are often problematic in practice. This paper presents a method for incorporating prior probabilities in remote sensing-based land cover classification using a supervised decision tree classification algorithm. The method exploits recent theory from the domain of statistics and machine learning that allows robust estimates of class membership to be estimated using a technique known as boosting. This approach allows poorly separable classes to be distinguished based on ancillary information, but does not penalize rare classes. Mark A. Friedl, Douglas K. McIver, Carla E. Brodley |
IGARSS | 3 |
| 2002 | Interactive Content-Based Image Retrieval Using Relevance Feedback
Sean D. MacArthur, Carla E. Brodley, Avinash C. Kak, Lynn S. Broderick |
Comput. Vis. Image Underst. | 2 |
| 2002 | Using Human Perceptual Categories for Content-Based Retrieval from a Medical Image Database
Chi-Ren Shyu, Christina Pavlopoulou, Avinash C. Kak, Carla E. Brodley, Lynn S. Broderick |
Comput. Vis. Image Underst. | 4 |
| 2001 | The Effect of Instance-Space Partition on Significance
Jeffrey P. Bradford, Carla E. Brodley |
Mach. Learn. | 2 |
| 2001 | Focusing attention on objects of interest using multiple matched filtersabstractIn order to be of use to scientists, large image databases need to be analyzed to create a catalog of the objects of interest. One approach is to apply a multiple tiered search algorithm that uses reduction techniques of increasing computational complexity to select the desired objects from the database. The first tier of this type of algorithm, often called a focus of attention (FOA) algorithm, selects candidate regions from the image data and passes them to the next tier of the algorithm. In this paper we present a new approach to FOA that employs multiple matched filters (MMF), one for each object prototype, to detect the regions of interest. The MMFs are formed using k-means clustering on a set of image patches identified by domain experts as positive examples of objects of interest. An innovation of the approach is to radically reduce the dimensionality of the feature space, used by the k-means algorithm, by taking block averages (spoiling) the sample image patches. The process of spoiling is analyzed and its applicability to other domains is discussed. The combination of the output of the MMFs is achieved through the projection of the detections back into an empty image and then thresholding. This research was motivated by the need to detect small volcanos in the Magellan probe data from Venus. An empirical evaluation of the approach illustrates that a combination of the MMF plus the average filter results in a higher likelihood of 100% detection of the objects of interest at a lower false positive rate than a single matched filter alone. Timothy M. Stough, Carla E. Brodley |
IEEE Trans. Image Process. | 2 |
| 2000 | Feature Subset Selection and Order Identification for Unsupervised Learning
Jennifer G. Dy, Carla E. Brodley |
ICML | 2 |
| 2000 | Data Reduction Techniques for Instance-Based Learning from Human/Computer Interface Data
Terran Lane, Carla E. Brodley |
ICML | 2 |
| 2000 | Visualization and interactive feature selection for unsupervised dataabstractFor many feature selection problems, a human denes the features that are potentially useful, and then a subset is chosen from the original pool of features using an automated feature selection algorithm. In contrast to supervised learning, class information is not available to guide the feature search for unsupervised learning tasks. In this paper, we introduce Visual-FSSEM (Visual Feature Subset Selection using Expectation-Maximization Clustering), which incorporates visualization techniques, clustering, and user interaction to guide the feature subset search and to enable a deeper understanding of the data. Visual-FSSEM, serves both as an exploratory and multivariate-data visualization tool. We illustrate Visual-FSSEM on a high-resolution computed tomography lung image data set. 1. INTRODUCTION Most research in unsupervised clustering assumes that when creating the target data set, the data analyst in conjunction with the domain expert was able to identify a small relevant set of ... Jennifer G. Dy, Carla E. Brodley |
KDD | 2 |
| 1999 | The Customized-Queries Approach to CBIR Using EMabstractThis paper makes two contributions. The first contribution is an approach called the "customized-queries" approach (CQA) to content-based image retrieval. The second is an algorithm called FSSEM that performs feature selection and clustering simultaneously. The customized queries approach first classifies a query using the features that best differentiate the major classes and then customizes the query to that class by using the features that best distinguish the images within the chosen major class. This approach is motivated by the observation that the features that are most effective in discriminating among images from different classes may not be the most effective for retrieval of visually similar images within a class. This occurs for domains in which not all pairs of images within one class have equivalent visual similarity, i.e., subclasses exists. Because we are not given subclass labels, we must simultaneously find the features that best discriminate the subclasses and at the same time find these subclasses. We use FSSEM to find these features. We apply this approach to content-based retrieval of high-resolution tomographic images of patients with lung disease and show that this approach radically improves the retrieval precision over the traditional approach that performs retrieval using a single feature vector. Jennifer G. Dy, Carla E. Brodley, Avinash C. Kak, Chi-Ren Shyu, Lynn S. Broderick |
CVPR | 2 |
| 1999 | Predictive Application-Performance Modeling in a Computational Grid EnvironmentabstractThis paper describes and evaluates the application of three local learning algorithms-nearest-neighbor, weighted-average, and locally-weighted polynomial regression-for the prediction of run-specific resource-usage on the basis of run-time input parameters supplied to tools. A two-level knowledge base allows the learning algorithms to track short-term fluctuations in the performances of computing systems, and the use of instance editing techniques improves the scalability of the performance-modeling system. The learning algorithms assist PUNCH, a network-computing system at Purdue University, in emulating an ideal user in terms of its resource management and usage policies. Nirav H. Kapadia, José A. B. Fortes, Carla E. Brodley |
HPDC | 3 |
| 1999 | A Hybrid Lazy-Eager Approach to Reducing the Computation and Memory Requirements of Local Parametric Learning Algorithms
Yuanhui Zhou, Carla E. Brodley |
ICML | 2 |
| 1999 | ASSERT: A Physician-in-the-Loop Content-Based Retrieval System for HRCT Image Databases
Chi-Ren Shyu, Carla E. Brodley, Avinash C. Kak, Akio Kosaka, Alex M. Aisen, Lynn S. Broderick |
Comput. Vis. Image Underst. | 2 |
| 1999 | Identifying Mislabeled Training DataabstractThis paper presents a new approach to identifying and eliminating mislabeled training instances for supervised learning. The goal of this approach is to improve classification accuracies produced by learning algorithms by improving the quality of the training data. Our approach uses a set of learning algorithms to create classifiers that serve as noise filters for the training data. We evaluate single algorithm, majority vote and consensus filters on five datasets that are prone to labeling errors. Our experiments illustrate that filtering significantly improves classification accuracy for noise levels up to 30 percent. An analytical and empirical evaluation of the precision of our approach shows that consensus filters are conservative at throwing away good data at the expense of retaining bad data and that majority filters are better at detecting bad data at the expense of throwing away good data. This suggests that for situations in which there is a paucity of data, consensus filters are preferable, whereas majority vote filters are preferable for situations with an abundance of data. Carla E. Brodley, Mark A. Friedl |
J. Artif. Intell. Res. | 1 |
| 1999 | Maximizing land cover classification accuracies produced by decision trees at continental to global scalesabstractClassification of land cover from remotely sensed data at continental to global scales requires sophisticated algorithms and feature selection techniques to optimize classifier performance. The authors examine methods to maximize classification accuracies using decision trees to map land cover from multitemporal AVHRR imagery at continental and global scales. As part of their analysis they test the utility of "boosting", a new technique developed to increase classification accuracy by forcing the learning (classification) algorithm to concentrate on those training observations that are most difficult to classify. Their results show that boosting consistently reduces misclassification rates by 20-50% depending on the data set in question, and that most of the benefit gained by boosting is achieved after seven boosting iterations. They also assess the utility of including phenological metrics and geographic position as additional features to the classification algorithm. They find that using derived phenological metrics produces little improvement in classification accuracy relative to using an annual time series of NDVI data, but that geographic position provides substantial power for predicting land cover types at continental and global scales. However, in order to avoid generating spurious classification accuracies using geographic position, training data must be distributed evenly in geographic space. Mark A. Friedl, Carla E. Brodley, Alan H. Strahler |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 1999 | Temporal Sequence Learning and Data Reduction for Anomaly DetectionabstractThe anomaly-detection problem can be formulated as one of learning to characterize the behaviors of an individual, system, or network in terms of temporal sequences of discrete data. We present an approach on the basis of instance-based learning (IBL) techniques. To cast the anomaly-detection task in an IBL framework, we employ an approach that transforms temporal sequences of discrete, unordered observations into a metric space via a similarity measure that encodes intra-attribute dependencies. Classification boundaries are selected from an a posteriori characterization of valid user behaviors, coupled with a domain heuristic. An empirical evaluation of the approach on user command data demonstrates that we can accurately differentiate the profiled user from alternative users when the available features encode sufficient information. Furthermore, we demonstrate that the system detects anomalous conditions quickly — an important quality for reducing potential damage by a malicious user. We present several techniques for reducing data storage requirements of the user profile, including instance-selection methods and clustering. As empirical evaluation shows that a new greedy clustering algorithm reduces the size of the user model by 70%, with only a small loss in accuracy. Terran Lane, Carla E. Brodley |
ACM Trans. Inf. Syst. Secur. | 2 |
| 1998 | Temporal Sequence Learning and Data Reduction for Anomaly DetectionabstractThe anom~y detection problem can be formtiated as one of learning to characterize the behaviom of an individud, system, or network in terms of temporal sequences of di~ crete data.We present an approach to this problem based on instance based learning (IBL) techniques.To cast the anom~y detection task in an IBL framework, we employ an approach that transforms temporal sequences of discrete, unordered observations into a metric space via a similarity measure that encodes intra-attribute dependence=.Classification boundaries are selected from an a posterior characterization of the v&d user's behaviors, coupled with a d~ main heuristic.An empirical evrduation of the approach on user command data demonstratees that we can accurately differentiate the profled user from alternative users when the avdable featura encode sticient information.Futherrnore, we demonstrate that the system detects anomalous conditions quickly -an important qufity for reducing potential damage by a mdcious user.We present several techniques for reducing the data storage requirements of the user profle, including instance selection methods and clustering.An empirical evahtation shows that a new greedy clustering dgonthm reduces the size of the user model by 70% with ordy a smd loss in accuracy.A comparison of the greedy clustering technique to clustering with K-centers shows that greedy clustering is preferable in terms of accuracy and computation time for this domain. Terran Lane, Carla E. Brodley |
CCS | 2 |
| 1998 | Pruning Decision Trees with Misclassification Costs
Jeffrey P. Bradford, Clayton Kunz, Ron Kohavi, Clifford Brunk, Carla E. Brodley |
ECML | 5 |
| 1998 | Approaches to Online Learning and Concept Drift for User Identification in Computer Security
Terran Lane, Carla E. Brodley |
KDD | 2 |
| 1998 | Resource-Usage Prediction for Demand-Based Network-ComputingabstractThis paper reports on an application of artificial intelligence to achieve demand-based scheduling within the context of a network-computing infrastructure. The described AI system uses tool-specific, run-time input to predict the resource-usage characteristics of runs. Instance-based learning with locally weighted polynomial regression is employed because of the need to simultaneously learn multiple polynomial concepts and the fact that knowledge is acquired incrementally in this domain. An innovative use of a two-level knowledge base allows the system to account for short-term variations in compute-server and network performance and exploit temporal and spatial locality of runs. Instance editing allows the approach to be tolerant to noise and computationally feasible for extended use. The learning system was tested on three tools during normal use of the Purdue University Network Computing Hubs. Results indicate that the described instance-based learning technique using locally weighted regression with a locally linear model works well for this domain. Nirav H. Kapadia, Carla E. Brodley, José A. B. Fortes, Mark S. Lundstrom |
SRDS | 2 |
| 1997 | Image Feature Reduction through Spoiling: Its Application to Multiple Matched Filters for Focus of Attention
Timothy M. Stough, Carla E. Brodley |
KDD | 2 |
| 1997 | Learning to Schedule Straight-Line Code
J. Eliot B. Moss, Paul E. Utgoff, John Cavazos, Doina Precup, Darko Stefanovic, Carla E. Brodley, David Scheeff |
NIPS | 6 |
| 1995 | Automatic Selection of Split Criterion during Tree Growing Based on Node Location
Carla E. Brodley |
ICML | 1 |
| 1995 | Recursive Automatic Bias Selection for Classifier Construction
Carla E. Brodley |
Mach. Learn. | 1 |
| 1995 | Multivariate Decision Trees
Carla E. Brodley, Paul E. Utgoff |
Mach. Learn. | 1 |
| 1994 | Goal-Directed Classification Using Linear Machine Decision TreesabstractRecent work in feature-based classification has focused on nonparametric techniques that can classify instances even when the underlying feature distributions are unknown. The inference algorithms for training these techniques, however, are designed to maximize the accuracy of the classifier, with all errors weighted equally. In many applications, certain errors are far more costly than others, and the need arises for nonparametric classification techniques that can be trained to optimize task-specific cost functions. This correspondence reviews the linear machine decision tree (LMDT) algorithm for inducing multivariate decision trees, and shows how LMDT can be altered to induce decision trees that minimize arbitrary misclassification cost functions (MCF's). Demonstrations of pixel classification in outdoor scenes show how MCF's can optimize the performance of embedded classifiers within the context of larger image understanding systems.> Bruce A. Draper, Carla E. Brodley, Paul E. Utgoff |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1993 | Automatic Algorith/Model Class Selection
Carla E. Brodley |
ICML | 1 |
| 1990 | An Incremental Method for Finding Multivariate Splits for Decision Trees
Paul E. Utgoff, Carla E. Brodley |
ML | 2 |