VLDB 2026 Research / reviewers in the wild / expert
Lawrence B. Holder
dblp:h/LawrenceBHolder · also Larry B. Holder, Larry Holder, Lawrence Holder
· DBLP profile ↗
83ranked-venue papers
10as first author
7since 2021 · last 2026
0000-0002-6586-3144ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 44 · 6 first-author · 2 since 2021Databases, data management, data science and information retrieval · 25 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 9 · 1 first-authorSecurity and privacy · 4Systems, architecture and hardware · 2Theory of computation · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Feature-Augmented Transformer Model to Recognize Functional Activities From in-the-Wild Smartwatch DataabstractHuman activity recognition (HAR) from wearable sensor data traditionally identifies atomic movements (e.g., sit, stand, walk). However, many medical fields require recognizing functional activities-higher-level, goal-directed behaviors (e.g., errands, socialize, work). Functional activity recognition is critical for cognitive health assessment, rehabilitation, post-surgical recovery, and chronic disease management, yet remains largely unexplored due to its inherent complexity and variability for in-the-wild settings. This work addresses these challenges by investigating methods for functional HAR and introducing a novel approach that augments feature representations with feature token-transformer embeddings to improve classification performance. We compare a range of machine learning and deep learning methods, analyzing their ability to generalize across a diverse population. Additionally, we present ArWISE, a large-scale functional activity dataset collected longitudinally from n = 503 participants, consisting of over 32 million labeled points. Our experiments demonstrate the advantages of incorporating feature embeddings into functional HAR models, particularly in handling real-world variability and data sparsity. By bridging the gap between atomic movement recognition and functional behavior modeling, this work lays the foundation for more advanced, behavior-aware applications in digital health and human-centered AI. Bryan David Minor, Colin Greeley, Ryan Holder, Lawrence B. Holder, Diane J. Cook |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Detecting and reacting to smart home novelties
Lawrence B. Holder, Baxter Eaves, Patrick Shafto, Christopher Pereyda, Brian L. Thomas, Diane J. Cook |
Data Min. Knowl. Discov. | 1 |
| 2023 | Hybrid deep learning approach to improve classification of low-volume high-dimensional dataabstractBACKGROUND: The performance of machine learning classification methods relies heavily on the choice of features. In many domains, feature generation can be labor-intensive and require domain knowledge, and feature selection methods do not scale well in high-dimensional datasets. Deep learning has shown success in feature generation but requires large datasets to achieve high classification accuracy. Biology domains typically exhibit these challenges with numerous handcrafted features (high-dimensional) and small amounts of training data (low volume). METHOD: A hybrid learning approach is proposed that first trains a deep network on the training data, extracts features from the deep network, and then uses these features to re-express the data for input to a non-deep learning method, which is trained to perform the final classification. RESULTS: The approach is systematically evaluated to determine the best layer of the deep learning network from which to extract features and the threshold on training data volume that prefers this approach. Results from several domains show that this hybrid approach outperforms standalone deep and non-deep learning methods, especially on low-volume, high-dimensional datasets. The diverse collection of datasets further supports the robustness of the approach across different domains. CONCLUSIONS: The hybrid approach combines the strengths of deep and non-deep learning paradigms to achieve high performance on high-dimensional, low volume learning tasks that are typical in biology domains. Pegah Mavaie, Lawrence B. Holder, Michael K. Skinner |
BMC Bioinform. | 2 |
| 2022 | Multimodal Fusion of Smart Home and Text-based Behavior Markers for Clinical Assessment PredictionabstractNew modes of technology are offering unprecedented opportunities to unobtrusively collect data about people's behavior. While there are many use cases for such information, we explore its utility for predicting multiple clinical assessment scores. Because clinical assessments are typically used as screening tools for impairment and disease, such as mild cognitive impairment (MCI), automatically mapping behavioral data to assessment scores can help detect changes in health and behavior across time. In this article, we aim to extract behavior markers from two modalities, a smart home environment and a custom digital memory notebook app, for mapping to 10 clinical assessments that are relevant for monitoring MCI onset and changes in cognitive health. Smart-home-based behavior markers reflect hourly, daily, and weekly activity patterns, while app-based behavior markers reflect app usage and writing content/style derived from free-form journal entries. We describe machine learning techniques for fusing these multimodal behavior markers and utilizing joint prediction. We evaluate our approach using three regression algorithms and data from 14 participants with MCI living in a smart-home environment. We observed moderate to large correlations between predicted and ground-truth assessment scores, ranging from r = 0.601 to r = 0.871 for each clinical assessment. Gina Sprint, Diane J. Cook, Maureen Schmitter-Edgecombe, Lawrence B. Holder |
ACM Trans. Comput. Heal. | 4 |
| 2022 | ITeM: Independent temporal motifs to summarize and compare temporal networksabstractNetworks are a fundamental and flexible way of representing various complex systems. Many domains such as communication, citation, procurement, biology, social media, and transportation can be modeled as a set of entities and their relationships. Temporal networks are a specialization of general networks where every relationship occurs at a discrete time. The temporal evolution of such networks is as important to understand as the structure of the entities and relationships. We present the Independent Temporal Motif (ITeM) to characterize temporal graphs from different domains. ITeMs can be used to model the structure and the evolution of the graph. In contrast to existing work, ITeMs are edge-disjoint directed motifs that measure the temporal evolution of ordered edges within the motif. For a given temporal graph, we produce a feature vector of ITeM frequencies and the time it takes to form the ITeM instances. We apply this distribution to measure the similarity of temporal graphs. We show that ITeM has higher accuracy than other motif frequency-based approaches. We define various ITeM-based metrics that reveal salient properties of a temporal network. We also present importance sampling as a method to efficiently estimate the ITeM counts. We present a distributed implementation of the ITeM discovery algorithm using Apache Spark and GraphFrame. We evaluate our approach on both synthetic and real temporal networks. Sumit Purohit, George Chin, Lawrence B. Holder |
Intell. Data Anal. | 3 |
| 2021 | Towards a Unifying Framework for Formal Theories of NoveltyabstractManaging inputs that are novel, unknown, or out-of-distribution is critical as an agent moves from the lab to the open world. Novelty-related problems include being tolerant to novel perturbations of the normal input, detecting when the input includes novel items, and adapting to novel inputs. While significant research has been undertaken in these areas, a noticeable gap exists in the lack of a formalized definition of novelty that transcends problem domains. As a team of researchers spanning multiple research groups and different domains, we have seen, first hand, the difficulties that arise from ill-specified novelty problems, as well as inconsistent definitions and terminology. Therefore, we present the first unified framework for formal theories of novelty and use the framework to formally define a family of novelty types. Our framework can be applied across a wide range of domains, from symbolic AI to reinforcement learning, and beyond to open world image recognition. Thus, it can be used to help kick-start new research efforts and accelerate ongoing work on these important novelty-related problems. Terrance E. Boult, Przemyslaw A. Grabowicz, Derek S. Prijatelj, Roni Stern, Lawrence B. Holder, Joshua Alspector, Mohsen Jafarzadeh, Touqeer Ahmad, Akshay Raj Dhamija, Chunchun Li, Steve Cruz, Abhinav Shrivastava, Carl Vondrick, Walter J. Scheirer |
AAAI | 5 |
| 2021 | Predicting environmentally responsive transgenerational differential DNA methylated regions (epimutations) in the genome using a hybrid deep-machine learning approachabstractBACKGROUND: Deep learning is an active bioinformatics artificial intelligence field that is useful in solving many biological problems, including predicting altered epigenetics such as DNA methylation regions. Deep learning (DL) can learn an informative representation that addresses the need for defining relevant features. However, deep learning models are computationally expensive, and they require large training datasets to achieve good classification performance. RESULTS: One approach to addressing these challenges is to use a less complex deep learning network for feature selection and Machine Learning (ML) for classification. In the current study, we introduce a hybrid DL-ML approach that uses a deep neural network for extracting molecular features and a non-DL classifier to predict environmentally responsive transgenerational differential DNA methylated regions (DMRs), termed epimutations, based on the extracted DL-based features. Various environmental toxicant induced epigenetic transgenerational inheritance sperm epimutations were used to train the model on the rat genome DNA sequence and use the model to predict transgenerational DMRs (epimutations) across the entire genome. CONCLUSION: The approach was also used to predict potential DMRs in the human genome. Experimental results show that the hybrid DL-ML approach outperforms deep learning and traditional machine learning methods. Pegah Mavaie, Lawrence B. Holder, Daniel Beck, Michael K. Skinner |
BMC Bioinform. | 2 |
| 2020 | Graph Filtering to Remove the "Middle Ground" for Anomaly DetectionabstractDiscovering patterns and anomalies in a variety of voluminous data represented as a graph is challenging. Current research has demonstrated success discovering graph patterns using a sampling of the data, but there has been little work when it comes to discovering anomalies that are based upon understanding what is normative. In this work we present two approaches to reducing graph data: subgraph filtering and graph filtering. The idea behind the proposed algorithms is the removal of a "murky middle", where data that may not be normative or anomalous, is removed from the discovery process. We empirically validate the proposed approach on real-world, pseudo-real-world, and synthetic data, as well as compare against a similar approach. William Eberle, Lawrence B. Holder |
IEEE BigData | 2 |
| 2019 | Efficient frequent subgraph mining on large streaming graphsabstractWe propose an efficient, approximate algorithm to solve the problem of finding frequent subgraphs in large streaming graphs. The graph stream is treated as batches of labeled nodes and edges. Our proposed algorithm finds the set of frequent subgraphs as the graph evolves after each batch. The compu tational complexity is bounded to linear limits by looking only at the changes made by the most recent batch, and the historical set of frequent subgraphs. As a part of our approach, we also propose a novel sampling algorithm that samples regions of the graph that have been changed by the most recent update to the graph. The performance of the proposed approach is evaluated using five large graph datasets, and our approach is shown to be faster than the state of the art large graph miners while maintaining their accuracy. We also compare our sampling algorithm against a well known sampling algorithm for network motif mining, and show that our sampling algorithm is faster, and capable of discovering more types of patterns. We provide theoretical guarantees of our algorithm’s accuracy using the well known Chernoff bounds, as well as an analysis of the computational complexity of our approach. Abhik Ray, Lawrence B. Holder, Albert Bifet |
Intell. Data Anal. | 2 |
| 2018 | Percolator: Scalable Pattern Discovery in Dynamic GraphsabstractWe demonstrate \perco, a distributed system for graph pattern discovery in dynamic graphs. In contrast to conventional mining systems, Percolator advocates efficient pattern mining schemes that (1) support pattern detection with keywords; (2) integrate incremental and parallel pattern mining; and (3) support analytical queries such as trend analysis. The core idea of \perco is to dynamically decide and verify a small fraction of patterns and their instances that must be inspected in response to buffered updates in dynamic graphs, with a total mining cost independent of graph size. We demonstrate a( the feasibility of incremental pattern mining by walking through each component of \perco, b) the efficiency and scalability of \perco over the sheer size of real-world dynamic graphs, and c) how the user-friendly \gui of \perco interacts with users to support keyword-based queries that detect, browse and inspect trending patterns. We demonstrate how \perco effectively supports event and trend analysis in social media streams and research publication, respectively. Sutanay Choudhury, Sumit Purohit, Yinghui Wu 0001, Lawrence B. Holder, Khushbu Agarwal |
WSDM | 5 |
| 2018 | Cross-environment activity recognition using a shared semantic vocabulary
Zachary Wemlinger, Lawrence B. Holder |
Pervasive Mob. Comput. | 2 |
| 2017 | Application-specific graph sampling for frequent subgraph mining and community detectionabstractGraph mining is an important data analysis methodology, but struggles as the input graph size increases. The scalability and usability challenges posed by such large graphs make it imperative to sample the input graph and reduce its size. The critical challenge in sampling is to identify the appropriate algorithm to insure the resulting analysis does not suffer heavily from the data reduction. Predicting the expected performance degradation for a given graph and sampling algorithm is also useful. In this paper, we present different sampling approaches for graph mining applications such as Frequent Subgrpah Mining (FSM), and Community Detection (CD). We explore graph metrics such as PageRank, Triangles, and Diversity to sample a graph and conclude that for heterogeneous graphs Triangles and Diversity perform better than degree based metrics. We also present two new sampling variations for targeted graph mining applications. We present empirical results to show that knowledge of the target application, along with input graph properties can be used to select the best sampling algorithm. We also conclude that performance degradation is an abrupt, rather than gradual phenomena, as the sample size decreases. We present the empirical results to show that the performance degradation follows a logistic function. Original Datasets, implementation of sampling algorithms, and results are available online. Sumit Purohit, Sutanay Choudhury, Lawrence B. Holder |
IEEE BigData | 3 |
| 2017 | Thyme: Improving Smartphone Prompt Timing Through Activity AwarenessabstractSmartphone prompts and notifications are popular because they provide users with timely and important information. However, they can also be an annoyance if they pop up at inopportune times and interrupt important tasks. In this paper, we introduce Thyme, an intelligent notification front end that uses activity recognition and machine learning to identify the best times to prompt smartphone users. We evaluate the performance of an activity-aware prompting approach based on 47 participants with fixed time and Thyme-based prompts. Our results show that responsiveness improves from 12.8% to 93.2% using this intelligent approach to the timing of smartphone-based prompts. Samaneh Aminikhanghahi, Ramin Fallahzadeh, Matthew Sawyer, Diane J. Cook, Lawrence B. Holder |
ICMLA | 5 |
| 2017 | Deep learning approach to link weight predictionabstractDeep learning has been successful in various domains including image recognition, speech recognition and natural language processing. However, the research on its application in graph mining is still in an early stage. Here we present Model R, a neural network model created to provide a deep learning approach to link weight prediction problem. This model extracts knowledge of nodes from known links' weights and uses this knowledge to predict unknown links' weights. We demonstrate the power of Model R through experiments and compare it with stochastic block model and its derivatives. Model R shows that deep learning can be successfully applied to link weight prediction and it outperforms stochastic block model and its derivatives by up to 73% in terms of prediction accuracy. We anticipate this new approach to provide effective solutions to more graph mining tasks. Yuchen Hou, Lawrence B. Holder |
IJCNN | 2 |
| 2016 | Classification in dynamic streaming networksabstractTraditional network classification techniques will become computationally intractable when applied on a network which is presented in a streaming fashion with continuous updates. In this paper, we examine the problem of classification in dynamic streaming networks, or graphs. Two scenarios have been considered: the graph transaction scenario and the one large graph scenario. We propose a unified framework consisting of three components: a subgraph extraction method, an online version of an existing graph kernel, and two kernel-based incremental learners. We demonstrate the advantages of our framework via empirical evaluations on several real-world network datasets. Yibo Yao, Lawrence B. Holder |
ASONAM | 2 |
| 2016 | StarIso: Graph Isomorphism Through Lossy CompressionabstractSummary form only given: Graph data has gained importance as social networks, shopping habits, and travel patterns are recorded in much greater detail and quantity. An important step in making this information useful is the ability to compare two different portions of this data. In this paper, we explore a method for fast compression of graph data and how that can be used for comparison. We show that when performing one-to-many matching it performs quite well against VF2, currently one of the best strict graph matching algorithms. Jason Fairey, Lawrence B. Holder |
DCC | 2 |
| 2016 | Network of Spiking Neurons Driven by CompressionabstractOur work aims to design an intelligent agent that chooses its actions based on compression as a reward signal. The design falls into the category of spiking neural network. In the scenarios tested, the goal of the agent is to compress an input stream of bytes. Neurons are organized by layers and connected to other neurons at adjacent layers. A neuron receives an increase in voltage based on how well its associated action compresses the input stream. The agent performs actions, based on the neuron, according to a probability that's calculated based on how successful the associated action has been at compressing the input stream and the strength of the neuron's connections to neighboring neurons that recently fired. The agent has been successful at learning to perform an action out of a subset of actions that leads to the most compression, and at learning sequences of actions that lead to significant or optimal compression. In some specific scenarios, the agent performs close to optimal actions that lead to significant compression rather than the optimal sequence of actions. There are limitations on the length of sequences of actions that can be learned, and some specific types of sequences could not be learned given the current structure, e.g., sequences that change with each iteration could not be learned. This agent compresses data in a novel and effective way and shows promise for simultaneously displaying intelligent behavior and compressing data intelligently. Alexander Gain, Lawrence B. Holder |
DCC | 2 |
| 2016 | Incremental SVM-based classification in dynamic streaming networksabstractWith the emergence of networked data, graph classification has received considerable interest during the past years. Most approaches to graph classification focus on designing effective kernels to compute similarities for static graphs. However, they become computationally intractable in terms of t ime and space when a graph is presented in an incremental fashion with continuous updates, i.e., insertions of nodes and edges. In this paper, we examine the problem of classification in large-scale and incrementally changing graphs. We propose a framework combining an incremental support vector machine (SVM) with the Weisfeiler-Lehman (W-L) graph kernel. By retaining the support vectors from each learning step, the classification model is incrementally updated whenever new changes are made to the graph. We design an entropy-based subgraph extraction strategy, that selects informative neighbor nodes and discards those with less discriminative power, to facilitate the classification of nodes in a dynamic network. We validate the advantages of our learning techniques by conducting an empirical evaluation on several large-scale real-world graph datasets in comparison with other graph classification methods. The experimental results also validate the benefits of our subgraph extraction method when combined with the incremental learning techniques. Yibo Yao, Lawrence B. Holder |
Intell. Data Anal. | 2 |
| 2015 | Scalable classification for large dynamic networksabstractWe examine the problem of node classification in large-scale and dynamically changing graphs. An entropy-based subgraph extraction method has been developed for extracting subgraphs surrounding the nodes to be classified. We introduce an online version of an existing graph kernel to incrementally compute the kernel matrix for a unbounded stream of these extracted subgraphs. After obtaining the kernel values, we adopt a kernel perceptron to learn a discriminative classifier and predict the class labels of the target nodes with their corresponding subgraphs. We demonstrate the advantages of our learning techniques by conducting empirical evaluations on two real-world graph datasets. Yibo Yao, Lawrence B. Holder |
IEEE BigData | 2 |
| 2015 | Fast and Accurate Support Vector Machines on Large Scale SystemsabstractSupport Vector Machines (SVM) is a supervised Machine Learning and Data Mining (MLDM) algorithm, which has become ubiquitous largely due to its high accuracy and obliviousness to dimensionality. The objective of SVM is to find an optimal boundary -- also known as hyperplane -- which separates the samples (examples in a dataset) of different classes by a maximum margin. Usually, very few samples contribute to the definition of the boundary. However, existing parallel algorithms use the entire dataset for finding the boundary, which is sub-optimal for performance reasons. In this paper, we propose a novel distributed memory algorithm to eliminate the samples which do not contribute to the boundary definition in SVM. We propose several heuristics, which range from early (aggressive) to late (conservative) elimination of the samples, such that the overall time for generating the boundary is reduced considerably. In a few cases, a sample may be eliminated (shrunk) pre-emptively -- potentially resulting in an incorrect boundary. We propose a scalable approach to synchronize the necessary data structures such that the proposed algorithm maintains its accuracy. We consider the necessary trade-offs of single/multiple synchronization using in-depth time-space complexity analysis. We implement the proposed algorithm using MPI and compare it with libsvm -- de facto sequential SVM software -- which we enhance with OpenMP for multi-core/many-core parallelism. Our proposed approach shows excellent efficiency using up to 4096 processes on several large datasets such as UCI HIGGS Boson dataset and Offending URL dataset. Abhinav Vishnu, Jeyanthi Narasimhan, Lawrence B. Holder, Darren J. Kerbyson, Adolfy Hoisie |
CLUSTER | 3 |
| 2015 | A Selectivity based approach to Continuous Pattern Detection in Streaming Graphs
Sutanay Choudhury, Lawrence B. Holder, George Chin, Khushbu Agarwal, John Feo |
EDBT | 2 |
| 2015 | Scalable anomaly detection in graphsabstractThe advantage of graph-based anomaly detection is that the relationships between elements can be analyzed for structural oddities that could represent activities such as fraud, network intrusions, or suspicious associations in a social network. Tradi William Eberle, Lawrence B. Holder |
Intell. Data Anal. | 2 |
| 2014 | A partitioning approach to scaling anomaly detection in graph streamsabstractDue to potentially complex relationships among heterogeneous data sets, recent research efforts have involved the representation of this type of complex data as a graph. For instance, in the case of computer network traffic, a graph representation of the traffic might consist of nodes representing computers and edges representing communications between the corresponding computers. However, computer network traffic is typically voluminous, or acquired in real-time as a stream of information. In previous work on static graphs, we have used a compression-based measure to find normative patterns, and then analyzed the close matches to the normative patterns to indicate potential anomalies. However, while our approach has demonstrated its effectiveness in a variety of domains, the issue of scalability has limited this approach when dealing with domains containing millions of nodes and edges. To address this issue, we propose a novel approach called Pattern Learning and Anomaly Detection on Streams, or PLADS, that is not only scalable to real-world data that is streaming, but also maintains reasonable levels of effectiveness in detecting anomalies. In this paper we present a partitioning and windowing approach that partitions the graph as it streams in over time and maintains a set of normative patterns and anomalies. We then empirically evaluate our approach using publicly available network data as well as a dataset that represents e-commerce traffic. William Eberle, Lawrence B. Holder |
IEEE BigData | 2 |
| 2014 | Scalable SVM-Based Classification in Dynamic GraphsabstractWith the emergence of networked data, graph classification has received considerable interest during the past years. Most approaches to graph classification focus on designing effective kernels to compute similarities for static graphs. However, they become computationally intractable in terms of time and space when a graph is presented in a incremental fashion with continuous updates, i.e., Insertions of nodes and edges. In this paper, we examine the problem of classification in large-scale and incrementally changing graphs. To this end, a framework combining an incremental Support Vector Machine (SVM) with the Weisfeiler-Lehman (W-L) graph kernel has been proposed to study this problem. By retaining the support vectors from each learning step, the classification model is incrementally updated whenever new changes are made to the subject graph. Furthermore, we design an entropy-based sub graph extraction strategy to select informative neighbor nodes and discard those with less discriminative power, to facilitate an effective classification process. We demonstrate the advantages of our learning techniques by conducting an empirical evaluation on two large-scale real-world graph datasets. The experimental results also validate the benefits of our sub graph extraction method when combined with the incremental learning techniques. Yibo Yao, Lawrence B. Holder |
ICDM | 2 |
| 2014 | Activity Recognition Using Graphical FeaturesabstractActivity Recognition is important in order to facilitate elderly residents' and their caregivers' needs. This problem has been widely investigated using different methods including probabilistic and Markovian approaches. The focus of this paper is to perform activity recognition more accurately than existing approaches using non-intrusive sensors. We represent motion sensors of smart environments in a graph and resident's movements as edges in the graph. Then graph-based features are extracted and used as input for a Support Vector Machine. These features have been combined with motion-sensor based features. This method has been compared with three other widely used approaches, Naive Bayes, Hidden Markov Model (HMM) and Conditional Random Fields (CRF) on three different datasets from three smart apartments. In all cases, the method based on graphical features outperformed one of the state of the art methods for activity recognition. Syeda Selina Akter, Lawrence B. Holder |
ICMLA | 2 |
| 2014 | Improving Activity Recognition in Smart Environments with Ontological Modeling
Zachary Wemlinger, Lawrence B. Holder |
ICOST | 2 |
| 2014 | Special issue on data mining in pervasive environments
Nirmalya Roy, Parisa Rashidi, Lawrence B. Holder, Liming Chen 0001 |
Pervasive Mob. Comput. | 3 |
| 2013 | Towards a network-of-networks framework for cyber securityabstractNetwork-of-networks (NoN) is a graph-theoretic model of interdependent networks that have distinct dynamics at each network (layer). By adding special edges to represent relationships between nodes in different layers, NoN provides a unified mechanism to study interdependent systems intertwined in a complex relationship. While NoN based models have been proposed for cyber-physical systems, in this position paper we build towards a three-layered NoN model for an enterprise cyber system. Each layer captures a different facet of a cyber system. We present in-depth discussion for four major graph-theoretic applications to demonstrate how the three-layered NoN model can be leveraged for continuous system monitoring and mission assurance. A longer version of this paper can be accessed from arXiv [1]. Mahantesh Halappanavar, Sutanay Choudhury, Emilie Hogan, Peter Hui, John R. Johnson, Indrajit Ray, Lawrence B. Holder |
ISI | 7 |
| 2013 | StreamWorks: a system for dynamic graph searchabstractActing on time-critical events by processing ever growing social media, news or cyber data streams is a major technical challenge. Many of these data sources can be modeled as multi-relational graphs. Mining and searching for subgraph patterns in a continuous setting requires an efficient approach to incremental graph search. The goal of our work is to enable real-time search capabilities for graph databases. This demonstration will present a dynamic graph query system that leverages the structural and semantic characteristics of the underlying multi-relational graph. Sutanay Choudhury, Lawrence B. Holder, George Chin, Abhik Ray, Sherman Beus, John Feo |
SIGMOD Conference | 2 |
| 2013 | Generalized Query-Based Active Learning to Identify Differentially Methylated Regions in DNAabstractActive learning is a supervised learning technique that reduces the number of examples required for building a successful classifier, because it can choose the data it learns from. This technique holds promise for many biological domains in which classified examples are expensive and time-consuming to obtain. Most traditional active learning methods ask very specific queries to the Oracle (e.g., a human expert) to label an unlabeled example. The example may consist of numerous features, many of which are irrelevant. Removing such features will create a shorter query with only relevant features, and it will be easier for the Oracle to answer. We propose a generalized query-based active learning (GQAL) approach that constructs generalized queries based on multiple instances. By constructing appropriately generalized queries, we can achieve higher accuracy compared to traditional active learning methods. We apply our active learning method to find differentially DNA methylated regions (DMRs). DMRs are DNA locations in the genome that are known to be involved in tissue differentiation, epigenetic regulation, and disease. We also apply our method on 13 other data sets and show that our method is better than another popular active learning technique. Md. Muksitul Haque, Lawrence B. Holder, Michael K. Skinner, Diane J. Cook |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2012 | Using smart phones for context-aware prompting in smart environmentsabstractIndividuals with cognitive impairment have difficulty successfully performing activities of daily living, which can lead to decreased independence. In order to help these individuals age in place and decrease caregiver burden, technologies for assistive living have gained popularity over the last decade. In this work, a context-aware prompting system is implemented, augmented by a smart phone to determine prompt situations in a smart home environment. While context-aware systems use temporal and environmental information to determine context, we additionally use ambulatory information from accelerometer data of a phone which also acts as a mobile prompting device. A pilot study with healthy young adults is conducted to examine the feasibility of using a smart phone interface for prompt delivery during activity completion in a smart home environment. Barnan Das, Adriana M. Seelye, Brian L. Thomas, Diane J. Cook, Lawrence B. Holder, Maureen Schmitter-Edgecombe |
CCNC | 5 |
| 2012 | Context-aware prompting from your smart phoneabstractIndividuals with cognitive impairment have difficulty successfully performing activities of daily living, which can lead to decreased independence. In order to help these individuals age in place and decrease caregiver burden, technologies for assistive living have gained popularity over the last decade. This demo illustrates the implementation of a context-aware prompting system augmented by a smart phone to determine prompt situations in a smart home environment. While context-aware systems use temporal and environmental information to determine context, we additionally use ambulatory information from accelerometer data of a phone which also acts as a mobile prompting device. Barnan Das, Brian L. Thomas, Adriana M. Seelye, Diane J. Cook, Lawrence B. Holder, Maureen Schmitter-Edgecombe |
CCNC | 5 |
| 2012 | Graph based MRI brain scan classification and correlation discoveryabstractThe shape of the human brain is correlated with many life events and psychological conditions. In this paper, we use a graph-based approach to represent the shape of the brain, including the shape of the ventricular system and shape relative to the skull. This graph representation is applied to classification of individuals based on level of cognitive impairment due to Alzheimer's Disease, level of education, and gender. The portions of the graph which are important to each distinction are found and visualized as an overlay on structural magnetic resonance images (MRI). We find that whole-brain analysis in this manner allows automatic classification of images based on gender if the whole brain is included, but not strictly based on the ventricular system. Alzheimer's Disease is found to strongly affect ventricle shape. Education is found to correlate with the shape of the medial longitudinal fissure and the Sylvian fissure, which may be due to increases in overall brain mass due to education. Gender is predicted primarily by information in the MRI regarding facial structure and head shape. Finally, age is found to be easier to classify than any of the above distinctions. The classifier is found to have 90.9% accuracy differentiating scans of individuals 40 and younger from those of individuals 60 or older. S. Seth Long, Lawrence B. Holder |
CIBCB | 2 |
| 2011 | Using Graphs to Improve Activity Prediction in Smart Environments Based on Motion Sensor Data
S. Seth Long, Lawrence B. Holder |
ICOST | 2 |
| 2011 | The COSE Ontology: Bringing the Semantic Web to Smart Environments
Zachary Wemlinger, Lawrence B. Holder |
ICOST | 2 |
| 2011 | Discovering Activities to Recognize and Track in a Smart EnvironmentabstractThe machine learning and pervasive sensing technologies found in smart homes offer unprecedented opportunities for providing health monitoring and assistance to individuals experiencing difficulties living independently at home. In order to monitor the functional health of smart home residents, we need to design technologies that recognize and track activities that people normally perform as part of their daily routines. Although approaches do exist for recognizing activities, the approaches are applied to activities that have been pre-selected and for which labeled training data is available. In contrast, we introduce an automated approach to activity tracking that identifies frequent activities that naturally occur in an individual's routine. With this capability we can then track the occurrence of regular activities to monitor functional health and to detect changes in an individual's patterns and lifestyle. In this paper we describe our activity mining and tracking approach and validate our algorithms on data collected in physical smart environments. Parisa Rashidi, Diane J. Cook, Lawrence B. Holder, Maureen Schmitter-Edgecombe |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2009 | Mining for insider threats in business transactions and processesabstractProtecting and securing sensitive information are critical challenges for businesses. Deliberate and intended actions such as malicious exploitation, theft or destruction of data, are not only harmful and difficult to detect, but frequently these threats are propagated by an insider. Unfortunately, current efforts to identify unauthorized access to information such as what is found in document control and management systems are limited in scope and capabilities. This paper presents an approach to detecting anomalies in business transactions and processes using a graph representation. In our graph-based anomaly detection (GBAD) approach, anomalous instances of structural patterns are discovered in data that represent entities, relationships and actions. A definition of graph-based anomalies and a brief description of the GBAD algorithms are presented, followed by empirical results using a discrete-event simulation of real-world business transactions and processes. William Eberle, Lawrence B. Holder |
CIDM | 2 |
| 2009 | Empirical comparison of graph classification algorithmsabstractThe graph classification problem is learning to classify separate, individual graphs in a graph database into two or more categories. A number of algorithms have been introduced for the graph classification problem. We present an empirical comparison of the major approaches for graph classification introduced in literature, namely, SubdueCL, frequent subgraph mining in conjunction with SVMs, walk-based graph kernel, frequent subgraph mining in conjunction with AdaBoost and DT-CLGBI. Experiments are performed on five real world data sets from the mutagenesis and predictive toxicology domain which are considered benchmark data sets for the graph classification problem. Additionally, experiments are performed on a corpus of artificial data sets constructed to investigate the performance of the algorithms across a variety of parameters of interest. Our conclusions are as follows. In datasets where the underlying concept has a high average degree, walk-based graph kernels perform poorly as compared to other approaches. The hypothesis space of the kernel is walks and it is insufficient at capturing concepts involving significant structure. In datasets where the underlying concept is disconnected, SubdueCL performs poorly as compared to other approaches. The hypothesis space of SubdueCL is connected graphs and it is insufficient at capturing concepts which consist of a disconnected graph. FSG+SVM, FSG+AdaBoost, DT-CLGBI have comparable performance in most cases. Nikhil S. Ketkar, Lawrence B. Holder, Diane J. Cook |
CIDM | 2 |
| 2009 | Faster computation of the direct product kernel for graph classificationabstractThe direct product kernel, introduced by Gartner et al. for graph classification, is based on defining a feature for every possible label sequence in a labelled graph and counting how many label sequences in two given graphs are identical. Although the direct product kernel has achieved promising results in terms of accuracy, the kernel computation is not feasible for large graphs. This is because computing the direct product kernel for two graphs is essentially computing either the inverse of or by diagonalizing the adjacency matrix of the direct product of these two graphs. For two graphs with adjacency matrices of sizes m and n, the adjacency matrix of their direct product graph can be of size mn in the worst case. As both matrix inversion or matrix diagonalizing in the general case is O(n3), computing the direct product kernel is O((mn)3). Our survey of data sets in graph classification indicates that most graphs have adjacency matrices of sizes in the order of hundreds which often leads to adjacency matrices of direct product graphs (of two graphs) having sizes in the order of thousands. In this work we show how the direct product kernel can be computed in O((m + n)3). The key insight behind our result is that the language of label sequences in a labeled graph is a regular language and that regular languages are closed under union and intersection. Nikhil S. Ketkar, Lawrence B. Holder, Diane J. Cook |
CIDM | 2 |
| 2009 | gRegress: Extracting Features from Graph Transactions for Regression
Nikhil S. Ketkar, Lawrence B. Holder, Diane J. Cook |
IJCAI | 2 |
| 2009 | Applying graph-based anomaly detection approaches to the discovery of insider threatsabstractThe ability to mine data represented as a graph has become important in several domains for detecting various structural patterns. One important area of data mining is anomaly detection, but little work has been done in terms of detecting anomalies in graph-based data. In this paper we present graph-based approaches to uncovering anomalies in applications containing information representing possible insider threat activity: e-mail, cell-phone calls, and order processing. William Eberle, Lawrence B. Holder |
ISI | 2 |
| 2009 | Learning patterns in the dynamics of biological networksabstractOur dynamic graph-based relational mining approach has been developed to learn structural patterns in biological networks as they change over time. The analysis of dynamic networks is important not only to understand life at the system-level, but also to discover novel patterns in other structural data. Most current graph-based data mining approaches overlook dynamic features of biological networks, because they are focused on only static graphs. Our approach analyzes a sequence of graphs and discovers rules that capture the changes that occur between pairs of graphs in the sequence. These rules represent the graph rewrite rules that the first graph must go through to be isomorphic to the second graph. Then, our approach feeds the graph rewrite rules into a machine learning system that learns general transformation rules describing the types of changes that occur for a class of dynamic biological networks. The discovered graph-rewriting rules show how biological networks change over time, and the transformation rules show the repeated patterns in the structural changes. In this paper, we apply our approach to biological networks to evaluate our approach and to understand how the biosystems change over time. We evaluate our results using coverage and prediction metrics, and compare to biological literature. Chang Hun You, Lawrence B. Holder, Diane J. Cook |
KDD | 2 |
| 2009 | Introduction to the special issue on homeland and global security
Lawrence B. Holder, Mohan Kumar, Raffaele Bruno 0001 |
Pervasive Mob. Comput. | 1 |
| 2008 | Temporal and structural analysis of biological networks in combination with microarray dataabstractWe introduce a graph-based relational learning approach using graph-rewriting rules for temporal and structural analysis of biological networks changing over time. The analysis of dynamic biological networks is necessary to understand life at the system-level, because biological networks continuously change their structures and properties, while an organism performs various biological activities. A dynamic graph represents dynamic properties as well as structural properties of biological networks. Microarray data can reflect dynamic properties of biological processes. Biological networks, which contain various molecules and relationships between molecules, show structural properties representing various relationships between entities. Most current graph-based data mining approaches overlook dynamic features of biological networks, because they are focused on only static graphs. Most approaches for analysis of microarray data disregard structural properties on biological systems. But our dynamic graph-based relational learning approach describes how the graphs temporally and structurally change over time in the dynamic graph representing biological networks in combination with microarray data. Chang Hun You, Lawrence B. Holder, Diane J. Cook |
CIBCB | 2 |
| 2008 | Game-Based Simulation for the Evaluation of Threat Detection in a Seaport Environment
Allen Christiansen, Damian Johnson, Lawrence B. Holder |
ICEC | 3 |
| 2008 | Strategic Path Planning on the Basis of Risk vs. Time
Ashish C. Singh, Lawrence B. Holder |
ICEC | 2 |
| 2007 | Anomaly detection in data represented as graphs
William Eberle, Lawrence B. Holder |
Intell. Data Anal. | 2 |
| 2007 | Inference of node replacement graph grammars
Jacek P. Kukluk, Lawrence B. Holder, Diane J. Cook |
Intell. Data Anal. | 2 |
| 2007 | Graph-Based Analysis of Human Transfer Learning Using a Game TestbedabstractThe ability to transfer knowledge learned in one environment in order to improve performance in a different environment is one of the hallmarks of human intelligence. Insights into human transfer learning help us to design computer-based agents that can better adapt to new environments without the need for substantial reprogramming. In this paper, we study the transfer of knowledge by humans playing various scenarios in a graphically realistic urban setting that are specifically designed to test various levels of transfer. We determine the amount and type of transfer that is being performed based on the performance of trained and untrained human players. In addition, we use a graph-based relational learning algorithm to extract patterns from player graphs. These analyses reveal that indeed humans are transferring knowledge from on 3 set of games to another and the amount and type of transfer varies according to player experience and scenario complexity. The results of this analysis help us understand the nature of human transfer in such environments and shed light on how we might endow computer-based agents with similar capabilities. The game simulator and human data collection also represent a significant testbed in which other Al capabilities can be tested and compared to human performance. Diane J. Cook, Lawrence B. Holder, G. Michael Youngblood |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2006 | Graph Grammar Induction on Structural Data for Visual ProgrammingabstractComputer programs that can be expressed in two or more dimensions are typically called visual programs. The underlying theories of visual programming languages involve graph grammars. As graph grammars are usually constructed manually, construction can be a time-consuming process that demands technical knowledge. Therefore, a technique for automatically constructing graph grammars - at least in part - is desirable. An induction method is given to infer node replacement graph grammars. The method operates on labeled graphs of broad applicability. It is evaluated by its performance on inferring graph grammars from various structural representations. The correctness of an inferred grammar is verified by parsing graphs not present in the training set Keven Ates, Jacek P. Kukluk, Lawrence B. Holder, Diane J. Cook |
ICTAI | 3 |
| 2006 | Detecting Anomalies in Cargo Using Graph Properties
William Eberle, Lawrence B. Holder |
ISI | 2 |
| 2006 | Classification of Threats Via a Multi-sensor Security Portal
Lawrence B. Holder |
ISI | 2 |
| 2006 | Inference of Node Replacement Recursive Graph GrammarsabstractIn this paper we describe an approach to learning node replacement graph grammars. This approach is based on previous research in frequent isomorphic subgraphs discovery. We extend the search for frequent subgraphs by checking for overlap among the instances of the subgraphs in the input graph. If subgraphs overlap by one node we propose a node replacement grammar production. We also can infer a hierarchy of productions by compressing portions of a graph described by a production and then infer new productions on the compressed graph. We validate this approach in experiments where we generate graphs from known grammars and measure how well our system infers the original grammar from the generated graph. Jacek P. Kukluk, Lawrence B. Holder, Diane J. Cook |
SDM | 2 |
| 2005 | A Learning Architecture for Automating the Intelligent Environment
G. Michael Youngblood, Diane J. Cook, Lawrence B. Holder |
AAAI | 3 |
| 2005 | Automation Intelligence for the Smart Environment
G. Michael Youngblood, Edwin O. Heierman III, Lawrence B. Holder, Diane J. Cook |
IJCAI | 3 |
| 2005 | Managing Adaptive Versatile EnvironmentsabstractThe goal of the MavHome project is to develop technologies to manage adaptive versatile environments. In this paper, we present a complete agent architecture for a single inhabitant intelligent environment and discuss the development, deployment, and techniques utilized in our working intelligent environments. Empirical evaluation of our approach has proven its effectiveness at reducing inhabitant interactions by 72.2% G. Michael Youngblood, Lawrence B. Holder, Diane J. Cook |
PerCom | 2 |
| 2005 | Seamlessly engineering a smart environmentabstractDeveloping technologies and systems for automated control of home and workplace environments is a challenging problem. We present a complete agent architecture for learning to automate a smart environment and discuss integration of AT and middleware technologies necessary to achieve the goals of this project. Results are demonstrated using the MavPad and MavLab intelligent environments. G. Michael Youngblood, Diane J. Cook, Lawrence B. Holder |
SMC | 3 |
| 2005 | Graph-based Relational Learning with Application to Security
Lawrence B. Holder, Diane J. Cook, Jeffrey Coble, Maitrayee Mukherjee |
Fundam. Informaticae | 1 |
| 2005 | Current And Future Trends In Feature Selection And Extraction For Classification ProblemsabstractIn this article, we describe some of the important currently used methods for solving classification problems, focusing on feature selection and extraction as parts of the overall classification task. We then go on to discuss likely future directions for research in this area, in the context of the other articles from this special issue. We propose that the next major step is the elaboration of a theory of how the methods of selection and extraction interact during the classification process for particular problem domains, along with any learning that may be part of the algorithms. Preferably this theory should be tested on a set of well-established benchmark challenge problems. Using this theory, we will be better able to identify the specific combinations that will achieve best classification performance for new tasks. Lawrence B. Holder, Ingrid Russell, Zdravko Markov, Anthony G. Pipe, Brian Carse |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2005 | Editorial
Ingrid Russell, Zdravko Markov, Brian Carse, Anthony G. Pipe, Lawrence B. Holder |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2005 | Managing Adaptive Versatile environments
G. Michael Youngblood, Diane J. Cook, Lawrence B. Holder |
Pervasive Mob. Comput. | 3 |
| 2003 | Using a Graph-Based Data Mining System to Perform Web SearchabstractThe World Wide Web provides an immense source of information. Accessing information of interest presents a challenge to scientists and analysts, particularly if the desired information is structural in nature. Our goal is to design a structural search engine that uses the hyperlink structure of the Web, in addition to textual information, to search for sites of interest. Our structural search engine, called WebSUBDUE, searches not only for particular words or topics but also for a desired hyperlink structure. Enhanced by WordNet text functions, our search engine retrieves sites corresponding to structures formed by graph-based user queries. We hypothesize that this system can form the heart of a structural query engine, and demonstrate the approach on a number of structural web queries. Diane J. Cook, Nitish Manocha, Lawrence B. Holder |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2002 | Graph-Based Relational Concept Learning
Jesus A. Gonzalez, Lawrence B. Holder, Diane J. Cook |
ICML | 2 |
| 2002 | Experimental Comparison of Graph-Based Relational Concept Learning with Inductive Logic Programming Systems
Jesus A. Gonzalez, Lawrence B. Holder, Diane J. Cook |
ILP | 2 |
| 2001 | Graph-Based Hierarchical Conceptual Clustering
Istvan Jonyer, Diane J. Cook, Lawrence B. Holder |
J. Mach. Learn. Res. | 3 |
| 2001 | Approaches to Parallel Graph-Based Knowledge Discovery
Diane J. Cook, Lawrence B. Holder, Gehad Galal, Ron Maglothin |
J. Parallel Distributed Comput. | 2 |
| 1999 | Experimentation-Driven Knowledge Acquisition for PlanningabstractKnowledge engineering for planning is expensive and the resulting knowledge can be imperfect. To autonomously learn a plan operator definition from environmental feedback, our learning system WISER explores an instantiated literal space using a breadth‐first search technique. Each node of the search tree represents a state, a unique subset of the instantiated literal space. A state at the root node is called a seed state. WISER can generate seed states with or without utilizing imperfect expert knowledge. WISER experiments with an operator at each node. The positive state, in which an operator can be successfully executed, constitutes initial preconditions of an operator. We analyze the number of required experiments as a function of the number of missing preconditions in a seed state. We introduce a naive domain assumption to test only a subset of the exponential state space. Since breadth‐first search is expensive, WISER introduces two search techniques to reorder literals at each level of the search tree. We demonstrate performance improvement using the naive domain assumption and literal‐ordering heuristics. To learn the effects of an operator, WISER computes the delta state, composed of the add list and the delete list, and parameterizes it. Unlike previous systems, WISER can handle unbound objects in the delta state. We show that machine‐generated effects definitions are often simpler in representation than expert‐provided definitions. Kang Soo Tae, Diane J. Cook, Lawrence B. Holder |
Comput. Intell. | 3 |
| 1999 | Knowledge discovery in molecular biology: Identifying structural regularities in proteinsabstractIn recent years, there has been an explosive amount of molecular biology information obtained and deposited in various databases. Identifying and interpreting interesting patterns from this massive amount of information has become an essential component in directing further molecular biology research. The goal of this research is to discover structural regularities in protein sequences by applying the SUBDUE discovery system to databases found in the Brookhaven Protein Data Bank. In this paper we discuss issues relevant to this application including data preparation and representation. We report on the results of applying SUBDUE to several classes of protein structures and discuss the potential significance of these results in the study of proteins. Shaobing Su, Diane J. Cook, Lawrence B. Holder |
Intell. Data Anal. | 3 |
| 1999 | Exploiting Parallelism in a Structural Scientific Discovery System to Improve ScalabilityabstractThe large amount of data collected today is quickly overwhelming researchers' abilities to interpret the data and discover interesting patterns. Knowledge discovery and data mining approaches hold the potential to automate the interpretation process, but these approaches frequently utilize computationally expensive algorithms. In particular, scientific discovery systems focus on the utilization of richer data representation, sometimes without regard for scalability. This research investigates approaches for scaling a particular knowledge discovery in databases (KDD) system, SUBDUE, using parallel and distributed resources. SUBDUE has been used to discover interesting and repetitive concepts in graph-based databases from a variety of domains, but requires a substantial amount of processing time. Experiments that demonstrate scalability of parallel versions of the SUBDUE system are performed using CAD circuit databases and artificially-generated databases, and potential achievements and obstacles are discussed. Gehad Galal, Diane J. Cook, Lawrence B. Holder |
J. Am. Soc. Inf. Sci. | 3 |
| 1997 | Improving Scalability in a Scientific Discovery System by Exploiting Parallelism
Gehad Galal, Diane J. Cook, Lawrence B. Holder |
KDD | 3 |
| 1997 | An Emprirical Study of Domain Knowledge and Its Benefits to Substructure DiscoveryabstractDiscovering repetitive, interesting, and functional substructures in a structural database improves the ability to interpret and compress the data. However, scientists working with a database in their area of expertise often search for predetermined types of structures or for structures exhibiting characteristics specific to the domain. The paper presents a method for guiding the discovery process with domain specific knowledge. The SUBDUE discovery system is used to evaluate the benefits of using domain knowledge to guide the discovery process. Domain knowledge is incorporated into SUBDUE following a single general methodology to guide the discovery process. Results show that domain specific knowledge improves the search for substructures that are useful to the domain and leads to greater compression of the data. To illustrate these benefits, examples and experiments from the computer programming, computer aided design circuit, and artificially generated domains are presented. Surnjani Djoko, Diane J. Cook, Lawrence B. Holder |
IEEE Trans. Knowl. Data Eng. | 3 |
| 1996 | Decision-Theoretic Cooperative Sensor PlanningabstractThis paper describes a decision-theoretic approach to cooperative sensor planning between multiple autonomous vehicles executing a military mission. For this autonomous vehicle application, intelligent cooperative reasoning must be used to select optimal vehicle viewing locations and select optimal camera pan and tilt angles throughout the mission. Decisions are made in such a way as to maximize the value of information gained by the sensors while maintaining vehicle stealth. Because the mission involves multiple vehicles, cooperation can be used to balance the work load and to increase information gain. This paper presents the theoretical foundations of our cooperative sensor planning research and describes the application of these techniques to ARPA's Unmanned Ground Vehicle program. Diane J. Cook, Piotr J. Gmytrasiewicz, Lawrence B. Holder |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1995 | Intermediate Decision Trees
Lawrence B. Holder |
IJCAI | 1 |
| 1995 | Analyzing the Benefits of Domain Knowledge in Substructure Discovery
Surnjani Djoko, Diane J. Cook, Lawrence B. Holder |
KDD | 3 |
| 1995 | Knowledge Discovery from Structural Data
Diane J. Cook, Lawrence B. Holder, Surnjani Djoko |
J. Intell. Inf. Syst. | 2 |
| 1994 | An Empirical Approach to Solving the General Utility Problem in Speedup Learning
Anurag Chaudhry, Lawrence B. Holder |
IEA/AIE | 2 |
| 1994 | Substructure Discovery Using Minimum Description Length and Background KnowledgeabstractThe ability to identify interesting and repetitive substructures is an essential component to discovering knowledge in structural data. We describe a new version of our SUBDUE substructure discovery system based on the minimum description length principle. The SUBDUE system discovers substructures that compress the original data and represent structural concepts in the data. By replacing previously-discovered substructures in the data, multiple passes of SUBDUE produce a hierarchical description of the structural regularities in the data. SUBDUE uses a computationally-bounded inexact graph match that identifies similar, but not identical, instances of a substructure and finds an approximate measure of closeness of two substructures when under computational constraints. In addition to the minimum description length principle, other background knowledge can be used by SUBDUE to guide the search towards more appropriate substructures. Experiments in a variety of domains demonstrate SUBDUE's ability to find substructures capable of compressing the original data and to discover structural concepts important to the domain. Description of Online Appendix: This is a compressed tar file containing the SUBDUE discovery system, written in C. The program accepts as input databases represented in graph form, and will output discovered substructures with their corresponding value. Diane J. Cook, Lawrence B. Holder |
J. Artif. Intell. Res. | 2 |
| 1993 | Discovery of Inexact Concepts from Structural DataabstractConcept discovery in structural data requires the identification of repetitive substructures in the data. A method for discovering substructures in data using an inexact graph match is described. An implementation of the authors' SUBDUE system that employs an inexact graph match to discover substructures which occur often in the data, but not always in the same form, is described. This inexact substructure discovery can be used to formulate fuzzy concepts, compress the data description, and discover interesting structures in data that are found either in an identical or in a slightly convoluted form. Examples from the domains of scene analysis and chemical compound analysis demonstrate the benefits of the inexact discovery technique.> Lawrence B. Holder, Diane J. Cook |
IEEE Trans. Knowl. Data Eng. | 1 |
| 1992 | Empirical Analysis of the General Utility Problem in Machine Learning
Lawrence B. Holder |
AAAI | 1 |
| 1992 | Fuzzy Substructure Discovery
Lawrence B. Holder, Diane J. Cook, Horst Bunke |
ML | 1 |
| 1990 | The General Utility Problem in Machine Learning
Lawrence B. Holder |
ML | 1 |
| 1989 | Empirical Substructure Discovery
Lawrence B. Holder |
ML | 1 |
| 1988 | Towards Intelligent Machine Learning Algorithms
Robert E. Stepp, Bradley L. Whitehall, Lawrence B. Holder |
ECAI | 3 |