Howard J. Hamilton

dblp:h/HowardJHamilton · also Howard John Hamilton · DBLP profile ↗
← Back
76ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0003-1475-0980ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 49 · 6 first-author · 3 since 2021Databases, data management, data science and information retrieval · 39 · 2 first-author · 2 since 2021Computer networks · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 2Security and privacy · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
YearPublicationVenuePosition
2026 An Early Conflict Resolution Mechanism for Blockchain-Based Delay-Sensitive IoT Networks
abstract
Blockchain technology, particularly Hyperledger Fabric (HLF), has emerged as a promising solution to enhance security and privacy in various domains, including Internet of Things (IoT) networks. Conflicting transactions in a HLF-based IoT network occur when multiple transactions attempt to modify the same asset or data concurrently. Conflicting transactions can lead to data inconsistencies, because the network may be unable to determine the correct order or the most preferred valid transaction. Existing conflict resolution mechanisms in HLF-based IoT networks often introduce considerable transaction latency, detect and resolve conflicting transactions in the late stages of the transaction lifecycle (ordering and validation), or require significant changes to the underlying HLF blockchain platform. To overcome these limitations, we propose an Early Conflict Resolution (ECR) mechanism that detects and resolves conflicts during the endorsement stage. The ECR mechanism uses a local cache (Sync.Map) and a dependency graph to efficiently detect conflicts by analyzing the Read-Sets (RS) and Write-Sets (WS) of transactions. ECR resolves conflicts in the detected conflicting transactions through transaction reordering or sequential processing. It also executes non-conflicting transactions in parallel to speed their processing. Our results show that the ECR mechanism improves transaction latency and the success rate for varying conflict rates, block sizes, and IoT devices compared to existing mechanisms.
Aditya Pathak, Irfan Al-Anbagi, Howard J. Hamilton
IEEE Trans. Netw. Serv. Manag.3
2025 Early-Stage Conflict Resolution Mechanism for HLF-Based Delay-Critical IoT Network
abstract
Conflicting transactions pose significant challenges in Hyperledger Fabric (HLF)-based IoT networks, affecting performance and introducing security vulnerabilities that can facilitate malicious attacks. Traditional conflict resolution mechanisms resolve conflicts in the later stages of transaction processing (i.e., the ordering or validation stages), resulting in increased transaction latency, which impacts delay-critical IoT applications. This paper proposes an Early Conflict Resolution (ECR) mechanism that integrates conflict detection and resolution at the endorsement stage, enhancing throughput and reducing transaction latency. This paper also explores the impact of conflicting transactions on blockchain attack vectors, focusing on four pivotal attacks-block withholding, double spending, balance attacks, and Distributed Denial-of-Service (DDoS)-simulated to analyze their exploitation of transaction conflicts and their impact on IoT networks. The results show that the ECR mechanism significantly improves the success rate and transaction latency compared to existing mechanisms.
Aditya Pathak, Irfan Al-Anbagi, Howard J. Hamilton
ICC3
2024 Privacy-Preserving Authentication Mechanism for P2P Energy Trading in Smart Grid Networks
abstract
Peer-to-Peer (P2P) energy trading, facilitated by prosumers who both produce and consume energy, provides a new type of for energy trading. Prosumers generate renewable energy in various environments, from industrial to residential. Traditional centralized energy trading methods pose risks, such as single point of failure and security issues. In contrast, de-centralized energy trading methods that use blockchains provide high security and reliability. However, the blockchain technology is not without limitations; in particular, the blockchain-based authentication mechanisms face three limitations, namely, they do not fully protect prosumer privacy due to unencrypted transactions, they are susceptible to multiple security attacks, and their authentication processes demand high computational and communication resources. To address these limitations, this paper proposes a novel Privacy-Preserving Mutual Authentication (PPMA) mechanism for P2P energy trading in smart grid networks. By employing Elliptic Curve Cryptography (ECC), symmetric encryption, and hash functions, the PPMA mechanism provides secure, privacy-preserving, and cost-effective mutual authentication for prosumers in P2P energy trading. When integrated with a permissioned blockchain and smart contract, PPMA aims to facilitate secure and scalable P2P energy trading. The efficacy of the PPMA mechanism is evaluated through comprehensive security and cost analyses.
Aditya Pathak, Irfan Al-Anbagi, Howard J. Hamilton
ICC3
2024 SATI: Sidechain-Based Access Control & Trust Mechanism for IoT Networks
abstract
Providing low latency, high security, and high resource utilization for Internet of Things (IoT) networks is challenging due to the heterogeneous nature of these networks and the need for more standardization in security algorithms. Current edge computing-based IoT solutions decrease network latency and improve resource utilization but do not provide adequate security because they offer multiple attack surfaces for adversaries. Recent work uses blockchain technology to provide better security in IoT networks. However, blockchain-based solutions suffer from scalability problems and can increase latency. Sidechains are parallel blockchain networks typically used to increase the scalability of blockchain networks. We propose a novel Sidechain-based Access control and Trust evaluation mechanism for IoT networks (SATI) to decrease network latency and improve scalability, security, and energy efficiency. SATI uses a sidechain with the blockchain network to improve its scalability. It also uses edge computing to provide low network latency and high resource utilization in terms of CPU and memory usage. In addition, trust evaluation and attribute-based access control mechanisms are used to improve the security of the IoT network. We compare our work with existing mechanisms in terms of scalability, security, latency, and CPU and memory usage. In addition, we perform a formal security analysis of the SATI mechanism using reduction-based analysis and the Scyther verification tool.
Aditya Pathak, Irfan Al-Anbagi, Howard J. Hamilton
IEEE Trans. Netw. Serv. Manag.3
2022 Efficient Removal of Weak Associations in Consensus Clustering
N. C. Ruckiya Sinorina, Howard J. Hamilton, Sandra Zilles
ICAART (3)2
2022 An Adaptive QoS and Trust-Based Lightweight Secure Routing Algorithm for WSNs
abstract
The limited resources and low computational power of wireless sensor networks (WSNs) make them vulnerable to various security attacks. Conventional security mechanisms require too many resources to allow the reliable operation of WSNs due to their resource-constrained nature. In addition, multihop communication in WSNs creates a requirement for guaranteed Quality of Service (QoS). Therefore, providing security while maintaining QoS and energy efficiency in WSNs are important design considerations. To further increase the performance of WSNs, there is a need to overcome the energy-hole problem, which leads to poor coverage of the field of interest. An energy-hole problem is created because of using poor deployment strategies. In this article, we define a multiobjective WSN optimization problem and present a novel algorithm known as lightweight secure routing (LSR) to manage WSNs that directly addresses the multiobjective WSN optimization problem. Our LSR algorithm uses ant colony optimization (ACO), an adaptive security model based on direct and indirect trust calculations, an adaptive QoS model, a hybrid deployment model based on 2-D Gaussian and uniform distributions, and an adaptive connectivity model that uses an appropriate communicational radius to ensure high connectivity between sensor nodes to solve the multiobjective WSN optimization problem. We divide our simulation results into three analyses, namely, trust model analysis, network scalability analysis, and security risk analysis to show that LSR outperforms the existing techniques in terms of energy consumed to calculate trust values, trust values convergence, network lifetime, average routing delay, and packet delivery ratio.
Aditya Pathak, Irfan Al-Anbagi, Howard J. Hamilton
IEEE Internet Things J.3
2021 Wildfire Occurrence Prediction Using Time Series Classification: A Comparative Study
abstract
We compare the effectiveness of four machine learning models at predicting wildfire occurrence from multivariate time series containing hourly weather data, vegetation data, and fire occurrence data. Strategies to improve performance on highly imbalanced datasets are investigated, including adapting KNN and HMM to consider cost effectiveness. Two different training regimes are compared: the imbalanced training regime varies the class imbalance in the training and testing datasets together, and the balanced training regime keeps the imbalance ratio 50:50 for every training dataset. FCN and ResNet outperform KNN and HMM across all class imbalances tested. We tested the methods on the SaskFire dataset, which is an extensive, new dataset describing wildfires in Saskatchewan, Canada. The two models that performed best on highly imbalanced datasets are FCN and ResNet trained with the imbalanced training regime. On our dataset with a non-fire to fire class imbalance of 99:1, FCN and ResNet have precisions of 0.190 and 0.250, respectively, and recalls of 0.800.
Ryan Laube, Howard J. Hamilton
IEEE BigData2
2021 Mining high utility patterns in interval-based event sequences
S. Mohammad Mirbagheri, Howard J. Hamilton
Data Knowl. Eng.2
2020 High-Utility Interval-Based Sequences
S. Mohammad Mirbagheri, Howard J. Hamilton
DaWaK2
2020 FIBS: A Generic Framework for Classifying Interval-Based Temporal Sequences
S. Mohammad Mirbagheri, Howard J. Hamilton
DaWaK2
2017 Real-Time Validation of Retail Gasoline Prices
Mondelle Simeon, Howard J. Hamilton
DS2
2017 Lossy Compression of Pattern Databases Using Acyclic Random Hypergraphs
abstract
A domain-independent heuristic function created by an abstraction is usually implemented using a Pattern Database (PDB), which is a lookup table of (abstract state, heuristic value) pairs. PDBs containing high quality heuristic values generally require substantial memory space and therefore need to be compressed. In this paper, we introduce Acyclic Random Hypergraph Compression (ARHC), a domain-independent approach to compressing PDBs using acyclic random r-partite r-uniform hypergraphs. The ARHC algorithm, which comes in Base and Extended versions, provides fast lookup and a high compression rate. ARHC-Extended achieves higher quality heuristics than ARHC-Base by decreasing the heuristic information loss at the cost of some decrease in the compression rate. ARHC shows higher performance than level-by-level Bloom filter PDB compression in all experiments conducted so far.
Mehdi Sadeqi, Howard J. Hamilton
IJCAI2
2016 Word Segmentation Algorithms with Lexical Resources for Hashtag Classification
abstract
We present a novel method for classifying hashtag types. Specifically, we apply word segmentation algorithms and lexical resources in order to classify two types of hashtags: those with sentiment information and those without. However, the complex structure of hashtags increases the difficulty of identifying sentiment information. In order to solve this problem, we segment hashtags into smaller semantic units using word segmentation algorithms in conjunction with lexical resources to classify hashtag types. Our experimental results demonstrate that our approach achieves a 14% increase in accuracy over baseline methods for identifying hashtags with sentiment information. Additionally, we achieve over 94% recall using this hashtag type for the subjectivity detection of tweets.
Credell Simeon, Howard J. Hamilton, Robert J. Hilderman
DSAA2
2015 Discovery of Parameters for Animation of Midge Swarms
Judith Bjorndahl, Ashley Herman, Richard Hamilton, Howard J. Hamilton, Mark Brigham
Discovery Science4
2014 Variable-sized, circular bokeh depth of field effects
Johannes Moersch, Howard J. Hamilton
Graphics Interface2
2013 Learning Models of Activities Involving Interacting Objects
Cristina E. Manfredotti, Kim Steenstrup Pedersen, Howard J. Hamilton, Sandra Zilles
IDA3
2012 A density-based spatial clustering for physical constraints
Xin Wang 0004, Camilo Rostoker, Howard J. Hamilton
J. Intell. Inf. Syst.3
2011 Simultaneous Tracking and Activity Recognition
abstract
Many tracking problems involve several distinct objects interacting with each other. We develop a framework that takes into account interactions between objects allowing the recognition of complex activities. In contrast to classic approaches that consider distinct phases of tracking and activity recognition, our framework performs these two tasks simultaneously. In particular, we adopt a Bayesian standpoint where the system maintains a joint distribution of the positions, the interactions and the possible activities. This turns out to be advantegeous, as information about the ongoing activities can be used to improve the prediction step of the tracking, while, at the same time, tracking information can be used for online activity recognition. Experimental results in two different settings show that our approach 1) decreases the error rate and improves the identity maintenance of the positional tracking and 2) identifies the correct activity with higher accuracy than standard approaches.
Cristina E. Manfredotti, David J. Fleet, Howard J. Hamilton, Sandra Zilles
ICTAI3
2010 An ontology-based framework for geospatial clustering
abstract
Geospatial clustering is an important topic in knowledge discovery research and geospatial information systems. However, current clustering research emphasizes the development of more efficient and effective clustering methods without paying much attention to domain knowledge and users' goals during the clustering process. Making better use of geospatial and clustering knowledge to select proper methods and datasets will help achieve clustering results that better meet users' requirements. In this article, we present the GEO_CLUST framework for performing geospatial clustering. The framework consists of the GeoCO ontology for geospatial clustering and the ontology reasoner reasoning mechanism. The GeoCO ontology is used to represent geospatial and clustering domain knowledge. The ontology reasoner uses classification and decomposition techniques to specify users' tasks. Using the framework, users can identify the appropriate geospatial data and clustering method based on their specific goals. To demonstrate the framework, two case studies on finding population density clusters in Western Canada and locating five hospitals in South Carolina are discussed. The results show that the framework can select the proper datasets and clustering methods with respect to users' goals.
Xin Wang 0004, Danielle Ziébelin, Howard J. Hamilton
Int. J. Geogr. Inf. Sci.4
2009 The Multi-Tree Cubing algorithm for computing iceberg cubes
abstract
The computation of data cubes is one of the most expensive operations in on-line analytical processing (OLAP). To improve efficiency, an iceberg cube represents only the cells whose aggregate values are above a given threshold (minimum support). Top-down and bottom-up approaches are used to compute the iceberg cube for a data set, but both have performance limitations. In this paper, a new algorithm, called Multi-Tree Cubing (MTC), is proposed for computing an iceberg cube. The Multi-Tree Cubing algorithm is an integrated top-down and bottom-up approach. Overall control is handled in a top-down manner, so MTC features shared computation. By processing the orderings in the opposite order from the Top-Down Computation algorithm, the MTC algorithm is able to prune attributes. The Bottom Up Computation (BUC) algorithm and its variations also perform pruning by relying on the processing of intermediate partitions. The MTC algorithm, however, prunes without processing such partitions. The MTC algorithm is based on a specialized type of prefix tree data structure, called an Attribute–Partition tree (AP-tree), consisting of attribute and partition nodes. The AP-tree facilitates fast, in-memory sorting and APRIORI-like pruning. We report on five series of experiments, which confirm that MTC is consistently as fast or faster than BUC, while finding the same iceberg cubes.
Howard J. Hamilton, Kamran Karimi, Liqiang Geng
J. Intell. Inf. Syst.2
2008 Mining functional dependencies from data
Hong Yao, Howard J. Hamilton
Data Min. Knowl. Discov.2
2007 Expectation Propagation in GenSpace Graphs for Summarization
Liqiang Geng, Howard J. Hamilton, Larry Korba
DaWaK2
2006 Searching for Pattern Rules
abstract
We address the problem of finding a set of pattern rules, from a transaction dataset given a statistical metric. A new data structure, called an incrementally counting suffix tree (ICST), is proposed for online computation of estimates of the support of any pattern or itemset. Using an ICST, our approach directly generates a set of pattern rules by a single scan of the whole dataset in partitions without the generation of frequent itemsets. Non-redundant rules can be found by removing redundancies from the pattern rules. The PPMCR algorithm first finds pattern rules and then non-redundant rules by generating valid candidates while traversing the ICST. Experimental results show that the PPMCR algorithm can be used for efficiently mining fewer non-redundant rules.
Guichong Li, Howard J. Hamilton
ICDM2
2006 The PDD Framework for Detecting Categories of Peculiar Data
abstract
Peculiar data are objects that are relatively few in number and significantly different from the other objects in a data set. In this paper, we propose the PDD framework for detecting multiple categories of peculiar data. This framework provides an extensible set of perspectives for viewing data, currently including viewing data as a set of records, attributes, frequencies, intervals, sequences, or sequences of changes. By using these six views of the data, multiple categories of peculiar data can be detected to reveal different aspects of the data. For each view, the framework provides an extensible set of peculiarity measures to detect outliers and other kinds of peculiar data. The PDD framework has been implemented for Oracle and Access. Experiments are reported for data sets concerning Regina weather and NHL hockey.
Mahesh Shrestha, Howard J. Hamilton, Yiyu Yao, Ken Konkel, Liqiang Geng
ICDM2
2006 Mining itemset utilities from transaction databases
Hong Yao, Howard J. Hamilton
Data Knowl. Eng.2
2005 The TIMERS II Algorithm for the Discovery of Causality
Howard J. Hamilton, Kamran Karimi
PAKDD1
2005 A machine-discovery approach to the evaluation of hashing techniques
abstract
This paper, describes an inference technique based on machine discovery for drawing conclusions from experimental results. Given access to the results of a full-factorial experiment, the inference technique finds three types of empirical generalizations. First, the best and worst values for each independent attribute, in terms of their effect on the dependent attribute, are identified. Second, direct and inverse relationships are found by applying regression to rank frequencies. Finally, cases where restricting a variable to a single value yields different behaviour from usual are identified. These three types of generalizations are produced in the form of English sentences. Experimental results using a Prolog implementation indicate that the inference technique finds many of the same generalizations as human researchers did in a fundamental study of the performance of hashing techniques.
Howard J. Hamilton, Demyen Doug
J. Exp. Theor. Artif. Intell.1
2004 Density-Based Spatial Clustering in the Presence of Obstacles and Facilitators
Xin Wang 0004, Camilo Rostoker, Howard J. Hamilton
PKDD3
2004 Basic Association Rules
abstract
Previous approaches for mining association rules generate large sets of association rules. Such sets are difficult for users to understand and manage. Here, the concept of a restricted conditional probability distribution is used to explain an association rule. Based on this concept, a new type of association rules, called basic association rules, is defined. We propose the GenBR algorithm to generate the set of classes of basic association rules. Theoretical analysis shows that the search space of the algorithm can be translated to an n-cube graph. The set of classes of basic association rules generated by GenBR is easy for users to understand and manage. Our experiments on synthetic and real datasets show that GenBR is either faster than previous approaches or generates fewer rules or both.
Guichong Li, Howard J. Hamilton
SDM2
2004 A Foundational Approach to Mining Itemset Utilities from Databases
abstract
Most approaches to mining association rules implicitly consider the utilities of the itemsets to be equal. We assume that the utilities of itemsets may differ, and identify the high utility itemsets based on information in the transaction database and external information about utilities. Our theoretical analysis of the resulting problem lays the foundation for future utility mining algorithms.
Hong Yao, Howard J. Hamilton, Cory J. Butz
SDM2
2003 Distinguishing Causal and Acausal Temporal Relations
Kamran Karimi, Howard J. Hamilton
PAKDD2
2003 DBRS: A Density-Based Spatial Clustering Method with Random Sampling
Xin Wang 0004, Howard J. Hamilton
PAKDD2
2003 Spatio-Temporal Data Mining with Expected Distribution Domain Generalization Graphs
abstract
We describe a method for spatio-temporal data mining based on expected distribution domain generalization (ExGen) graphs. Using familiar calendar and geographical concepts, such as workdays, weeks, climatic regions, and countries, spatio-temporal data can be aggregated into summaries in many ways. We automatically search for a summary with a distribution that is anomalous, i.e., far from user expectations. We repeatedly ranked possible summaries according to current expectations, and then allow the user to adjust these expectations.
Howard J. Hamilton, Liqiang Geng, Leah Findlater, Dee Jay Randall
TIME1
2003 Extracting Share Frequent Itemsets with Infrequent Subsets
Brock Barber, Howard J. Hamilton
Data Min. Knowl. Discov.2
2003 Iceberg-cube algorithms: An empirical evaluation on synthetic and real data
Leah Findlater, Howard J. Hamilton
Intell. Data Anal.2
2003 Measuring the interestingness of discovered knowledge: A principled approach
Robert J. Hilderman, Howard J. Hamilton
Intell. Data Anal.2
2002 ESRS: A Case Selection Algorithm Using Extended Similarity-based Rough Sets
abstract
A case selection algorithm selects representative cases from a large data set for future case-based reasoning tasks. This paper proposes the ESRS algorithm, based on extended similarity-based rough set theory, which selects a reasonable number of the representative cases while maintaining satisfactory classification accuracy. It also can handle noise and inconsistent data. Experimental results on synthetic and real sets of cases showed that its predictive accuracy is similar to that of well-known machine learning systems on standard data sets, while it has the advantage of being applicable to any data set where a similarity function can be defined.
Liqiang Geng, Howard J. Hamilton
ICDM2
2002 FD_Mine: Discovering Functional Dependencies in a Database Using Equivalences
abstract
The discovery of FDs from databases has recently become a significant research problem. In this paper, we propose a new algorithm, called FD-Mine. FD-Mine takes advantage of the rich theory of FDs to reduce both the size of the dataset and the number of FDs to be checked by using discovered equivalences. We show that the pruning does not lead to loss of information. Experiments on 15 UCI datasets show that FD-Mine can prune more candidates than previous methods.
Hong Yao, Howard J. Hamilton, Cory J. Butz
ICDM2
2002 TimeSleuth: A Tool for Discovering Causal and Temporal Rules
abstract
Discovering causal and temporal relations in a system is essential to understanding how it works, and to learning to control the behaviour of the system. TimeSleuth is a causality miner that uses association relations as the basis for the discovery of causal and temporal relations. It does so by introducing time into the observed data. TimeSleuth uses C4.5 as its association discoverer, and by using a series of preprocessing and post-processing techniques to enable the user to try different scenarios for mining causality. The data to be mined should originate sequentially from a single system. TimeSleuth's use of a standard decision tree builder such as C4.5 puts it outside the current mainstream method of discovering causality, which is based on conditional independencies and causal Bayesian networks. This paper introduces TimeSleuth as a tool, and describes its functionality. It is an unsupervised tool that can handle and interpret temporal data. It also helps the user in analyzing the relationships among the attributes. There is also a mechanism to distinguish between causality and acausal relations. The user is thus encouraged to perform experiments and discover the nature of relationships among the data.
Kamran Karimi, Howard J. Hamilton
ICTAI2
2002 Discovering Temporal Rules from Temporally Ordered Data
Kamran Karimi, Howard J. Hamilton
IDEAL2
2002 Design of knowledge-based systems with the ontology-domain-system approach
abstract
An ontology is a comprehensive knowledge model that enables a developer to practice a "higher" level of reuse, namely knowledge reuse. To achieve knowledge reuse instead of software reuse, we propose forging a closer mapping between the knowledge and software models in the development process. In this paper, we first present UML as an ontology modeling language and then describe the Ontology-Domain-System approach to deriving a system model from a UML-based ontology model.
Xin Wang 0004, Christine W. Chan, Howard J. Hamilton
SEKE3
2001 Evaluation of Interestingness Measures for Ranking Discovered Knowledge
Robert J. Hilderman, Howard J. Hamilton
PAKDD2
2001 Parametric Algorithms for Mining Share Frequent Itemsets
Brock Barber, Howard J. Hamilton
J. Intell. Inf. Syst.2
2000 Principles for mining summaries using objective measures of interestingness
abstract
An important problem in the area of data mining is the development of effective measures of interestingness for ranking discovered knowledge. The authors propose five principles that any measure must satisfy to be considered useful for ranking the interestingness of summaries generated from databases. We investigate the problem within the context of summarizing a single dataset which can be generalized in many different ways and to many levels of granularity. We perform a comparative sensitivity analysis of fifteen well-known diversity measures to identify those which satisfy the proposed principles. The fifteen diversity measures have previously been utilized in various disciplines, such as information theory, statistics, ecology, and economics. Their use as objective measures of interestingness for ranking summaries generated from databases is novel. The objective of this work is to gain some insight into the behaviour that can be expected from each of the diversity measures in practice, and to begin to develop a theory of interestingness against which the utility of new measures can be assessed.
Robert J. Hilderman, Howard J. Hamilton
ICTAI2
2000 Support based measures applied to ice hockey scoring summaries
abstract
We present the hockey line extraction (HLE) algorithm, which examines ice hockey scoring summaries in an attempt to determine a team's lines. The players on a hockey team are divided into units called "lines" that appear together on the ice. The HLE algorithm uses single link clustering, support based measures and positional information to identify lines of players. The hockey lines software enables users to view relationships between players on a team based on time period and either support or confidence.
Bradley P. Kram, James A. Hall, Howard J. Hamilton
ICTAI3
2000 Logical Decision Rules: Teaching C4.5 to Speak Prolog
Kamran Karimi, Howard J. Hamilton
IDEAL2
2000 Parametric Algorithms for Mining Share-Frequent Itemsets
Brock Barber, Howard J. Hamilton
ISMIS2
2000 Finding Temporal Relations: Causal Bayesian Networks vs. C4.5
Kamran Karimi, Howard J. Hamilton
ISMIS2
2000 Algorithms for Mining Share Frequent Itemsets Containing Infrequent Subsets
Brock Barber, Howard J. Hamilton
PKDD2
2000 Applying Objective Interestingness Measures in Data Mining Systems
Robert J. Hilderman, Howard J. Hamilton
PKDD2
1999 Heuristic Selection of Aggregated Temporal Data for Knowledge Discovery
Howard J. Hamilton, Dee Jay Randall
IEA/AIE1
1999 Learning English Grapheme Segmentation Using the Iterated Version Space Algorithm
Jianna Jian Zhang, Howard J. Hamilton, Nick Cercone
ISMIS2
1999 Heuristic for Ranking the Interestigness of Discovered Knowledge
Robert J. Hilderman, Howard J. Hamilton
PAKDD2
1999 Heuristic Measures of Interestingness
abstract
The tuples in a generalized relation (i.e., a summary generated from a database) are unique, and therefore, can be considered to be a population with a structure that can be described by some probability distribution. In this paper, we present and empirically compare sixteen heuristic measures that evaluate the structure of a summary to assign a single real-valued index that represents its interestingness relative to other summaries generated from the same database. The heuristics are based upon well-known measures of diversity, dispersion, dominance, and inequality used in several areas of the physical, social, ecological, management, information, and computer sciences. Their use for ranking summaries generated from databases is a new application area. All sixteen heuristics rank less complex summaries (i.e., those with few tuples and/or few non-ANY attributes) as most interesting. We demonstrate that for sample data sets, the order in which some of the measures rank summaries is highly correlated.
Robert J. Hilderman, Howard J. Hamilton
PKDD2
1999 Determining the incremental worth of members of an aggregate set through difference-based induction
abstract
Calculating the incremental worth or weight of individual components of an aggregate set when only the whole set's total worth or weight is known is a problem common to several domains. Here we describe an algorithm that induces such incremental worth from a database of similar (not identical) aggregate sets. The algorithm focuses on finding aggregate sets in the database exhibiting minimal differences in corresponding components (attributes and values). This procedure isolates dissimilarities between nearly similar aggregate sets so any difference in worth between sets is attributed to them. The algorithm builds a classification tree similar to those of ID3 and C4.5 distributes all aggregate sets in the database according to their attributes and values; and groups together those with the same attributes and values. Each leaf of the classification tree then contains a group of aggregate sets identical to each other insofar as their attribues and values. Groups' members belonging to two sibling leaves (having the same immediate parent) differ from each other in the value of exactly one attribute. Thus, any difference in the worth of sets in those groups can be attributed to that difference. The worth of the aggregate sets in these groups can be averaged when data are noisy. This algorithm works well when applied to real-estate appraisal domain. ©1999 John Wiley & Sons, Inc.
Avelino J. Gonzalez, Sylvia Daroszewski, Howard J. Hamilton
Int. J. Intell. Syst.3
1999 Temporal Generalization with Domain Generalization Graphs
abstract
This paper addresses the problem of using domain generalization graphs to generalize temporal data extracted from relational databases. A domain generalization graph associated with an attribute defines a partial order which represents a set of generalization relations for the attribute. We propose formal specifications for domain generalization graphs associated with calendar (date and time) attributes. These graphs are reusable (i.e. can be used to generalize any calendar attributes), adaptable (i.e. can be extended or restricted as appropriate for particular applications), and transportable (i.e. can be used with any database containing a calendar attribute).
Dee Jay Randall, Howard J. Hamilton, Robert J. Hilderman
Int. J. Pattern Recognit. Artif. Intell.2
1999 Data Mining in Large Databases Using Domain Generalization Graphs
Robert J. Hilderman, Howard J. Hamilton, Nick Cercone
J. Intell. Inf. Syst.2
1998 Mining Market Basket Data Using Share Measures and Characterized Itemsets
Robert J. Hilderman, Colin L. Carter, Howard J. Hamilton, Nick Cercone
PAKDD3
1998 Generalization Lattices
Howard J. Hamilton, Robert J. Hilderman, Liangchun Li, Dee Jay Randall
PKDD1
1998 Efficient Attribute-Oriented Generalization for Knowledge Discovery from Large Databases
abstract
We present GDBR (Generalize DataBase Relation) and FIGR (Fast, Incremental Generalization and Regeneralization), two enhancements of Attribute Oriented Generalization, a well known knowledge discovery from databases technique. GDBR and FIGR are both O(n) and, as such, are optimal. GDBR is an online algorithm and requires only a small, constant amount of space. FIGR also requires a constant amount of space that is generally reasonable, although under certain circumstances, may grow large. FIGR is incremental, allowing changes to the database to be reflected in the generalization results without rereading input data. FIGR also allows fast regeneralization to both higher and lower levels of generality without rereading input. We compare GDBR and FIGR to two previous algorithms, LCHR and AOI, which are O(n log n) and O(np), respectively, where n is the number of input tuples and p the number of tuples in the generalized relation. Both require O(n) space that, for large input, causes memory problems. We implemented all four algorithms and ran empirical tests, and we found that GDBR and FIGR are faster. In addition, their runtimes increase only linearly as input size increases, while the runtimes of LCHR and AOI increase greatly when input size exceeds memory limitations.
Colin L. Carter, Howard J. Hamilton
IEEE Trans. Knowl. Data Eng.2
1997 Inducing and Using Decision Rules in the GRG Knowledge Discovery System
Ning Shan, Howard J. Hamilton, Nick Cercone
ECML2
1997 Data Visualization in the DB-Discover System
abstract
The DB-Discover system is a research software tool for knowledge discovery from databases. Utilizing an efficient attribute oriented induction algorithm based upon the climbing tree and dropping condition methods, called multi attribute generalization, it generates many possible summaries of data from a database. We present the design of the data visualization capabilities of DB-Discover and describe how these capabilities help us manage the summaries generated.
Robert J. Hilderman, Liangchun Li, Howard J. Hamilton
ICTAI3
1997 A Comparison of Attribute Selection Strategies for Attribute-Oriented Generalization
Brock Barber, Howard J. Hamilton
ISMIS2
1997 Learning English Syllabification for Words
Howard J. Hamilton
ISMIS2
1997 Share Based Measures for Itemsets
Colin L. Carter, Howard J. Hamilton, Nick Cercone
PKDD2
1997 Parallel Knowledge Discovery Using Domain Generalization Graphs
Robert J. Hilderman, Howard J. Hamilton, Robert J. Kowalchuk, Nick Cercone
PKDD2
1997 A Note on Regeneration with Virtual Copies
abstract
Regeneration with virtual copies (RVC) is a voting-based consistency control algorithm for replicated data objects in a distributed computing system. Proposed by Adam and Tewari (ibid., vol. 19, no. 6, pp. 594-602, 1993), it utilizes selective regeneration and recovery mechanisms for maintaining the availability and consistency of copies. This paper describes some problems with the original paper and proposes solutions.
Robert J. Hilderman, Howard J. Hamilton
IEEE Trans. Software Eng.2
1996 Attribute-oriented Induction Using Domain Generalization Graphs
abstract
Attribute-oriented induction summarizes the information in a relational database by repeatedly replacing specific attribute values with more general concepts according to user-defined concept hierarchies. We show how domain generalization graphs can be constructed from multiple concept hierarchies associated with an attribute, describe how these graphs can be used to control the generalization of a set of attributes, and present the Multi-Attribute Generalization algorithm for attribute-oriented induction using domain generalization graphs. Based upon a generate-and-test approach, the algorithm generates all possible combinations of nodes from the domain generalization graphs associated with the individual attributes, to produce all possible generalized relations for the set of attributes. We rant the interestingness of the resulting generalized relations using measures based upon relative entropy and variance. Our experiments show that these measures provide a basis for analyzing summary data from relational databases. Variance appears more useful because it tends to rank the less complex generalized relations (i.e., those with few attributes and/or few tuples) as more interesting.
Howard J. Hamilton, Robert J. Hilderman, Nick Cercone
ICTAI1
1996 Induction of Classification Rules from Imperfect Data
Ning Shan, Howard J. Hamilton, Nick Cercone
ISMIS2
1996 Discovering Classification Knowledge in Databases Using Rough Sets
Ning Shan, Wojciech Ziarko, Howard J. Hamilton, Nick Cercone
KDD3
1996 It's About Time: An Introduction to the Special Issue on Temporal Representation and Reasoning
Scott D. Goodwin, Howard J. Hamilton
Comput. Intell.2
1995 Performance evaluation of attribute-oriented algorithms for knowledge discovery from databases
abstract
Practical tools for knowledge discovery from databases must be efficient enough to handle large data sets found in commercial environments. Attribute-oriented induction has proved to be a useful method for knowledge discovery. Three algorithms are AOI, LCHR and GDBR. We have implemented efficient versions of each algorithm and empirically compared them on large commercial data sets. These tests show that GDBR is consistently faster than AOI and LCHR. GDBR's times increase linearly with increased input size, while times for AOI and LCHR increase non-linearly when memory is exceeded. Through better memory management, however, AOI can be improved to provide some advantages.
Colin L. Carter, Howard J. Hamilton
ICTAI2
1995 GRG: knowledge discovery using information generalization, information reduction, and rule generation
abstract
We present the three-step GRG approach for learning decision rules from large relational databases. In the first step, an attribute-oriented concept tree ascension technique is applied to generalize an information system. This step loses some information but substantially improves the efficiency of the following steps. In the second step, the reduction technique is applied to generate a minimized information system called a reduct which contains a minimal subset of the generalized attributes and the smallest number of distinct tuples for those attributes. Finally, a set of maximally general rules are derived directly from the reduct. These rules can be used to interpret and understand the active mechanisms underlying the database.
Ning Shan, Howard J. Hamilton, Nick Cercone
ICTAI2
1995 Using Rough Sets as Tools for Knowledge Discovery
Ning Shan, Wojciech Ziarko, Howard J. Hamilton, Nick Cercone
KDD3
1995 Performance Analysis of a Regeneration-Based Dynamic Voting Algorithm
abstract
RVC2 is a consistency control algorithm for replicated data objects in a distributed computing system. It is a dynamic voting algorithm which utilizes selective regeneration and recovery mechanisms for failed copies. Virtual copies which record information about the current state of a data object, but do not contain actual data, are used to reduce network and storage overhead. Experimental results for availability, storage cost, and message cost, obtained through simulation, are discussed. Our results show that the replacement of real copies with virtual copies has no significant impact on the availability of a data object. Neither does varying the generation threshold. We also show that high availability can be maintained without regeneration. We conclude that regeneration makes no significant contribution to the high availability of RVC2.
Robert J. Hilderman, Howard J. Hamilton
SRDS2
1995 Estimating DBLEARN's Potential for Knowledge Discovery in Databases
abstract
We propose a procedure for estimating DBLEARN's potential for knowledge discovery, given a relational database and concept hierarchies. This procedure is most useful for evaluating alternative concept hierarchies for the same database. The DBLEARN knowledge discovery program uses an attribute‐oriented inductive‐inference method to discover potentially significant high‐level relationships in a database. A concept forest, with at most one concept hierarchy for each attribute, defines the possible generalizations that DBLEARN can make for a database. The potential for discovery in a database is estimated by examining the complexity of the corresponding concept forest. Two heuristic measures are defined based on the number, depth, and height of the interior nodes. Higher values for these measures indicate more complex concept forests and arguably more potential for discovery. Experimental results using a variety of concept forests and four commercial databases show that in practice both measures permit quite reliable decisions to be made; thus, the simplest may be most appropriate.
Howard J. Hamilton, David R. Fudger
Comput. Intell.1