Nick Cercone

dblp:c/NickCercone · DBLP profile ↗
← Back
110ranked-venue papers
10as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 63 · 5 first-authorDatabases, data management, data science and information retrieval · 49 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 10Software engineering, systems software and programming languages · 6Human-computer interaction and ubiquitous computing · 5 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-authorSecurity and privacy · 2Systems, architecture and hardware · 1Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
16 papers
Information retrieval · 68% Data mining · 16% Query processing and optimization · 12%
Artificial intelligence
8 papers
Knowledge representation and reasoning · 66% Information extraction and text analysis · 22% Language models and text generation · 12%

Topics — the 30 heaviest of 47, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › ranking › text ranking
answer ranking
0.212015
Meaningful keyword search in relational databases with large and complex schema · ICDE 2015
Query processing and optimization
keyword query processing
0.212015
Meaningful keyword search in relational databases with large and complex schema · ICDE 2015
Information retrieval › keyword search
keyword search over relational databases
0.212015
Meaningful keyword search in relational databases with large and complex schema · ICDE 2015
Information retrieval › evaluation › effectiveness metrics
relevance measure
0.212015
Meaningful keyword search in relational databases with large and complex schema · ICDE 2015
Information retrieval
keyword search
0.212014
MeanKS: meaningful keyword search in relational databases with complex schema · SIGMOD Conference 2014
Information retrieval › keyword search
keyword search in relational databases
0.212014
MeanKS: meaningful keyword search in relational databases with complex schema · SIGMOD Conference 2014
Information retrieval
ranking
0.212014
MeanKS: meaningful keyword search in relational databases with complex schema · SIGMOD Conference 2014
Data mining › pattern mining
association rule mining
0.122002
Discovery of Interesting Association Rules from Livelink Web Log Data · ICDM 2002
Mining Knowledge Rules from Databases: A Rough Set Approach · ICDE 1996
Data mining › predictive modeling › classification › rule learning
classification rule learning
0.032000
RuleViz: a model for visualizing knowledge discovery process · KDD 2000
Discovering Classification Knowledge in Databases Using Rough Sets · KDD 1996
An Attribute-Oriented Approach for Learning Classification Rules from Relational Databases · ICDE 1990
Data mining › pattern mining
rule mining
0.032000
RuleViz: a model for visualizing knowledge discovery process · KDD 2000
DBLearn: A System Prototype for Knowledge Discovery in Relational Databases · SIGMOD Conference 1994
Data-Driven Discovery of Quantitative Rules in Relational Databases · IEEE Trans. Knowl. Data Eng. 1993
Information retrieval › cross-language information retrieval
chinese text retrieval
0.012002
Using self-supervised word segmentation in Chinese information retrieval · SIGIR 2002
Information retrieval
cross-language information retrieval
0.012002
Using self-supervised word segmentation in Chinese information retrieval · SIGIR 2002
Information retrieval
indexing
0.012002
Using self-supervised word segmentation in Chinese information retrieval · SIGIR 2002
Data mining › pattern mining › association rule mining
interestingness measures
0.012002
Discovery of Interesting Association Rules from Livelink Web Log Data · ICDM 2002
Data mining
pattern mining
0.012002
Discovery of Interesting Association Rules from Livelink Web Log Data · ICDM 2002
Data mining
attribute-oriented induction
0.041996
Data-Driven Discovery of Quantitative Rules in Relational Databases · IEEE Trans. Knowl. Data Eng. 1993
Knowledge Discovery in Databases: An Attribute-Oriented Approach · VLDB 1992
An Attribute-Oriented Approach for Learning Classification Rules from Relational Databases · ICDE 1990
Data mining › granular computing
rough set theory
0.021996
Mining Knowledge Rules from Databases: A Rough Set Approach · ICDE 1996
Rough Sets Similarity-Based Learning from Databases · KDD 1995
Knowledge, reasoning and agents › Knowledge representation and reasoning
case-based reasoning
0.011999
Rule-Induction and Case-Based Reasoning: Hybrid Architectures Appear Advantageous · IEEE Trans. Knowl. Data Eng. 1999
Knowledge, reasoning and agents › Knowledge representation and reasoning
rule learning
0.011999
Rule-Induction and Case-Based Reasoning: Hybrid Architectures Appear Advantageous · IEEE Trans. Knowl. Data Eng. 1999
Natural language and speech › Information extraction and text analysis
slot filling
0.012006
Segment-Based Hidden Markov Models for Information Extraction · ACL 2006
Knowledge graphs
concept hierarchy
0.011996
Intelligent Query Answering by Knowledge Discovery Techniques · IEEE Trans. Knowl. Data Eng. 1996
Data mining
similarity-based learning
0.011995
Rough Sets Similarity-Based Learning from Databases · KDD 1995
Web and social media mining
web usage mining
0.012002
Discovery of Interesting Association Rules from Livelink Web Log Data · ICDM 2002
Natural language and speech › Language models and text generation
text generation
0.011993
Decision-Theoretic Salience Interactions in Language Generation · IJCAI 1993
Data integration and cleaning
missing data
0.021988
Providing Quality Responses with Natural Language Interfaces: The Null Value Problem · IEEE Trans. Software Eng. 1988
What do you mean "Null"? Turning Null Responses into Quality Responses · ICDE 1987
Data mining › structured data mining
relational data mining
0.011990
An Attribute-Oriented Approach for Learning Classification Rules from Relational Databases · ICDE 1990
Data models and query languages › relational model
extended relational model
0.021988
What do you mean "Null"? Turning Null Responses into Quality Responses · ICDE 1987
Providing Quality Responses with Natural Language Interfaces: The Null Value Problem · IEEE Trans. Software Eng. 1988
Data models and query languages › natural language interface
natural language interface to database
0.011988
Providing Quality Responses with Natural Language Interfaces: The Null Value Problem · IEEE Trans. Software Eng. 1988
Knowledge, reasoning and agents › Knowledge representation and reasoning › ontology
concept hierarchy
0.021993
Data-Driven Discovery of Quantitative Rules in Relational Databases · IEEE Trans. Knowl. Data Eng. 1993
An Attribute-Oriented Approach for Learning Classification Rules from Relational Databases · ICDE 1990
Data integration and cleaning
data preprocessing
0.011996
Mining Knowledge Rules from Databases: A Rough Set Approach · ICDE 1996

Methods — techniques the papers use, named apart from their topics

user study · 0.2information content measure · 0.2segment retrieval · 0.1hidden markov model · 0.1parallel coordinates · 0.1rough sets · 0.0self-supervised learning · 0.0attribute-oriented induction · 0.0association rule learning · 0.0hybrid architecture · 0.0rough set theory · 0.0data summarization · 0.0concept clustering · 0.0decision theory · 0.0learning from examples · 0.0concept tree ascending · 0.0knowledge base construction · 0.0formal representation · 0.0
YearPublicationVenuePosition
2016 Deep parallelization of parallel FP-growth using parent-child MapReduce
abstract
MapReduce is an important programming model for processing in distributed environments. Compared to other distributed programming models, MapReduce reduces communication overheads between computers and improves fault tolerance. However, the MapReduce model does not allow for automatic synchronization between jobs. A large number of data analytics algorithms use a recursive divide-and-conquer approach, which inherently allows for parallelism at each level of recursion. However, it is often difficult to parallelize such algorithms using the traditional MapReduce model if the process requires synchronization. In this paper we introduce Parent-Child MapReduce, a version of the MapReduce programming model that allows for MapReduce tasks to be created dynamically and synchronized in a hierarchical parent-child fashion. Using the Parallel FP-Growth (PFP) algorithm for mining frequent patterns as a reference, we show that Parent-Child MapReduce can be used to parallelize recursive divide-and-conquer algorithms using the MapReduce model and that this can lead to significant speed ups in the computational speed of such algorithms. Our evaluation shows that we can achieve 68% (or 3 times) performance gain when used with PFP.
Adetokunbo Makanju, Zahra Farzanyar, Aijun An, Nick Cercone, Zane Zhenhua Hu, Yonggang Hu
IEEE BigData4
2016 Direction histogram: novel discriminative global feature for Thai offline handwritten OCR
Ekawat Chaowicharat, Kanlaya Naruedomkul, Nick Cercone
Pattern Anal. Appl.3
2015 Meaningful keyword search in relational databases with large and complex schema
abstract
Keyword search over relational databases offers an alternative way to SQL to query and explore databases that is effective for lay users who may not be well versed in SQL or the database schema. This becomes more pertinent for databases with large and complex schemas. An answer in this context is a join tree spanning tuples containing the query's keywords. As there are potentially many answers to the query, and the user is often only interested in seeing the top-k answers, how to rank the answers based on their relevance is of paramount importance. We focus on the relevance of join as the fundamental means to rank answers. We devise means to measure relevance of relations and foreign keys in the schema over the information content of the database. This can be done offline with no need for external models. We compare the proposed measures against a gold standard we derive from a real workload over TPC-E and evaluate the effectiveness of our methods. Finally, we test the performance of our measures against existing techniques to demonstrate a marked improvement, and perform a user study to establish naturalness of the ranking of the answers.
Mehdi Kargar, Aijun An, Nick Cercone, Parke Godfrey, Jarek Szlichta, Xiaohui Yu 0001
ICDE3
2015 Goal and Preference Identification through natural language
abstract
Goal models allow efficient representation of stakeholder goals and alternative ways by which these can be satisfied. Preferences over goals in the goal model are then used to specify criteria for selecting alternatives that fit specific contexts, situations and strategies. Given such preferences, automated reasoning tools allow for efficient exploration of such alternatives. Nevertheless, to be amenable to such automated processing, goals and preferences need to be specified in a formal language, making automated processing inaccessible to the very bearers of goals and preferences, i.e., the stakeholders. We combine natural language processing techniques to allow specification of preferences through natural language statements. The natural language statement is first matched through regular expressions to distinguish between the preference component and the goal component. The former is then mapped to a preferential strength measure, while the latter is used to identify the relevant goal in the goal model through statistical semantic similarity techniques. The result constitutes a formal representation that can be used for alternatives analysis. In this way, stakeholders can access advanced goal reasoning techniques through simple natural language preference expressions, facilitating their decision making in various requirements analysis contexts. An experimental evaluation with human participants shows that the proposed system is of substantial precision and that a mapping from natural preferential verbalizations to predefined preferential strength labels is possible through sampling from crowds.
Fatima Alabdulkareem, Nick Cercone, Sotirios Liaskos
RE2
2014 How Complementary Are Different Information Retrieval Techniques? A Study in Biomedicine Domain
Xiangdong An 0001, Nick Cercone
CICLing (2)2
2014 MeanKS: meaningful keyword search in relational databases with complex schema
abstract
Keyword search in relational databases was introduced in the last decade to assist users who are not familiar with a query language, the schema of the database, or the content of the data. An answer is a join tree of tuples that contains the query keywords. When searching a database with a complex schema, there are potentially many answers to the query. Therefore, ranking answers based on their relevance is crucial in this context. Prior work has addressed relevance based on the size of the answer or the IR scores of the tuples. However, this is not sufficient when searching a complex schema. We demonstrate MeanKS, a new system for meaningful keyword search over relational databases. The system first captures the user's interest by determining the roles of the keywords. Then, it uses schema-based ranking to rank join trees that cover the keyword roles. This uses the relevance of relations and foreign-key relationships in the schema over the information content of the database. In the demonstration, attendees can execute queries against the TPC-E warehouse and compare the proposed measures against a gold standard derived from a real workload over TPC-E to test the effectiveness of our methods.
Mehdi Kargar, Aijun An, Nick Cercone, Parke Godfrey, Jarek Szlichta, Xiaohui Yu 0001
SIGMOD Conference3
2013 Efficient mining of frequent itemsets in social network data based on MapReduce framework
abstract
Social Networks promote information sharing between people everywhere and at all times. Mining data produced in this data-rich environment can be extremely useful. Frequent itemset mining plays an important role in mining associations, correlations, sequential patterns, causality, episodes, multidimensional patterns, max-patterns, partial periodicity, emerging patterns, and many other significant data mining tasks in social networks. With the exponential growth of social network data towards a terabyte or more, most of the traditional frequent itemset mining algorithms become ineffective due to either huge resource requirements or large communications overhead. Cloud computing has proved that processing very large datasets over commodity clusters can be done by providing the right programming model. As a parallel programming model, MapReduce, one of most important techniques for cloud computing, has emerged in the mining of datasets of terabyte scale or larger on clusters of computers. In this paper, we propose an efficient frequent itemset mining algorithm, called IMRApriori, based on MapReduce framework which deals with Hadoop cloud, a parallel store and computing platform. The paper demonstrates experimental results to corroborate the theoretical claims.
Zahra Farzanyar, Nick Cercone
ASONAM2
2013 P2P-FISM: Mining (recently) frequent item sets from distributed data streams over P2P network
Zahra Farzanyar, Mohammadreza Kangavari, Nick Cercone
Inf. Process. Lett.3
2012 A Study on Novelty Evaluation in Biomedical Information Retrieval
Xiangdong An 0001, Nick Cercone
SPIRE2
2011 From QoD to QoS - Data Quality Issues in Cloud Computing
Przemyslaw Pawluk, Marin Litoiu, Nick Cercone
CLOSER3
2011 Towards Automatic Acquisition of a Fully Sense Tagged Corpus for Persian
Bahareh Sarrafzadeh, Nikolay Yakovets, Nick Cercone, Aijun An
ISMIS3
2011 Finding best evidence for evidence-based best practice recommendations in health care: the initial decision support system design
Nick Cercone, Xiangdong An 0001, Jiye Li, Zhenmei Gu, Aijun An
Knowl. Inf. Syst.1
2010 Correlating CpG islands, motifs, and sequence variants in human chromosome 21
abstract
CpG islands are important regions in DNA. They usually appear at the 5' end of genes containing GC-rich dinucleotides. When DNA methylation occurs, gene regulation is affected and it sometimes leads to carcinogenesis. We propose a new detection program using a hidden Markov model alongside the Viterbi algorithm. Our solution provides a graphical user interface not seen in many of the other CGI detection programs and we unify the detection and analysis under one program to allow researchers to scan a genetic sequence, detect the significant CGIs, and analyze the sequence once the scan is complete for any noteworthy findings. Using human chromosome 21, we run an analysis on a dataset of promoters and discover that the characteristics of methylated and unmethylated CGIs are significantly different. Finally, we detected significantly different motifs between methylated and unmethylated CGI promoters using MEME and MAST.
Leah Spontaneo, Nick Cercone
BIBM2
2010 Optimal IR: How Far Away?
Xiangdong An 0001, Jimmy Huang 0001, Nick Cercone
CICLing3
2010 Computer Support for Mathematical Word Problem Solving - Guided by Thai Teachers' Views
Wanintorn Supap, Kanlaya Naruedomkul, Nick Cercone
CSEDU (1)3
2010 Thai Visually Impaired's Requirements to Access Mathematics Via an Automatic Math Reader
Wararat Wongkia, Kanlaya Naruedomkul, Nick Cercone
CSEDU (1)3
2010 A method of discovering important rules using rules as attributes
abstract
Use of rough sets theory to select essential attributes that can represent the original data set is well known. A reduct is the subset of the original data set that contains the essential attributes. Decision rules generated from reducts can fully describe a data set. We introduce a new method of evaluating important rules by taking advantage of rough sets theory. We consider rules generated from the original data set as attributes in the new constructed decision table. Reducts generated from this new decision table contain essential attributes, which are the rules. Only important rules are contained in the reducts. Experiments on an artificial data set, UCI data sets, and real-world data sets show that the reduct rules are more important, and this new method provides an automatic and effective way of ranking rules. © 2009 Wiley Periodicals, Inc.
Jiye Li, Nick Cercone
Int. J. Intell. Syst.2
2009 Efficient image retrieval in DCT domain by hypothesis testing
abstract
We consider a hypothesis testing approach to content-based image retrieval (CBIR) using discrete cosine transform (DCT) coefficients restored by partially decoding JPEG images. In order to further decorrelate DC coefficients from an image, a 2 × 2 DCT is performed on the sub-image constructed from all the DC coefficients. Assume that each DCT coefficient sequence is emitted from a memoryless source, and all these sources are independent of each other. For each target image we form a hypothesis that its DCT coefficient sequences are emitted from the same sources as the corresponding sequences in the query image. Testing these hypotheses by measuring the log-likelihoods leads to a simple yet efficient scheme that ranks each target image according to the Kullback-Leibler (KL) divergence between the empirical distribution of the DCT coefficient sequences in the query image and that in the target image. Experiments on two image datasets show that our approach achieves consistently better retrieval results than related methods in the literature.
Daan He, Zhenmei Gu, Nick Cercone
ICIP3
2009 Fault-Tolerant Multi-Agent Exact Belief Propagation
abstract
Multiply sectioned Bayesian networks (MSBNs) support multi‐agent probabilistic inference in distributed large problem domains, where agents (subdomains) are organized by a tree structure (called hypertree). In earlier work, all belief updating methods on a hypertree are made of two rounds of propagation, each of which is implemented as a recursive process. Both processes need to be started from the same designated (root) hypernode. Agents perform local belief updating at most in a partial parallel manner. Such methods may not be suitable for practical multi‐agent environments because they are easy to crush for the problems happened in communication or local belief updating. In this paper, we present a fault‐tolerant belief updating method for multi‐agent probabilistic inference. In this method, multiple agents concurrently perform exact belief updating in a complete parallel. Temporary problems happened from time to time at some agents or some communication channels would not prevent agents from eventually converging to the correct beliefs. Permanently disconnected communication channels would not keep the properly connected portions of the system from appropriately finishing their belief updating within portions. Compared to the previous traversal‐based belief updating, the proposed approach is not only fault‐tolerant but also robust and scalable.
Xiangdong An 0001, Nick Cercone
Comput. Intell.2
2009 Distributed parallel compilation of MSBNs
abstract
Abstract Multiply sectioned Bayesian networks (MSBNs) support multiagent probabilistic inference in distributed large problem domains. Inference with MSBNs can be performed using their compiled representations. The compilation involves moralization and triangulation of a set of local graphical structures. Privacy of agents may prevent us from compiling MSBNs at a central location. In earlier work, agents performed compilation sequentially via a depth‐first traversal of the hypertree that organizes local subnets, where communication failure between any two agents would crush the whole work. In this paper, we present an asynchronous compilation method by which multiple agents compile MSBNs in full parallel. Compared with the traversal compilation, the asynchronous one is robust, self‐adaptive, and fault‐tolerant. Experiments show that both methods provide similar quality compilation to simple MSBNs, but the asynchronous one provides much higher quality compilation to complex MSBNs. Empirical study also indicates that the asynchronous one is consistently faster than the traversal one. Copyright © 2009 John Wiley & Sons, Ltd.
Xiangdong An 0001, Nick Cercone
Concurr. Comput. Pract. Exp.2
2008 Mining Causal Knowledge from Diagnostic Knowledge
Xiangdong An 0001, Nick Cercone
ADMA2
2008 Dynamic multiagent probabilistic inference
Xiangdong An 0001, Nick Cercone
Int. J. Approx. Reason.3
2007 Keyword Extraction Strategy for Item Banks Text Categorization
abstract
We proposed a feature selection approach, Patterned Keyword in Phrase (PKIP), to text categorization for item banks. The item bank is a collection of textual question items that are short sentences. Each sentence does not contain enough relevant words for directly categorizing by the traditional approaches such as “bag‐of‐words.” Therefore, PKIP was designed to categorize such question item using only available keywords and their patterns. PKIP identifies the appropriate keywords by computing the weight of all words. In this paper, two keyword selection strategies are suggested to ensure the categorization accuracy of PKIP. PKIP was implemented and tested with the item bank of Thai high primary mathematics questions. The test results have proved that PKIP is able to categorize the question items correctly and the two keyword selection strategies can extract the very informative keywords.
Atorn Nuntiyagul, Kanlaya Naruedomkul, Nick Cercone, Damras Wongsawang
Comput. Intell.3
2006 Privacy intrusion detection using dynamic Bayesian networks
abstract
Concerns for personal information privacy could be produced during information collection, transmission and handling. In information handling, privacy could be compromised from both inside and outside of organizations. Within an organization, private data are generally protected by organizations' privacy policies and the corresponding platforms for privacy practices. However, private data could still be misused intentionally or unintentionally by individuals who have legitimate accesses to them. In general, activities of a database operator form a stochastic process, and at different time, privacy intrusion behavior may show different features. In particular, one's past activities can help determine the natures of his/her current practices. In this paper, we propose to use dynamic Bayesian networks to model such temporal environments and detect any privacy intrusions happened within them.
Xiangdong An 0001, Dawn N. Jutla, Nick Cercone
ICEC3
2006 Segment-Based Hidden Markov Models for Information Extraction
abstract
Hidden Markov models (HMMs) are powerful statistical models that have found successful applications in Information Extraction (IE). In current approaches to applying HMMs to IE, an HMM is used to model text at the document level. This modelling might cause undesired redundancy in extraction in the sense that more than one filler is identified and extracted. We propose to use HMMs to model text at the segment level, in which the extraction process consists of two steps: a segment retrieval step followed by an extraction step. In order to retrieve extraction-relevant segments from documents, we introduce a method to use HMMs to model and retrieve segments. Our experimental results show that the resulting segment HMM IE system not only achieves near zero extraction redundancy, but also has better overall extraction performance than traditional document HMM IE systems.
Zhenmei Gu, Nick Cercone
ACL2
2006 Naive Bayes Modeling with Proper Smoothing for Information Extraction
abstract
Information extraction (IE) summarizes a collection of documents into a structural representation by identifying specific facts from text. The naive Bayes model is one of the first statistical models that have been applied to IE for learning extraction patterns from labeled data. In spite of the simplicity and popularity of the naive Bayes model, we have observed a formulation problem in previous work on naive Bayes IE. In this paper, we present a formal naive Bayes modeling for IE, by which the derived formula for the filler probability estimation is more theoretically sound. We also address smoothing techniques in order to overcome the data sparseness problem. Our proposed smoothing strategy is shown to be critical to the robustness of a naive Bayes IE system. Experimental results show that our naive Bayes IE systems achieve better extraction performance compared to related work.
Zhenmei Gu, Nick Cercone
FUZZ-IEEE2
2006 Probabilistic Internal Privacy Intrusion Detection
abstract
Many organizations need to maintain a lot of private data to run their businesses. Private data could be violated by both the inside and the outside intruders. In this paper, we propose a probabilistic method to detect insider privacy intrusion in database systems
Xiangdong An 0001, Dawn N. Jutla, Nick Cercone
IDEAS3
2006 Dynamic inference control in privacy preference enforcement
abstract
In pervasive (ubiquitous) environments, context-aware agents are used to obtain, understand, and share local contexts with each other so that the environments could be integrated seamlessly. Context sharing among agents should be made privacy-conscious. Privacy preferences are generally specified to regulate the exchange of the contexts, where who have rights under what conditions to have what contexts are designated. However, released contexts could be used to infer those unreleased. In particular, different contexts released could endanger the security of different contexts unreleased. The existing privacy preference specification platforms do not have a mechanism to prevent inference. To date, there have been very few inference control mechanisms specifically tailored to context management in pervasive (ubiquitous) environments. A Bayesian network based mechanism has been proposed to prevent privacy-sensitive contexts from being inferred from those to be released. Nevertheless, contexts in pervasive (ubiquitous) environments could change from time to time and are history dependent. In this paper, we propose to use dynamic Bayesian networks to track the most updated beliefs of the adversaries about the dynamic domains in order to evaluate which contexts in the domains could be released safely in various situations.
Xiangdong An 0001, Dawn N. Jutla, Nick Cercone
PST3
2006 Privacy Preserving Multiagent Probabilistic Reasoning about Ambiguous Contexts: A Case Study
abstract
Contexts in ubiquitous environments, either sensed or interpreted, are usually ambiguous. However, to provide context-aware services and applications, agents in the environments need to have an as clear as possible understanding of their contexts. Ambiguous contexts can be made clearer by agents using inference based on their domain knowledge, local and global evidence. Bayesian networks have been proposed to represent and reason about uncertain contexts under the single agent paradigm. In distributed multiagent systems, multiply sectioned Bayesian networks (MSBNs) provide a coherent framework for distributed multiagent probabilistic inference, where agents' privacy is respected. In this paper, we propose to apply MSBNs to uncertain contexts representation and reasoning in ubiquitous environments
Xiangdong An 0001, Dawn N. Jutla, Nick Cercone
Web Intelligence3
2005 A Probabilistic Evaluation Function for Relaxed Unification
abstract
Classical unification is strict in the sense that it requires a perfect agreement between the terms being unified. In practise, data are seldom error-free and can contain incorrect information. Classical unification fails when the data are imperfect. Relaxed unification is a new formalism that relaxes the rigid constraints of classical unification and enables reasoning under uncertainty and in the presence of inconsistent data. We propose a probabilistic evaluation function to evaluate the degree of mismatches in relaxed terms and illustrate its use with an example.
Tony Abou-Assaleh, Nick Cercone, Vlado Keselj
COMPSAC (2)2
2005 Recovering "Lack of Words" in Text Categorization for Item Banks
abstract
PKIP, patterned keywords in phrase, is our feature selection approach to text categorization (TC) for item banks. An item bank is a collection of textual data in which each item consists of short sentences and has only a few relevant words for categorization. Traditional TC techniques cannot provide sufficiently accurate results because of a "lack of words" problem. PKIP improves categorization accuracy and recovers from the "lack of words" problem. Our sample item bank is the collection of Thai primary mathematics problems and we use SVM as our classifier. Classification results show that PKIP produces acceptable classification performance.
Atorn Nuntiyagul, Nick Cercone, Kanlaya Naruedomkul
COMPSAC (2)2
2005 Hybrid Intelligent Systems: Selecting Attributes for Soft-Computing Analysis
abstract
It is difficult to provide significant insight into any hybrid intelligent system design. We offer an informative account of the basic ideas underlying hybrid intelligent systems. We propose a balanced approach to constructing a hybrid intelligent system for a medical domain, along with arguments in favor of this balance and mechanisms for achieving a proper balance. This first of a series of contributions to hybrid intelligent systems design focuses on selecting attributes for soft-computing analysis. One part of this first contribution in our system is developed. Two definitions, probe and probe reducts, are introduced. Our CDispro algorithm can produce the core attribute and reducts that are essential condition attributes in data sets. Our initial study tests data from the UCI repository and geriatric data from DalMedix. The performance and utility of generated reducts are evaluated by 3-fold cross-validation that illustrates reduced dimensionality and complexity of data sets and processes.
Puntip Pattaraintakorn, Nick Cercone, Kanlaya Naruedomkul
COMPSAC (1)2
2005 TSTMT: step towards an accurate Thai sign translation
abstract
We propose Thai sign to Thai machine translation (TSTMT), an alternative approach to machine translation which is used for translating from Thai sign language into Thai text. TSTMT performs the translation by recognizing an image sequence of signs as sign words and synthesizing a target sentence from those sign words. TSTMT represents a hybrid of image recognition and natural language processing techniques. This approach is comprised of seven modules which are straightforward and extendable. The applications of TSTMT can be used to facilitate learning in deaf people and to help them to communicate with hearing people via their own language. In this paper the idea and initial experiment are presented.
Thammanoon Ditcharoen, Kanlaya Naruedomkul, Nick Cercone, Bundit Tipakorn
ICMLA3
2005 PKIP: Feature Selection in Text Categorization for Item Banks
abstract
We propose an alternative approach to text categorization for item banks. An item bank is a collection of textual data in which each item consists of short sentences and has only a few relevant words for categorization; some items could be categorized into many categories. The traditional categorization techniques cannot provide sufficiently accurate results because of a "lack of words" problem. From this observation, items in the same category always have the same group of terms (or keywords) and the similar locations of these terms in phrases suggest that the items have a high probability to be in the same category. Our new methodology PKIP, patterned keywords in phrase, is proposed to improve categorization accuracy and recover from the "lack of words "problem. The k-highest weight order words are selected as the keywords from each category and their patterns are mapped for feature selection. The value of k affects the classification result. The item bank categorization process is based on a supervised machine learning technique. The sample of the item bank that is used in this research is the collection of Thai primary mathematics problems item bank and we use SVM in the Weka machine learning software package as our classifier. The result of the classification shows that our approach produces acceptable classification results and the highest classification result is given when k = 12
Atorn Nuntiyagul, Kanlaya Naruedomkul, Nick Cercone, Damras Wongsawang
ICTAI3
2004 Thai Syllable-Based Information Extraction Using Hidden Markov Models
Lalita Narupiyakul, Calvin Thomas, Nick Cercone, Booncharoen Sirinaovakul
CICLing3
2004 Intelligent Query Answering Based on Neighborhood Systems and Data Mining Techniques
Tsau Young Lin, Nick Cercone, Xiaohua Hu 0001, Jianchao Han
IDEAS2
2004 Applying Association Rules for Interesting Recommendations Using Rule Templates
Jiye Li, Nick Cercone
PAKDD3
2004 Detection of New Malicious Code Using N-grams Signatures
Tony Abou-Assaleh, Nick Cercone, Vlado Keselj, Ray Sweidan
PST2
2004 A Weighted Freshness Metric for Maintaining Search Engine Local Repository
abstract
Current search engines maintain a local repository to improve the search efficiency. A crawler is used to periodically poll the remote web pages to update the contents of the local repository. Due to the resource limitations, some local pages may be stale. To maintain the high freshness of the repository, the crawler is expected to revisit remote web pages in optimized order and frequency. The intuitive metric of freshness of the local repository is defined as the fraction of up-to-date web pages in the repository, which is merely based on the repository content, and does not, unfortunately, reflect the perspective of the search engine users, e.g., how often is a web page queried? We propose a novel weighted metric of the repository freshness with the importance of web pages being the weights. This metric not only takes into account the local web pages themselves but also the perspectives of the search engine users. We study the repository synchronization policy under this new metric, compare this metric with others, analyze its features, and discuss how the web page importance is determined.
Jianchao Han, Nick Cercone, Xiaohua Hu 0001
Web Intelligence2
2004 A data warehouse/online analytic processing framework for web usage mining and business intelligence reporting
abstract
Web usage mining is the application of data mining techniques to discover usage patterns and behaviors from web data (clickstream, purchase information, customer information, etc.) in order to understand and serve e-commerce customers better and improve the online business. In this article, we present a general data warehouse/online analytic processing (OLAP) framework for web usage mining and business intelligence reporting. When we integrate the web data warehouse construction, data mining, and OLAP into the e-commerce system, this tight integration dramatically reduces the time and effort for web usage mining, business intelligence reporting, and mining deployment. Our data warehouse/OLAP framework consists of four phases: data capture, webhouse construction (clickstream marts), pattern discovery and cube construction, and pattern evaluation and deployment. We discuss data transformation operations for web usage mining and business reporting in clickstream, session, and customer levels; describe the problems and challenging issues in each phase in detail; provide plausible solutions to the issues; and demonstrate the framework with some examples from some real web sites. Our data warehouse/OLAP framework has been integrated into some commercial e-commerce systems. We believe this data warehouse/OLAP framework would be very useful for developing any real-world web usage mining and business intelligence reporting systems. © 2004 Wiley Periodicals, Inc.
Xiaohua Hu 0001, Nick Cercone
Int. J. Intell. Syst.2
2004 Toward better scoring metrics for pseudo-independent models
abstract
Learning belief networks from data is NP-hard in general. A common method used in heuristic learning is the single-link lookahead search. When the problem domain is pseudo-independent (PI), the method cannot discover the underlying probabilistic model. In learning these models, to explicitly trade model accuracy and model complexity, parameterization of PI models is necessary. Understanding of PI models also provides a new dimension of trade-off in learning even when the underlying model may not be PI. In this work, we adopt a hypercube perspective to analyze PI models and derive an improved result for computing the maximum number of parameters needed to specify a full PI model. We also present results on parameterization of a subclass of partial PI models. © 2004 Wiley Periodicals, Inc. Int J Int Syst 19: 749–768, 2004.
Jae-Hyuck Lee, Nick Cercone
Int. J. Intell. Syst.3
2003 Discovering Cyber Communities from the WWW
abstract
In this paper we present a novel method to discover cyber communities from the WWW. A cyber community is a set of highly connected web pages in the WWW, which share the similar topics or interests. The WWW can be modeled as a huge scale-free network graph. Discovering cyber communities is converted to trawling the corresponding web graph to search for approximate p-quasi complete graph. The algorithm presented in this paper considers the characteristics of the scale-free network graphs and is based on the neighborhood information of the vertex in the p-quasi complete graph in the web graph.
Xiaohua Hu 0001, Jianchao Han, Nick Cercone
COMPSAC3
2003 Towards the Theory of Relaxed Unification
Tony Abou-Assaleh, Nick Cercone, Vlado Keselj
ISMIS2
2003 Applying Machine Learning to Text Segmentation for Information Retrieval
Jimmy Huang 0001, Fuchun Peng, Dale Schuurmans, Nick Cercone, Stephen E. Robertson
Inf. Retr.4
2002 Comparison of interestingness functions for learning web usage patterns
abstract
Livelink is a collaborative intranet, extranet and e-business application that enables employees and business partners of an organization to capture, share and reuse business information and knowledge. The usage of the Livelink software has been recorded by the Livelink Web server in its log files. We present an application of data mining techniques to the Livelink Web usage data. In particular, we focus on how to find interesting association rules and sequential patterns from the Livelink log files. A number of interestingness measures are used in our application to identify interesting rules and patterns. We present a comparison of these measures based on the feedback from domain experts. Some of the interestingness measures are found to be better than others.
Jimmy Huang 0001, Nick Cercone, Aijun An
CIKM2
2002 Investigating the Relationship between Word Segmentation Performance and Retrieval Performance in Chinese IR
Fuchun Peng, Jimmy Huang 0001, Dale Schuurmans, Nick Cercone
COLING4
2002 An OLAM framework for Web usage mining and business intelligence reporting
abstract
Web usage mining describes the procedure, which mines Web data (clickstream, purchase information, etc.) to determine usage patterns. Web usage mining is needed in order to understand and serve e-commerce customers better. We present a general OLAM (on-line analytical mining) framework for Web usage mining and business intelligence reporting. Our OLAM framework consists of four phases: data capture, Webhouse construction (clickstream marts), pattern discovery and pattern evaluation. We describe the problems and challenging issues in each phase in detail and provide a general approach and guidelines to Web usage mining and business intelligence reporting for e-commerce.
Xiaohua Hu 0001, Nick Cercone
FUZZ-IEEE2
2002 Discovery of Interesting Association Rules from Livelink Web Log Data
abstract
We present our experience in mining web usage patterns from a large collection of Livelink log data. Livelink is a web-based product of Open Text, which provides automatic management and retrieval of different types of information objects over an intranet or extranet. We report our experience in preprocessing raw log data and post-processing the mining results for finding interesting rules. In particular we compare and evaluate a number of rule interestingness measures and find that two of the measures that have not been used in association rule learning work very well.
Jimmy Huang 0001, Aijun An, Nick Cercone, Gary Promhouse
ICDM3
2002 Interactive Construction of Classification Rules
Jianchao Han, Nick Cercone
PAKDD2
2002 Using self-supervised word segmentation in Chinese information retrieval
abstract
We propose a self-supervised word-segmentation technique for Chinese information retrieval. This method combines the advantages of traditional dictionary based approaches with character based approaches, while overcoming many of their shortcomings. Experiments on TREC data show comparable performance to both the dictionary based and the character based approaches. However, our method is language independent and unsupervised, which provides a promising avenue for constructing accurate multilingual information retrieval systems that are flexible and adaptive.
Fuchun Peng, Jimmy Huang 0001, Dale Schuurmans, Nick Cercone, Stephen E. Robertson
SIGIR4
2002 Introduction to the Special Issue
Nick Cercone, Booncharoen Sirinaovakul, Vilas Wuwongse
Comput. Intell.1
2002 Generate and Repair Machine Translation
abstract
We propose Generate and Repair Machine Translation (GRMT), a constraint–based approach to machine translation that focuses on accurate translation output. GRMT performs the translation by generating a Translation Candidate (TC), verifying the syntax and semantics of the TC and repairing the TC when required. GRMT comprises three modules: Analysis Lite Machine Translation (ALMT), Translation Candidate Evaluation (TCE) and Repair and Iterate (RI). The key features of GRMT are simplicity, modularity, extendibility, and multilinguality. An English–Thai translation system has been implemented to illustrate the performance of GRMT. The system has been developed and run under SWI–Prolog 3.2.8. The English and Thai grammars have been developed based on Head–Driven Phrase Structure Grammar (HPSG) and implemented on the Attribute Logic Engine (ALE). GRMT was tested to generate the translations for a number of sentences/phrases. Examples are provided throughout the article to illustrate how GRMT performs the translation process.
Kanlaya Naruedomkul, Nick Cercone
Comput. Intell.2
2001 Interactive Construction of Decision Trees
Jianchao Han, Nick Cercone
PAKDD2
2001 Extracting meaningful semantic information with EMATISE: an HPSG-based Internet search engine parser
abstract
We describe EMATISE, our prototype natural language interface that serves as a front-end to Internet search engines; this prototype is based on the HPSG (Head-Driven Phrase Structure Grammar) formalism. Our current work concentrates on parsing natural language phrases and extracting meaningful semantic information. We present the architecture of our system for Internet access and the reasons for our choice of parser. The implementation of some modules of the system is discussed and some experimental results with sample sessions are presented.
Lijun Hou, Nick Cercone
SMC2
2001 From Computational Intelligence to Web Intelligence: An Ensemble from Potpourri
Nick Cercone
Web Intelligence1
2001 Implementation Issues and Paradigms of Visual KDD Systems
Jianchao Han, Nick Cercone
Web Intelligence2
2001 Rule Quality Measures for Rule Induction Systems: Description and Evaluation
abstract
A rule quality measure is important to a rule induction system for determining when to stop generalization or specialization. Such measures are also important to a rule‐based classification procedure for resolving conflicts among rules. We describe a number of statistical and empirical rule quality formulas and present an experimental comparison of these formulas on a number of standard machine learning datasets. We also present a meta‐learning method for generating a set of formula‐behavior rules from the experimental results which show the relationships between a formula's performance and the characteristics of a dataset. These formula‐behavior rules are combined into formula‐selection rules that can be used in a rule induction system to select a rule quality formula before rule induction. We will report the experimental results showing the effects of formula‐selection on the predictive performance of a rule induction system.
Aijun An, Nick Cercone
Comput. Intell.2
2001 Discovering Maximal Generalized Decision Rules Through Horixontal and Vertical Data Reduction
abstract
We present a method to learn maximal generalized decision rules from databases by integrating discretization, generalization and rough set feature selection. Our method reduces the data horizontally and vertically. In the first phase, discretization and generalization are integrated and the numeric attributes are discretized into a few intervals. The primitive values of symbolic attributes are replaced by high level concepts and some obvious superfluous or irrelevant symbolic attributes are also eliminated. Horizontal reduction is accomplished by merging identical tuples after the substitution of an attribute value by its higher level value in a pre‐defined concept hierarchy for symbolic attributes, or the discretization of continuous (or numeric) attributes. This phase greatly decreases the number of tuples in the database. In the second phase, a novel context‐sensitive feature merit measure is used to rank the features, a subset of relevant attributes is chosen based on rough set theory and the merit values of the features. A reduced table is obtained by removing those attributes which are not in the relevant attributes subset and the data set is further reduced vertically without destroying the interdependence relationships between classes and the attributes. Then rough set‐based value reduction is further performed on the reduced table and all redundant condition values are dropped. Finally, tuples in the reduced table are transformed into a set of maximal generalized decision rules. The experimental results on UCI data sets and a real market database demonstrate that our method can dramatically reduce the feature space and improve learning accuracy.
Xiaohua Hu 0001, Nick Cercone
Comput. Intell.2
2000 Rule Quality Measures Improve the Accuracy of Rule Induction: An Experimental Approach
Aijun An, Nick Cercone
ISMIS2
2000 RuleViz: a model for visualizing knowledge discovery process
abstract
We propose an interactive model, RuleViz, for visualizing the entire process of knowledge discovery and data mining.The model consists of ve components according to the main ingredients of the knowledge discovery process: original data visualization, visual data reduction, visual data preprocess, visual rule discovery, and rule visualization.The RuleViz model for visualizing the process of knowledge discovery is introduced and each component is discussed.Two aspects are emphasized, human-machine interaction and process visualization.The interaction helps the KDD system navigate through the enormous search spaces and recognize the intentions of the user, and the visualization of the KDD process helps users gain better insight i n to the multidimensional data, understand the intermediate results, and interpret the discovered patterns.According to the RuleViz model, we implement a n i n teractive system, CViz, which exploits \parallel coordinates" technique to visualize the process of rule induction.The original data is visualized on the parallel coordinates, and can be interactively reduced both horizontally and vertically.Three approaches for discretizing numerical attributes are provided in the visual data preprocessing.CViz learns classi cation rules on the basis of a rule induction algorithm and presents the result as the algorithm proceeds.The discovered rules are nally visualized on the parallel coordinates with each rule being displayed as a directed \polygon", and the rule accuracy and quality are used to render the \polygons" and control the choice of rules to be displayed to avoid clutter.The CViz system has been experimented with the UCI data sets and synthesis data sets, and the results demonstrate that the RuleViz model and the implemented visualization system are useful and helpful for understanding the process of knowledge discovery and interpreting the nal results.
Jianchao Han, Nick Cercone
KDD2
2000 AViz: A Visualization System for Discovering Numeric Association Rules
Jianchao Han, Nick Cercone
PAKDD2
2000 Introduction To the Special Issue on the 1999 Pacific Association for Computational Linguistics Conference
Nick Cercone, Kiyoshi Kogure, Kanlaya Naruedomkul
Comput. Intell.1
2000 Probability-Based Chinese Text Processing and Retrieval
abstract
We discuss the use of probability‐based natural language processing for Chinese text retrieval. We focus on comparing different text extraction methods and probabilistic weighting methods. Several document processing methods and probabilistic weighting functions are presented. A number of experiments have been conducted on large standard text collections. We present the experimental results that compare a word‐based text processing method with a character‐based method. The experimental results also compare a number of term‐weighting functions including both single‐unit weighting and compound‐unit weighting functions.
Jimmy Huang 0001, Stephen E. Robertson, Nick Cercone, Aijun An
Comput. Intell.3
1999 ORTES: The Design of a Real-Time Control Expert System
Aijun An, Nick Cercone, Christine W. Chan
ISMIS2
1999 Learning English Grapheme Segmentation Using the Iterated Version Space Algorithm
Jianna Jian Zhang, Howard J. Hamilton, Nick Cercone
ISMIS3
1999 Discretization of Continuous Attributes for Learning Classification Rules
Aijun An, Nick Cercone
PAKDD2
1999 DVIZ: A System for Visualizing Data Mining
Jianchao Han, Nick Cercone
PAKDD2
1999 English-Thai Translation: Initial Experiments with a Multiphase Translation System
abstract
A multiphase machine translation approach, Generate and Repair Machine Translation (GRMT), is proposed. GRMT is designed to generate accurate translations that focus primarily on retaining the linguistic meaning of the source language sentence. GRMT presently incorporates a limited multilingual translation capability. The central idea behind the GRMT approach is to generate a translationcandidate (TC) by quick and dirty machine translation (QDMT), then investigate the accuracy of that TC by translation candidate evaluation (TCE), and, if necessary, revise the translation in the repair and iterate (RI) phase. To demonstrate the GRMT approach, a translation system that translates from English to Thai has been developed. This paper presents the design characteristics and some experimental results of QDMT and also the initial design, some experiments, and proposed ideas behind TCE and RI.
Kanlaya Naruedomkul, Nick Cercone, Booncharoen Sirinaovakul
Comput. Intell.2
1999 Data Mining in Large Databases Using Domain Generalization Graphs
Robert J. Hilderman, Howard J. Hamilton, Nick Cercone
J. Intell. Inf. Syst.3
1999 Data Mining via Discretization, Generalization and Rough Set Feature Selection
Xiaohua Hu 0001, Nick Cercone
Knowl. Inf. Syst.2
1999 Rule-Induction and Case-Based Reasoning: Hybrid Architectures Appear Advantageous
abstract
Researchers have embraced a variety of machine learning (ML) techniques in their efforts to improve the quality of learning programs. The recent evolution of hybrid architectures for machine learning systems has resulted in several approaches that combine rule induction methods with case-based reasoning techniques to engender performance improvements over more traditional single-representation architectures. We briefly survey several major rule-induction and case-based reasoning ML systems. We then examine some interesting hybrid combinations of these systems and explain their strengths and weaknesses as learning systems. We present a balanced approach to constructing a hybrid architecture, along with arguments in favor of this balance and mechanisms for achieving a proper balance. Finally, we present some initial empirical results from testing our ideas and draw some conclusions based on those results.
Nick Cercone, Aijun An, Christine W. Chan
IEEE Trans. Knowl. Data Eng.1
1998 Mining Market Basket Data Using Share Measures and Characterized Itemsets
Robert J. Hilderman, Colin L. Carter, Howard J. Hamilton, Nick Cercone
PAKDD4
1997 Inducing and Using Decision Rules in the GRG Knowledge Discovery System
Ning Shan, Howard J. Hamilton, Nick Cercone
ECML3
1997 Integrating Rule Induction and Case-Based Reasoning to Enhance Problem Solving
Aijun An, Nick Cercone, Christine W. Chan
ICCBR2
1997 Learning Maximal Generalized Decision Rules via Discretization, Generalization, and Rough Set Feature Selection
abstract
We present a method to mine maximal generalized decision rules from databases by integrating discretization, generalization and rough sets feature selection. Our method reduces the data horizontally and vertically. In the first phase, discretization and generalization are integrated and the numeric attributes are discretized into a few intervals. Primitive values of symbolic attributes are replaced by high level concepts and some obvious superfluous or irrelevant symbolic attributes are also eliminated. Horizontal reduction is accomplished by merging identical tuples after the substitution of an attribute value by its higher level value in a predefined concept hierarchy for symbolic attributes or the discretization of continuous (or numeric) attributes. In the second phase, a novel context sensitive feature merit measure is used to rank the features, a subset of relevant attributes is chosen based on rough sets theory and the merit values of the features. A reduced table is obtained by removing those attributes which are not in the relevant attributes subset and the data set is further reduced vertically without destroying the interdependence relationships between the classes and the attributes. Rough sets based value reduction is further performed on the reduced table and all redundant condition values are dropped, finally, tuples in the reduced table are transformed into a set of maximal generalized decision rules. The experimental results on UCI data sets and an actual market database shows that our method can dramatically reduce the feature space and improve the learning accuracy.
Xiaohua Hu 0001, Nick Cercone
ICTAI2
1997 Share Based Measures for Itemsets
Colin L. Carter, Howard J. Hamilton, Nick Cercone
PKDD3
1997 Parallel Knowledge Discovery Using Domain Generalization Graphs
Robert J. Hilderman, Howard J. Hamilton, Robert J. Kowalchuk, Nick Cercone
PKDD4
1997 A "Microscopic" Study of Minimum Entropy Search in Learning Decomposable Markov Networks
S. K. Michael Wong, Nick Cercone
Mach. Learn.3
1996 Mining Knowledge Rules from Databases: A Rough Set Approach
abstract
The principle and experimental results of an attribute oriented rough set approach for knowledge discovery in databases are described. Our method integrates the database operation, rough set theory and machine learning techniques. In this method the learning procedure consists of two phases: data generalization and data reduction. In the data generalization phase, attribute oriented induction is performed attribute by attribute using attribute removal and concept ascension, some undesirable attributes to the discovery task are removed and the primitive data is generalized to the desirable level; thus a set of tuples may be generalized to the same generalized tuple. This procedure substantially reduces the computational complexity of the database learning process. Subsequently, in data reduction phase, the rough set method is applied to the generalized relation to find a minimal attribute set relevant to the learning task. The generalized relation is reduced further by removing those attributes which are irrelevant and/or unimportant to the learning task. Finally the tuples in the reduced relation are transformed into different knowledge rules based on different knowledge discovery algorithms. Based upon these principles, a prototype knowledge discovery system, DBROUGH has been constructed. In DBROUGH, a variety of knowledge discovery algorithms are incorporated and different kinds of knowledge rules, such as characteristic rules, classification rules, decision rules, maximal generalized rules can be discovered efficiently and effectively from large databases.
Xiaohua Hu 0001, Nick Cercone
ICDE2
1996 Attribute-oriented Induction Using Domain Generalization Graphs
abstract
Attribute-oriented induction summarizes the information in a relational database by repeatedly replacing specific attribute values with more general concepts according to user-defined concept hierarchies. We show how domain generalization graphs can be constructed from multiple concept hierarchies associated with an attribute, describe how these graphs can be used to control the generalization of a set of attributes, and present the Multi-Attribute Generalization algorithm for attribute-oriented induction using domain generalization graphs. Based upon a generate-and-test approach, the algorithm generates all possible combinations of nodes from the domain generalization graphs associated with the individual attributes, to produce all possible generalized relations for the set of attributes. We rant the interestingness of the resulting generalized relations using measures based upon relative entropy and variance. Our experiments show that these measures provide a basis for analyzing summary data from relational databases. Variance appears more useful because it tends to rank the less complex generalized relations (i.e., those with few attributes and/or few tuples) as more interesting.
Howard J. Hamilton, Robert J. Hilderman, Nick Cercone
ICTAI3
1996 Induction of Classification Rules from Imperfect Data
Ning Shan, Howard J. Hamilton, Nick Cercone
ISMIS3
1996 Rule Discovery from Databases with Decision Matrices
Wojciech Ziarko, Nick Cercone, Xiaohua Hu 0001
ISMIS2
1996 Discovering Classification Knowledge in Databases Using Rough Sets
Ning Shan, Wojciech Ziarko, Howard J. Hamilton, Nick Cercone
KDD4
1996 Critical Remarks on Single Link Search in Learning Belief Networks
S. K. Michael Wong, Nick Cercone
UAI3
1996 Intelligent Query Answering by Knowledge Discovery Techniques
abstract
Knowledge discovery facilitates querying database knowledge and intelligent query answering in database systems. We investigate the application of discovered knowledge, concept hierarchies, and knowledge discovery tools for intelligent query answering in database systems. A knowledge-rich data model is constructed to incorporate discovered knowledge and knowledge discovery tools. Queries are classified into data queries and knowledge queries. Both types of queries can be answered directly by simple retrieval or intelligently by analyzing the intent of query and providing generalized, neighborhood or associated information using stored or discovered knowledge. Techniques have been developed for intelligent query answering using discovered knowledge and/or knowledge discovery tools, which includes generalization, data summarization, concept clustering, rule discovery, query rewriting, deduction, lazy evaluation, application of multiple-layered databases, etc. Our study shows that knowledge discovery substantially broadens the spectrum of intelligent query answering and may have deep implications on query answering in data- and knowledge-base systems.
Jiawei Han 0001, Yue Huang 0001, Nick Cercone, Yongjian Fu 0001
IEEE Trans. Knowl. Data Eng.3
1995 GRG: knowledge discovery using information generalization, information reduction, and rule generation
abstract
We present the three-step GRG approach for learning decision rules from large relational databases. In the first step, an attribute-oriented concept tree ascension technique is applied to generalize an information system. This step loses some information but substantially improves the efficiency of the following steps. In the second step, the reduction technique is applied to generate a minimized information system called a reduct which contains a minimal subset of the generalized attributes and the smallest number of distinct tuples for those attributes. Finally, a set of maximally general rules are derived directly from the reduct. These rules can be used to interpret and understand the active mechanisms underlying the database.
Ning Shan, Howard J. Hamilton, Nick Cercone
ICTAI3
1995 Rough Sets Similarity-Based Learning from Databases
Xiaohua Hu 0001, Nick Cercone
KDD2
1995 Using Rough Sets as Tools for Knowledge Discovery
Ning Shan, Wojciech Ziarko, Howard J. Hamilton, Nick Cercone
KDD4
1995 Learning in Relational Databases: A Rough Set Approach
abstract
Knowledge discovery in databases, or dala mining, is an important direction in the development of data and knowledge‐based systems. Because of the huge amount of data stored in large numbers of existing databases, and because the amount of data generated in electronic forms is growing rapidly, it is necessary to develop efficient methods to extract knowledge from databases. An attribute‐oriented rough set approach has been developed for knowledge discovery in databases. The method integrates machine‐learning paradigm, especially learning‐from‐examples techniques, with rough set techniques. An attribute‐oriented concept tree ascension technique is first applied in generalization, which substantially reduces the computational complexity of database learning processes. Then the cause‐effect relationship among the attributes in the database is analyzed using rough set techniques, and the unimportant or irrelevant attributes are eliminated. Thus concise and strong rules with little or no redundant information can be learned efficiently. Our study shows that attribute‐oriented induction combined with rough set theory provide an efficient and effective mechanism for knowledge discovery in database systems.
Xiaohua Hu 0001, Nick Cercone
Comput. Intell.2
1995 Quantification of Unvertainty in Classification Rules Discovered from Databases
abstract
We apply rough set constructs to inductive learning from a database. A design guideline is suggested, which provides users the option to choose appropriate attributes, for the construction of data classification rules. Error probabilities for the resultant rule are derived. A classification rule can be further generalized using concept hierarchies. The condition for preventing overgeneralization is derived. Moreover, given a constraint, an algorithm for generating a rule with minimal error probability is proposed.
S. K. Michael Wong, Nick Cercone
Comput. Intell.3
1995 A Knowledge-based System for Generating Informative Responsens to Indirect Database Queries
Nick Cercone, Tadao Ichikawa
J. Intell. Inf. Syst.2
1994 Discovery of Decision Rules in Relational Databases: A Rough Set Approach
abstract
We develop an attribute-oriented rough set approach for the discovery of decision rules in relational databases. Our approach combines machine learning techniques and rough set theory. We consider a learning procedure to consist of the two phases data generalization and data reduction. In the data generalization phase, utilizing knowledge about concept hierarchies and relevance of the data, an attribute-oriented induction is performed attribute by attribute. Some undesirable attributes of the discovery task are removed and the primitive data in the databases are generalized to the desirable level; this process greatly decreases the number of tuples which must be examined for the discovery task and substantially reduces the computational complexity of the database learning processes. Subsequently, in data reduction phase, rough set theory is applied to the generalized relation; the cause-effect relationships among the condition and decision attributes in the databases are analyzed and the non-essential or irrelevant attributes to the discovery task are eliminated without losing information of the original database system. This process further reduces the generalized relation. Thus very concise and more accurate decision rules for each class in the decision attribute with little or no redundancy information, can be extracted automatically from the reduced relation during the learning process. Our study shows that attribute-oriented induction combined with rough set theory provide an efficient and effective mechanism for discovering decision rules in database systems.
Xiaohua Hu 0001, Nick Cercone
CIKM2
1994 DBROUGH: A Rough Set Based Knowledge Discovery System
Xiaohua Hu 0001, Ning Shan, Nick Cercone, Wojciech Ziarko
ISMIS3
1994 DBLearn: A System Prototype for Knowledge Discovery in Relational Databases
abstract
A prototyped data mining system, DBLearn, has been developed, which efficiently and effectively extracts different kinds of knowledge rules from relational databases. It has the following features: high level learning interfaces, tightly integrated with commercial relational database systems, automatic refinement of concept hierarchies, efficient discovery algorithms and good performance. Substantial extensions of its knowledge discovery power towards knowledge mining in object-oriented, deductive and spatial databases are under research and development.
Jiawei Han 0001, Yongjian Fu 0001, Yue Huang 0001, Yandong Cai, Nick Cercone
SIGMOD Conference5
1993 Decision-Theoretic Salience Interactions in Language Generation
T. Pattabhiraman, Nick Cercone
IJCAI2
1993 Guest Editors' Introduction
Nick Cercone, Mas Tsuchiya
IEEE Trans. Knowl. Data Eng.1
1993 Data-Driven Discovery of Quantitative Rules in Relational Databases
abstract
A quantitative rule is a rule associated with quantitative information which assesses the representativeness of the rule in the database. An efficient induction method is developed for learning quantitative rules in relational databases. With the assistance of knowledge about concept hierarchies, data relevance, and expected rule forms, attribute-oriented induction can be performed on the database, which integrates database operations with the learning process and provides a simple, efficient way of learning quantitative rules from large databases. The method involves the learning of both characteristic rules and classification rules. Quantitative information facilitates quantitative reasoning, incremental learning, and learning in the presence of noise. Moreover, learning qualitative rules can be treated as a special case of learning quantitative rules. It is shown that attribute-oriented induction provides an efficient and effective mechanism for learning various kinds of knowledge rules from relational databases.>
Jiawei Han 0001, Yandong Cai, Nick Cercone
IEEE Trans. Knowl. Data Eng.3
1992 Knowledge Discovery in Databases: An Attribute-Oriented Approach
Jiawei Han 0001, Yandong Cai, Nick Cercone
VLDB3
1992 Introduction
T. Pattabhiraman, Nick Cercone
Comput. Intell.2
1991 Learning in relational databases: an attribute-oriented approach
abstract
The development of efficient algorithms for learning from large relational databases is an important task in applicative machine learning. In this paper, we study knowledge discovery in relational databases and develop an attribute‐oriented learning method which extracts generalization rules from relational databases. The method adopts the artificial intelligence “learning‐from‐examples” paradigm and applies in the learning process an attribute‐oriented concept tree ascending technique which integrates database operations with the learning process and provides a simple and efficient way of learning from databases. The method learns both characteristic rules and classification rules of a learning concept, where a characteristic rule characterizes the properties shared by all the facts of the class being learned; while a classification rule characterizes the properties that distinguish the class being learned from other classes. The learning result could be a conjunctive rule or a rule with a small number of disjuncts. Moreover, learning can be performed with databases containing noisy data and exceptional cases using database statistics. Our analysis of the algorithms shows that attribute‐oriented induction substantially reduces the computational complexity of the database learning process. Le développement d'algorithmes efficaces permettant l'apprentissage à partir de bases de donnees relationnelles est une fonction importante de l'apprentissage automatique applicatif. Dans cet article, les auteurs examinent la découverte des connaissances dans les bases de données relationnelles et élaborent une méthode d'apprentissage orientée sur l'attribut qui extrait des bases de données relationnelles les règies de généralisation. La méthode adopte le paradigme d'apprentissage à partir d'exemples et applique au processus d'apprentissage la technique de l'arbre des concepts orientés sur l'attribut qui incorpore les opérations de base de données au processus d'apprentissage, ce qui permet d'obtenir une méthode simple et efficace d'apprentissage à partir des bases de données. La méthode fait l'apprentissage des règies caractéristiques et des règies de classification d'un concept d'apprentissage; la règie caractéristique qualifie les pro‐priétés communes à tous les faits d'une categorie faisant l'objet d'un apprentissage alors que la règie de classification caractérise les propriétés qui distinguent la catégorie faisant l'objet d'un apprentissage des autres catégories. Le résultat peut ětre une règie conjonctive ou une règie ayant un petit nombre de disjonctifs. Qui plus est, 1′apprentissage peut se faire avec des bases de données contenant des donnees bruitees et des cas exceptionnels utilisant des statistiques de bases de données. L'analyse des algorithmes démontre que l'induction orientée sur l'attribut réduit considérablement la complexité informàtique du processus d'apprentissage des bases de données.
Yandong Cai, Nick Cercone, Jiawei Han 0001
Comput. Intell.2
1990 An Attribute-Oriented Approach for Learning Classification Rules from Relational Databases
abstract
A classification rule is a rule which characterizes the properties that distinguish one class from other classes. An attribute-oriented induction algorithm which extracts classification rules from relational databases is developed. The algorithm adopts the artificial intelligence learning from examples paradigm and applies an attribute-oriented concept tree ascending technique in the learning process. The technique integrates database operations with the learning process and provides a simple and efficient way of learning from large databases. The algorithm learns both conjunctive rules and restricted forms of disjunctive rules. Using database statistics, learning can be performed on databases containing noisy data and exceptions. An analysis and comparison with other algorithms show that attribute-oriented induction substantially reduces the complexity of database learning processes.>
Yandong Cai, Nick Cercone, Jiawei Han 0001
ICDE2
1990 Selection: Salience, Relevance and the Coupling between Domain-Level Tasks and Text Planning
T. Pattabhiraman, Nick Cercone
INLG2
1989 Non-Singular Concepts in Natural Language Discourse
Tomek Strzalkowski, Nick Cercone
Comput. Linguistics2
1988 Providing Quality Responses with Natural Language Interfaces: The Null Value Problem
abstract
An underlying relational database model and the database query language SQL are assumed, and methods are presented for responding with appropriate answers to null value responses. This is done by using a knowledge base based on RM/T, an extended relational model. The advantages of this approach are described. To demonstrate the utility of the knowledge base model, a simple knowledge base is constructed. The algorithms that provide additional information when a null answer is returned are detailed.>
Mimi Kao, Nick Cercone, Wo-Shun Luk
IEEE Trans. Software Eng.2
1987 What do you mean "Null"? Turning Null Responses into Quality Responses
abstract
When natural language front-ends are introduced to database management systems, generation of quality responses have proven problematical in situations when null values arise. In our work, in which we assume the database query language is SQL, we present methods for responding with appropriate answers to null value responses. To do so we use a knowledge base based on RM/T, an extended relational model proposed by E. F. Codd. The advantages of this approach are described. To demonstrate the utility of the knowledge base, a simple knowledge base is constructed. A detailed algorithm is given to provide additional information when a null answer is returned.
Mimi Kao, Nick Cercone, Wo-Shun Luk
ICDE2
1986 A framework for computing extrasentential references
abstract
We are concerned with developing a computational method for selecting possible antecedents of referring expressions over sentence boundaries. Our stratified model which uses a Λ‐categorial language for meaning representation incorporates valuable features of Fregean‐type semantics (a la Lewis, Montague, Partee, and others) along with features of situation semantics developed by Barwise and Perry. We consider a series of selected two‐sentence stories which we use to illustrate referential interdependencies between sentences. We explain the conditions under which such dependencies arise, explain the conditions under which various translations can be performed, and formalize a set of rules which specify how to compute the reference. We restrict our discussion to two‐sentence stories to avoid most of the problems inherent in where to look for the reference, that is, how to determine the proper antecedent. We restrict our considerations in this paper to situations where a reference, if it can be computed at all, has a unique antecedent. Thus we consider examples such as John wants to catch a fish. He (John) wants to eat it. and John interviewed a man. The man killed him (John). We then summarize the transformation which encompasses these rules and relate it to the stratified model. We discuss three aspects of this transformation that merit special attention from the computational viewpoint and summarize the contributions we have made. We also discuss the computational characteristics of the stratified model in general and present our ideas for a computer realization; there is no implementation of the t“ratified model at this time.
Tomek Strzalkowski, Nick Cercone
Comput. Intell.2
1984 Artificial intelligence: Underlying assumptions and basic objectives
abstract
Abstract Artificial intelligence (AI) research has recently captured media interest and it is fast becoming our newest “hot” technology. AI is an interdisciplinary field which derives from a multiplicity of roots. In this article we present our perspectives on methodological assumptions underlying research efforts in Al. We also discuss the goals (design objectives) of AI across the spectrum of subareas it comprises. We conclude by discussing why there is increased interest in AI and whether current predictions of the future importance of AI are well founded.
Nick Cercone, Gordon I. McCalla
J. Am. Soc. Inf. Sci.1
1981 Lexicon design using perfect hash functions
abstract
The research reported in this paper derives from the recent algorithm of Cichelli (1980) for computing machine-independent, minimal perfect hash functions of the form:hash value: hash key length + associated value of the key's first letter + associated value of the key's last letterA minimal perfect hash function is one which provides single probe retrieval from a minimally-sized table of hash identifiers [ keys]. Cichelli's hash function is machine-independent because the character code used by a particular machine never enters into the hash calculation.Cichelli's algorithm uses a simple backtracking process to find an assignment of non-negative integers to letters which results in a perfect minimal hash function. Cichelli employs a twofold ordering strategy which rearranges the static set of keys in such a way that hash value collisions will occur and be resolved as early as possible during the backtracking process. This double ordering provides a necessary reduction in the size of the potentially large search space, thus considerably speeding the computation of associated values.In spite of Cichelli's ordering strategies, his method is found to require excessive computation to find hash functions for sets of keys with more than about 40 members. Cichelli's method is also limited since two keys with the same first and last letters and the same length are not permitted.Alternative algorithms and their implementations will be discussed in the next section; these algorithms overcome some of the difficulties encountered when using Cichelli's original algorithm. Some experimental results are presented, followed by a discussion of the application of perfect hash functions to the problem of natural language lexicon design.
Nick Cercone, Max Krause, John Boates
CHI (2)1
1977 A Note on Representing Adjectives and Adverbs
Nick Cercone
IJCAI1
1975 Toward a State Based Conceptual Representation
Nick Cercone, Lenhart K. Schubert
IJCAI1