EDBT 2026 Demo / reviewers in the wild / expert
Sau Dan Lee
dblp:71/6044
· DBLP profile ↗
21ranked-venue papers
2as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 20 · 2 first-authorArtificial intelligence and machine learning · 8 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
13 papers |
Data mining · 77% Information retrieval · 8% Web and social media mining · 5% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 100% | |
| Theoretical computer science
2 papers |
Computational geometry · 58% Algorithms and data structures · 42% | |
| Computer graphics and multimedia
1 paper |
Visualization and visual analytics · 100% | |
| Artificial intelligence
2 papers |
Probabilistic and Bayesian machine learning · 100% |
Topics — the 30 heaviest of 34, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining
pattern mining |
0.4 | 6 | 2013 | Mining Order-Preserving Submatrices from Data with Repeated Measurements · IEEE Trans. Knowl. Data Eng. 2013 OLAP on sequence data · SIGMOD Conference 2008 Mining Order-Preserving Submatrices from Data with Repeated Measurements · ICDM 2008 |
Data mining › predictive modeling
classification |
0.3 | 3 | 2011 | Decision Trees for Uncertain Data · IEEE Trans. Knowl. Data Eng. 2011 Naive Bayes Classification of Uncertain Data · ICDM 2009 Decision Trees for Uncertain Data · ICDE 2009 |
Data mining › pattern mining › matrix pattern mining
order-preserving submatrix mining |
0.2 | 2 | 2013 | Mining Order-Preserving Submatrices from Data with Repeated Measurements · IEEE Trans. Knowl. Data Eng. 2013 Mining Order-Preserving Submatrices from Data with Repeated Measurements · ICDM 2008 |
Data mining
uncertain data mining |
0.2 | 2 | 2012 | Efficient Mining of Frequent Item Sets on Large Uncertain Databases · IEEE Trans. Knowl. Data Eng. 2012 Decision Trees for Uncertain Data · ICDE 2009 |
Data mining › predictive modeling › classification
decision tree learning |
0.2 | 2 | 2011 | Decision Trees for Uncertain Data · IEEE Trans. Knowl. Data Eng. 2011 Decision Trees for Uncertain Data · ICDE 2009 |
Data mining › predictive modeling › classification
uncertain data classification |
0.2 | 2 | 2011 | Decision Trees for Uncertain Data · IEEE Trans. Knowl. Data Eng. 2011 Naive Bayes Classification of Uncertain Data · ICDM 2009 |
Data mining
clustering |
0.2 | 2 | 2010 | Clustering Uncertain Data Using Voronoi Diagrams and R-Tree Index · IEEE Trans. Knowl. Data Eng. 2010 Clustering Uncertain Data Using Voronoi Diagrams · ICDM 2008 |
Data mining › clustering
uncertain data clustering |
0.2 | 2 | 2010 | Clustering Uncertain Data Using Voronoi Diagrams and R-Tree Index · IEEE Trans. Knowl. Data Eng. 2010 Clustering Uncertain Data Using Voronoi Diagrams · ICDM 2008 |
Visualization and visual analytics
scientific visualization |
0.2 | 1 | 2014 | ECplot: an online tool for making standardized plots from large datasets for bioinformatics publications · Bioinform. 2014 |
Data mining › pattern mining › itemset mining
frequent itemset mining |
0.1 | 1 | 2012 | Efficient Mining of Frequent Item Sets on Large Uncertain Databases · IEEE Trans. Knowl. Data Eng. 2012 |
Data mining
incremental mining |
0.1 | 1 | 2012 | Efficient Mining of Frequent Item Sets on Large Uncertain Databases · IEEE Trans. Knowl. Data Eng. 2012 |
Information retrieval › retrieval models › latent semantic models
latent semantic indexing |
0.1 | 1 | 2011 | CubeLSI: An effective and efficient method for searching resources in social tagging systems · ICDE 2011 |
Information retrieval
retrieval models |
0.1 | 1 | 2011 | CubeLSI: An effective and efficient method for searching resources in social tagging systems · ICDE 2011 |
Web and social media mining › social tagging
social tagging systems |
0.1 | 1 | 2011 | CubeLSI: An effective and efficient method for searching resources in social tagging systems · ICDE 2011 |
Computational geometry
voronoi diagram |
0.1 | 2 | 2010 | Clustering Uncertain Data Using Voronoi Diagrams · ICDM 2008 Clustering Uncertain Data Using Voronoi Diagrams and R-Tree Index · IEEE Trans. Knowl. Data Eng. 2010 |
Indexing and storage engines › spatial index
r-tree |
0.1 | 1 | 2010 | Clustering Uncertain Data Using Voronoi Diagrams and R-Tree Index · IEEE Trans. Knowl. Data Eng. 2010 |
Spatial and temporal data management
spatial indexing |
0.1 | 1 | 2010 | Clustering Uncertain Data Using Voronoi Diagrams and R-Tree Index · IEEE Trans. Knowl. Data Eng. 2010 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › bayesian network › bayesian network classifiers
naive bayes |
0.1 | 1 | 2009 | Naive Bayes Classification of Uncertain Data · ICDM 2009 |
Bioinformatics and computational biology › gene expression analysis
microarray data analysis |
0.1 | 1 | 2008 | Mining Order-Preserving Submatrices from Data with Repeated Measurements · ICDM 2008 |
Query processing and optimization
OLAP |
0.1 | 1 | 2008 | OLAP on sequence data · SIGMOD Conference 2008 |
Algorithms and data structures
pruning |
0.1 | 1 | 2008 | Clustering Uncertain Data Using Voronoi Diagrams · ICDM 2008 |
Data mining › structured data mining › relational data mining
inductive query answering |
0.1 | 2 | 2003 | An Algebra for Inductive Query Evaluation · ICDM 2003 A Theory of Inductive Query Answering · ICDM 2002 |
Bioinformatics and computational biology
gene expression analysis |
0.0 | 1 | 2013 | Mining Order-Preserving Submatrices from Data with Repeated Measurements · IEEE Trans. Knowl. Data Eng. 2013 |
Data mining › pattern mining
constraint-based mining |
0.0 | 1 | 2003 | An Algebra for Inductive Query Evaluation · ICDM 2003 |
Data models and query languages
query algebra |
0.0 | 1 | 2003 | An Algebra for Inductive Query Evaluation · ICDM 2003 |
Web and social media mining
social tagging |
0.0 | 1 | 2011 | CubeLSI: An effective and efficient method for searching resources in social tagging systems · ICDE 2011 |
Data mining › pattern mining
association rule mining |
0.0 | 1 | 2002 | Effect of Data Skewness and Workload Balance in Parallel Data Mining · IEEE Trans. Knowl. Data Eng. 2002 |
Data mining › pattern mining › association rule mining
parallel association rule mining |
0.0 | 1 | 2002 | Effect of Data Skewness and Workload Balance in Parallel Data Mining · IEEE Trans. Knowl. Data Eng. 2002 |
Parallel and multicore computing
load balancing |
0.0 | 1 | 2002 | Effect of Data Skewness and Workload Balance in Parallel Data Mining · IEEE Trans. Knowl. Data Eng. 2002 |
Parallel and multicore computing
parallel data mining |
0.0 | 1 | 2002 | Effect of Data Skewness and Workload Balance in Parallel Data Mining · IEEE Trans. Knowl. Data Eng. 2002 |
Methods — techniques the papers use, named apart from their topics
pruning techniques · 0.5expected distance computation · 0.4benchmarking · 0.4XML-based formatting · 0.4pruning · 0.3repeated measurement aggregation · 0.3probability density function · 0.3decision tree induction · 0.2possible world semantics · 0.1poisson-binomial distribution · 0.1tensor decomposition · 0.1latent semantic indexing · 0.1class conditional probability estimation · 0.1repeated measurement handling · 0.1mining algorithm optimization · 0.1k-means · 0.1bounding-box pruning · 0.1entropy-based metrics · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | ECplot: an online tool for making standardized plots from large datasets for bioinformatics publicationsabstractMOTIVATION AND RESULTS: We have implemented ECplot, an online tool for plotting charts from large datasets. This tool supports a variety of chart types commonly used in bioinformatics publications. In our benchmarking, it was able to create a Box-and-Whisker plot with about 67 000 data points and 8 MB total file size within several seconds. The design of the tool makes common formatting operations easy to perform. It also allows more complex operations to be achieved by advanced XML (Extensible Markup Language) and programming options. Data and formatting styles are stored in separate files, such that style templates can be made and applied to new datasets. The text-based file formats based on XML facilitate efficient manipulation of formatting styles for a large number of data series. These file formats also provide a means to reproduce published figures from raw data, which complement parallel efforts in making the data and software involved in published analysis results accessible. We demonstrate this idea by using ECplot to replicate some complex figures from a previous publication. AVAILABILITY AND IMPLEMENTATION: ECplot and its source code (under MIT license) are available at https://yiplab.cse.cuhk.edu.hk/ecplot/. CONTACT: [email protected]. Alex Chun-Hong Fok, Sunny Siu-Nam Mok, Sau Dan Lee, Kevin Y. Yip |
Bioinform. | 3 |
| 2013 | Model-based probabilistic frequent itemset miningabstractData uncertainty is inherent in emerging applications such as location-based services, sensor monitoring systems, and data integration. To handle a large amount of imprecise information, uncertain databases have been recently developed. In this paper, we study how to efficiently discover frequent itemsets from large uncertain databases, interpreted under the Possible World Semantics. This is technically challenging, since an uncertain database induces an exponential number of possible worlds. To tackle this problem, we propose a novel methods to capture the itemset mining process as a probability distribution function taking two models into account: the Poisson distribution and the normal distribution. These model-based approaches extract frequent itemsets with a high degree of accuracy and support large databases. We apply our techniques to improve the performance of the algorithms for (1) finding itemsets whose frequentness probabilities are larger than some threshold and (2) mining itemsets with the $$k$$ highest frequentness probabilities. Our approaches support both tuple and attribute uncertainty models, which are commonly used to represent uncertain databases. Extensive evaluation on real and synthetic datasets shows that our methods are highly accurate and four orders of magnitudes faster than previous approaches. In further theoretical and experimental studies, we give an intuition which model-based approach fits best to different types of data sets. Thomas Bernecker, Reynold Cheng, David Wai-Lok Cheung, Hans-Peter Kriegel, Sau Dan Lee, Matthias Renz, Florian Verhein, Andreas Züfle |
Knowl. Inf. Syst. | 5 |
| 2013 | Mining Order-Preserving Submatrices from Data with Repeated MeasurementsabstractOrder-preserving submatrices (OPSM's) have been shown useful in capturing concurrent patterns in data when the relative magnitudes of data items are more important than their exact values. For instance, in analyzing gene expression profiles obtained from microarray experiments, the relative magnitudes are important both because they represent the change of gene activities across the experiments, and because there is typically a high level of noise in data that makes the exact values untrustable. To cope with data noise, repeated experiments are often conducted to collect multiple measurements. We propose and study a more robust version of OPSM, where each data item is represented by a set of values obtained from replicated experiments. We call the new problem OPSM-RM (OPSM with repeated measurements). We define OPSM-RM based on a number of practical requirements. We discuss the computational challenges of OPSM-RM and propose a generic mining algorithm. We further propose a series of techniques to speed up two time dominating components of the algorithm. We show the effectiveness and efficiency of our methods through a series of experiments conducted on real microarray data. Kevin Y. Yip, Ben Kao, Xinjie Zhu, Chun Kit Chui, Sau Dan Lee, David Wai-Lok Cheung |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2012 | Efficient Mining of Frequent Item Sets on Large Uncertain DatabasesabstractThe data handled in emerging applications like location-based services, sensor monitoring systems, and data integration, are often inexact in nature. In this paper, we study the important problem of extracting frequent item sets from a large uncertain database, interpreted under the Possible World Semantics (PWS). This issue is technically challenging, since an uncertain database contains an exponential number of possible worlds. By observing that the mining process can be modeled as a Poisson binomial distribution, we develop an approximate algorithm, which can efficiently and accurately discover frequent item sets in a large uncertain database. We also study the important issue of maintaining the mining result for a database that is evolving (e.g., by inserting a tuple). Specifically, we propose incremental mining algorithms, which enable Probabilistic Frequent Item set (PFI) results to be refreshed. This reduces the need of re-executing the whole mining algorithm on the new database, which is often more expensive and unnecessary. We examine how an existing algorithm that extracts exact item sets, as well as our approximate algorithm, can support incremental mining. All our approaches support both tuple and attribute uncertainty, which are two common uncertain database models. We also perform extensive evaluation on real and synthetic data sets to validate our approaches. David Wai-Lok Cheung, Reynold Cheng, Sau Dan Lee, Xuan S. Yang |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2011 | CubeLSI: An effective and efficient method for searching resources in social tagging systemsabstractIn a social tagging system, resources (such as photos, video and web pages) are associated with tags. These tags allow the resources to be effectively searched through tag-based keyword matching using traditional IR techniques. We note that in many such systems, tags of a resource are often assigned by a diverse audience of causal users (taggers). This leads to two issues that gravely affect the effectiveness of resource retrieval: (1) Noise: tags are picked from an uncontrolled vocabulary and are assigned by untrained taggers. The tags are thus noisy features in resource retrieval. (2) A multitude of aspects: different taggers focus on different aspects of a resource. Representing a resource using a flattened bag of tags ignores this important diversity of taggers. To improve the effectiveness of resource retrieval in social tagging systems, we propose CubeLSI - a technique that extends traditional LSI to include taggers as another dimension of feature space of resources. We compare CubeLSI against a number of other tag-based retrieval models and show that CubeLSI significantly outperforms the other models in terms of retrieval accuracy. We also prove two interesting theorems that allow CubeLSI to be very efficiently computed despite the much enlarged feature space it employs. Bin Bi, Sau Dan Lee, Ben Kao, Reynold Cheng |
ICDE | 2 |
| 2011 | Metric and trigonometric pruning for clustering of uncertain data in 2D geometric space
Wang Kay Ngai, Ben Kao, Reynold Cheng, Michael Chau, Sau Dan Lee, David Wai-Lok Cheung, Kevin Y. Yip |
Inf. Syst. | 5 |
| 2011 | Decision Trees for Uncertain DataabstractTraditional decision tree classifiers work with data whose values are known and precise. We extend such classifiers to handle data with uncertain information. Value uncertainty arises in many applications during the data collection process. Example sources of uncertainty include measurement/quantization errors, data staleness, and multiple repeated measurements. With uncertainty, the value of a data item is often represented not by one single value, but by multiple values forming a probability distribution. Rather than abstracting uncertain data by statistical derivatives (such as mean and median), we discover that the accuracy of a decision tree classifier can be much improved if the "complete information” of a data item (taking into account the probability density function (pdf)) is utilized. We extend classical decision tree building algorithms to handle data tuples with uncertain values. Extensive experiments have been conducted which show that the resulting classifiers are more accurate than those using value averages. Since processing pdfs is computationally more costly than processing single values (e.g., averages), decision tree construction on uncertain data is more CPU demanding than that for certain data. To tackle this problem, we propose a series of pruning techniques that can greatly improve construction efficiency. Smith Tsang, Ben Kao, Kevin Y. Yip, Wai-Shing Ho, Sau Dan Lee |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2010 | Accelerating probabilistic frequent itemset mining: a model-based approachabstractData uncertainty is inherent in emerging applications such as location-based services, sensor monitoring systems, and data integration. To handle a large amount of imprecise information, uncertain databases have been recently developed. In this paper, we study how to efficiently discover frequent itemsets from large uncertain databases, interpreted under the Possible World Semantics. This is technically challenging, since an uncertain database induces an exponential number of possible worlds. To tackle this problem, we propose a novel method to capture the itemset mining process as a Poisson binomial distribution. This model-based approach extracts frequent itemsets with a high degree of accuracy, and supports large databases. We apply our techniques to improve the performance of the algorithms for: (1) finding itemsets whose frequentness probabilities are larger than some threshold; and (2) mining itemsets with the k highest frequentness probabilities. Our approaches support both tuple and attribute uncertainty models, which are commonly used to represent uncertain databases. Extensive evaluation on real and synthetic datasets shows that our methods are highly accurate. Moreover, they are orders of magnitudes faster than previous approaches. Reynold Cheng, Sau Dan Lee, David Wai-Lok Cheung |
CIKM | 3 |
| 2010 | Clustering Uncertain Data Using Voronoi Diagrams and R-Tree IndexabstractAbstract-We study the problem of clustering uncertain objects whose locations are described by probability density functions (pdfs). We show that the UK-means algorithm, which generalizes the k-means algorithm to handle uncertain objects, is very inefficient. The inefficiency comes from the fact that UK-means computes expected distances (EDs) between objects and cluster representatives. For arbitrary pdfs, expected distances are computed by numerical integrations, which are costly operations. We propose pruning techniques that are based on Voronoi diagrams to reduce the number of expected distance calculations. These techniques are analytically proven to be more effective than the basic bounding-box-based technique previously known in the literature. We then introduce an R-tree index to organize the uncertain objects so as to reduce pruning overheads. We conduct experiments to evaluate the effectiveness of our novel techniques. We show that our techniques are additive and, when used in combination, significantly outperform previously known methods. Ben Kao, Sau Dan Lee, Foris K. F. Lee, David Wai-Lok Cheung, Wai-Shing Ho |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2009 | Decision Trees for Uncertain DataabstractTraditional decision tree classifiers work with data whose values are known and precise. We extend such classifiers to handle data with uncertain information, which originates from measurement/quantisation errors, data staleness, multiple repeated measurements, etc. The value uncertainty is represented by multiple values forming a probability distribution function (pdf). We discover that the accuracy of a decision tree classifier can be much improved if the whole pdf, rather than a simple statistic, is taken into account. We extend classical decision tree building algorithms to handle data tuples with uncertain values. Since processing pdf's is computationally more costly, we propose a series of pruning techniques that can greatly improve the efficiency of the construction of decision trees. Smith Tsang, Ben Kao, Kevin Y. Yip, Wai-Shing Ho, Sau Dan Lee |
ICDE | 5 |
| 2009 | Naive Bayes Classification of Uncertain DataabstractTraditional machine learning algorithms assume that data are exact or precise. However, this assumption may not hold in some situations because of data uncertainty arising from measurement errors, data staleness, and repeated measurements, etc. With uncertainty, the value of each data item is represented by a probability distribution function (pdf). In this paper, we propose a novel naive Bayes classification algorithm for uncertain data with a pdf. Our key solution is to extend the class conditional probability estimation in the Bayes model to handle pdf’s. Extensive experiments on UCI datasets show that the accuracy of naive Bayes model can be improved by taking into account the uncertainty information. Jiangtao Ren, Sau Dan Lee, Xianlu Chen, Ben Kao, Reynold Cheng, David Wai-Lok Cheung |
ICDM | 2 |
| 2008 | Mining Order-Preserving Submatrices from Data with Repeated MeasurementsabstractOrder-preserving submatrices (OPSM's) have been shown useful in capturing concurrent patterns in data when the relative magnitudes of data items are more important than their absolute values. To cope with data noise, repeated experiments are often conducted to collect multiple measurements. We propose and study a more robust version of OPSM, where each data item is represented by a set of values obtained from replicated experiments. We call the new problem OPSM-RM (OPSM with repeated measurements). We define OPSM-RM based on a number of practical requirements. We discuss the computational challenges of OPSM-RM and propose a generic mining algorithm. We further propose a series of techniques to speed up two time-dominating components of the algorithm. We clearly show the effectiveness of our methods through a series of experiments conducted on real microarray data. Chun Kit Chui, Ben Kao, Kevin Y. Yip, Sau Dan Lee |
ICDM | 4 |
| 2008 | Clustering Uncertain Data Using Voronoi DiagramsabstractWe study the problem of clustering uncertain objects whose locations are described by probability density functions (pdf). We show that the UK-means algorithm, which generalises the k-means algorithm to handle uncertain objects, is very inefficient. The inefficiency comes from the fact that UK-means computes expected distances (ED) between objects and cluster representatives. For arbitrary pdf's, expected distances are computed by numerical integrations, which are costly operations. We propose pruning techniques that are based on Voronoi diagrams to reduce the number of expected distance calculation. These techniques are analytically proven to be more effective than the basic bounding-box-based technique previous known in the literature. We conduct experiments to evaluate the effectiveness of our pruning techniques and to show that our techniques significantly outperform previous methods. Ben Kao, Sau Dan Lee, David Wai-Lok Cheung, Wai-Shing Ho, K. F. Chan |
ICDM | 2 |
| 2008 | OLAP on sequence dataabstractAbstract. Many kinds of real-life data exhibit logical ordering among their data items and are thus sequential in nature. However, traditional online analytical processing (OLAP) systems and techniques were not designed for sequence data and they are incapable of supporting sequence data analysis. In this paper, we propose the concept of Sequence OLAP, or S-OLAP for short. The biggest distinction of S-OLAP from traditional OLAP is that a sequence can be characterized not only by the attributes ’ values of its constituting items, but also by the subsequence/substring patterns it possesses. This paper studies many aspects related to Sequence OLAP. The concepts of sequence cuboid and sequence data cube are introduced. A prototype S-OLAP system is built in order to validate the proposed concepts. The prototype is able to support “pattern-based ” grouping and aggregation, which is currently not supported by any OLAP system. The implementation details of the prototype system as well as experimental results are presented. 1 Eric Lo 0001, Ben Kao, Wai-Shing Ho, Sau Dan Lee, Chun Kit Chui, David Wai-Lok Cheung |
SIGMOD Conference | 4 |
| 2003 | An Algebra for Inductive Query EvaluationabstractInductive queries are queries that generate pattern sets. We study properties of Boolean inductive queries, i.e. queries that are Boolean expressions over monotonic and antimonotonic constraints. More specifically, we introduce and study algebraic operations on the answer sets of such queries and show how these can be used for constructing and optimizing query plans. Special attention is devoted to the dimension of the queries, i.e. the minimum number of version spaces needed to represent the answer sets. The framework has been implemented for the pattern domain of strings and experimentally validated. Sau Dan Lee, Luc De Raedt |
ICDM | 1 |
| 2002 | A Theory of Inductive Query AnsweringabstractWe introduce the Boolean inductive query evaluation problem, which is concerned with answering inductive queries that are arbitrary Boolean expressions over monotonic and anti-monotonic predicates. Secondly, we develop a decomposition theory for inductive query evaluation in which a Boolean query Q is reformulated into k sub-queries Q/sub i/ = Q/sub A/ /spl and/ Q/sub M/ that are the conjunction of a monotonic and an anti-monotonic predicate. The solution to each subquery can be represented using a version space. We investigate how the number of version spaces k needed to answer the query can be minimized. Thirdly, for the pattern domain of strings, we show how the version spaces can be represented using a novel data structure, called the version space tree, and can be computed using a variant of the famous a priori algorithm. Finally, we present experiments that validate the approach. Luc De Raedt, Manfred Jaeger, Sau Dan Lee, Heikki Mannila |
ICDM | 3 |
| 2002 | Effect of Data Skewness and Workload Balance in Parallel Data MiningabstractTo mine association rules efficiently, we have developed a new parallel mining algorithm FPM on a distributed share-nothing parallel system in which data are partitioned across the processors. FPM is an enhancement of the FDM algorithm, which we previously proposed for distributed mining of association rules (Cheung et al., 1996). FPM requires fewer rounds of message exchanges than FDM and, hence, has a better response time in a parallel environment. The algorithm has been experimentally found to outperform CD, a representative parallel algorithm for the same goal (Agrawal and Srikant, 1994). The efficiency of FPM is attributed to the incorporation of two powerful candidate sets pruning techniques: distributed and global prunings. The two techniques are sensitive to two data distribution characteristics, data skewness, and workload balance. Metrics based on entropy are proposed for these two characteristics. The prunings are very effective when both the skewness and balance are high. In order to increase the efficiency of FPM, we have developed methods to partition a database so that the resulting partitions have high balance and skewness. Experiments have shown empirically that our partitioning algorithms can achieve these aims very well, in particular, the results are consistently better than a random partitioning. Moreover, the partitioning algorithms incur little overhead. So, using our partitioning algorithms and FPM together, we can mine association rules from a database efficiently. David Wai-Lok Cheung, Sau Dan Lee, Yongqiao Xiao |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2001 | Towards the building of a dense-region-based OLAP system
David Wai-Lok Cheung, Ben Kao, Kan Hu, Sau Dan Lee |
Data Knowl. Eng. | 5 |
| 1999 | DROLAP - A Dense-Region Based Approach to On-Line Analytical Processing
David Wai-Lok Cheung, Ben Kao, Kan Hu, Sau Dan Lee |
DEXA | 5 |
| 1998 | Is Sampling Useful in Data Mining? A Case in the Maintenance of Discovered Association Rules
Sau Dan Lee, David Wai-Lok Cheung, Ben Kao |
Data Min. Knowl. Discov. | 1 |
| 1997 | A General Incremental Technique for Maintaining Discovered Association Rules
David Wai-Lok Cheung, Sau Dan Lee, Ben Kao |
DASFAA | 2 |