Philippe Fournier-Viger

dblp:76/2649 · DBLP profile ↗
← Back
113ranked-venue papers in the field
26as first author
47since 2021 · last 2026
0000-0002-7680-9899ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 48 (18 first)Knowledge Engineering, Semantic Web & Information Systems · 27 (3 first)Database Systems & Data Management · 26 (5 first)Other / Interdisciplinary · 5Information Retrieval & Web Search · 4Big Data, Cloud & Distributed Data Systems · 3
YearPublicationVenuePosition
2026 SeqRFM: Fast RFM analysis in sequence data
Yanxin Zheng, Wensheng Gan, Pinlyu Zhou, Philippe Fournier-Viger
Inf. Sci.5
2025 SpaPool: Soft Partition Assignment Pooling for Graph Neural Networks
Rodrigue Govan, Romane Scherrer, Philippe Fournier-Viger, Nazha Selmaoui-Folcher
DaWaK3
2025 TK-RNSP: Efficient Top-K Repetitive Negative Sequential Pattern mining
abstract
Repetitive Negative Sequential Patterns (RNSPs) can provide critical insights into the importance of sequences. However, most current RNSP mining methods require users to set an appropriate support threshold to obtain the expected number of patterns, which is a very difficult task for the users without prior experience . To address this issue, we propose a new algorithm, TK-RNSP, to mine the Top- K RNSPs with the highest support, without the need to set a support threshold. In detail, we achieve a significant breakthrough by proposing a series of definitions that enable RNSP mining to satisfy anti-monotonicity. Then, we propose a bitmap-based Depth-First Backtracking Search (DFBS) strategy to decrease the heavy computational burden by increasing the speed of support calculation. Finally, we propose the algorithm TK-RNSP in an one-stage process, which can effectively reduce the generation of unnecessary patterns and improve computational efficiency comparing to those two-stage process algorithms. To the best of our knowledge, TK-RNSP is the first algorithm to mine Top- K RNSPs. Extensive experiments on eight datasets show that TK-RNSP has better flexibility and efficiency to mine Top- K RNSPs.
Dun Lan, Chuanhou Sun, Xiangjun Dong 0001, Ping Qiu, Yongshun Gong, Xinwang Liu 0002, Philippe Fournier-Viger, Chengqi Zhang
Inf. Process. Manag.7
2025 A novel multi-source weighted naive Bayes classifier
Guiliang Ou, Yu-Lin He, Philippe Fournier-Viger, Joshua Zhexue Huang
Inf. Sci.3
2025 OUTO-Miner: Detecting outlying occurrences in maximal frequent order-preserving patterns in time series
Youxi Wu, Siqi Lou, Yan Li 0087, Lei Guo 0015, Philippe Fournier-Viger, Xindong Wu 0001
Inf. Sci.5
2025 Mining Repetitive Negative Sequential Patterns with Gap Constraints
abstract
Sequential pattern mining (SPM) with gap constraints (or repetitive SPM or tandem repeat discovery in bioinformatics) can find frequent repetitive subsequences satisfying gap constraints, which are called positive sequential patterns with gap constraints (PSPGs). However, classical SPM with gap constraints cannot find the frequent missing items in the PSPGs. To tackle this issue, this article explores negative sequential patterns with gap constraints (NSPGs). We propose an efficient NSPG-Miner algorithm that can mine both frequent PSPGs and NSPGs simultaneously. To effectively reduce candidate patterns, we propose a pattern join strategy with negative patterns which can generate both positive and negative candidate patterns at the same time. To calculate the support (frequency of occurrence) of a pattern in each sequence, we explore a NegPair algorithm that employs a key-value pair array structure to deal with the gap constraints and the negative items simultaneously and can avoid redundant rescanning of the original sequence, thus improving the efficiency of the algorithm. To report the performance of NSPG-Miner, 11 competitive algorithms and 11 datasets are employed. The experimental results not only validate the effectiveness of the strategies adopted by NSPG-Miner but also verify that NSPG-Miner can discover more valuable information than the state-of-the-art algorithms. Algorithms and datasets can be downloaded from https://github.com/wuc567/Pattern-Mining/tree/master/NSPG-Miner .
Yan Li 0087, Zhulin Wang, Jing Liu 0066, Lei Guo 0015, Philippe Fournier-Viger, Youxi Wu, Xindong Wu 0001
ACM Trans. Knowl. Discov. Data5
2025 Mining Cross-Level High Utility Itemsets in Unstable and Negative Profit Databases
abstract
High utility itemset mining (HUIM) is one of the most compelling problems in data mining, extending frequent itemset mining (FIM) and serving as a crucial method for analyzing customer behavior. Many HUIM algorithms have been proposed to improve execution time and memory consumption. However, most assume that the profit is fixed for each item in a database, which is unrealistic. Some algorithms address products with unstable transaction profits but still need to run faster due to ineffective pruning strategies. Additionally, generalizing items into categories is often neglected. To address these issues, this paper considers a more practical database type that integrates unstable profits with a taxonomy of items. The proposed algorithm, CLHUN (Cross-level High Utility Itemset Mining in a Database with Unstable and Negative Profits), combines efficient techniques such as item sorting and tighter upper bounds to prune the search space. Furthermore, it introduces strategies to eliminate unpromising items during mining and reduce the number of transaction scans. Several experiments were conducted to evaluate the algorithm's performance. Results demonstrate that CLHUN is efficient with these techniques and strategies.
N. T. Tung, Trinh D. D. Nguyen, Loan T. T. Nguyen, Duc-Lung Vu, Philippe Fournier-Viger, Bay Vo
IEEE Trans. Knowl. Data Eng.5
2024 Special Issue Editorial on "The Innovative Use of Data Science to Transform How We Work and Live"
Yee Ling Boo, Manik Gupta, Weijia Zhang 0001, Philippe Fournier-Viger
Data Sci. Eng.4
2024 HUSM: High utility subgraph mining in single graph databases
Zhaoming Chen 0001, Guoting Chen, Wensheng Gan, Philippe Fournier-Viger
Inf. Sci.5
2024 Targeted mining of contiguous sequential patterns
Kaixia Hu, Wensheng Gan, Shan Huang 0009, Philippe Fournier-Viger
Inf. Sci.5
2024 MRI-CE: Minimal rare itemset discovery using the cross-entropy method
Wei Song 0004, Philippe Fournier-Viger, Youxi Wu
Inf. Sci.3
2024 Data heterogeneity's impact on the performance of frequent itemset mining algorithms
abstract
Frequent itemset mining (FIM) is a widely used task that extracts frequently occurring itemsets from data. Plenty of deterministic algorithms are available for this daunting task. However, experimental studies have not considered that data heterogeneity significantly impacts the algorithms' performance, giving rise to unfair comparisons and biased conclusions. This paper seeks to advance by comparing cutting-edge algorithms using various frequency thresholds, considering the resulting data heterogeneity. An extensive experimental study is carried out, including the number of itemsets mined per second as the performance quality measure to compare algorithms. The experiments include defining eight metrics to quantify data heterogeneity, and their values vary the algorithms' performance. The results revealed that some techniques (hypercube decomposition and k-items machine) are essential to achieve excellent performance on any dataset, and most algorithms behave similarly well when they include those techniques. As a final important point, different threshold values produce dissimilar data subsets (data heterogeneity is not an immutable data characteristic), so a previous study on the database characteristics with a few minimum support thresholds could be beneficial to select the best-suited FIM algorithm beforehand.
Antonio Manuel Trasierras, José María Luna, Philippe Fournier-Viger, Sebastián Ventura
Inf. Sci.3
2024 Efficient high utility itemset mining without the join operation
Yihe Yan, Xinzheng Niu, Philippe Fournier-Viger, Libin Ye, Fan Min 0001
Inf. Sci.4
2024 An efficient approach for incremental erasable utility pattern mining from non-binary data
Yoonji Baek, Hanju Kim, Myungha Cho, Hyeonmo Kim, Chanhee Lee 0005, Taewoong Ryu, Heonho Kim, Bay Vo, Vincent W. Gan, Philippe Fournier-Viger, Jerry Chun-Wei Lin, Witold Pedrycz, Unil Yun
Knowl. Inf. Syst.10
2024 CG-FHAUI: an efficient algorithm for simultaneously mining succinct pattern sets of frequent high average utility itemsets
Hai Duong 0001, Tin Truong 0001, Bac Le, Philippe Fournier-Viger
Knowl. Inf. Syst.4
2024 RNP-Miner: Repetitive Nonoverlapping Sequential Pattern Mining
abstract
Sequential pattern mining (SPM) is an important branch of knowledge discovery that aims to mine frequent sub-sequences (patterns) in a sequential database. Various SPM methods have been investigated, and most of them are classical SPM methods, since these methods only consider whether or not a given pattern occurs within a sequence. Classical SPM can only find the common features of sequences, but it ignores the number of occurrences of the pattern in each sequence, i.e., the degree of interest of specific users. To solve this problem, this paper addresses the issue of repetitive nonoverlapping sequential pattern (RNP) mining and proposes the RNP-Miner algorithm. To reduce the number of candidate patterns, RNP-Miner adopts an itemset pattern join strategy. To improve the efficiency of support calculation, RNP-Miner utilizes the candidate support calculation algorithm based on the position dictionary. To validate the performance of RNP-Miner, 10 competitive algorithms and 20 sequence databases were selected. The experimental results verify that RNP-Miner outperforms the other algorithms, and using RNPs can achieve a better clustering performance than raw data and classical frequent patterns. All the algorithms were developed using the PyCharm environment and can be downloaded fromhttps://github.com/wuc567/Pattern-Mining/tree/master/RNP-Miner.
Meng Geng, Youxi Wu, Yan Li 0087, Jing Liu 0066, Philippe Fournier-Viger, Xingquan Zhu 0001, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.5
2024 COPP-Miner: Top-k Contrast Order-Preserving Pattern Mining for Time Series Classification
abstract
Recently, order-preserving pattern (OPP) mining, a new sequential pattern mining method, has been proposed to mine frequent relative orders in a time series. Although frequent relative orders can be used as features to classify a time series, the mined patterns do not reflect the differences between two classes of time series well. To effectively discover the differences between time series, this paper addresses the top-kcontrast OPP (COPP) mining and proposes a COPP-Miner algorithm to discover the top-kcontrast patterns as features for time series classification, avoiding the problem of improper parameter setting. COPP-Miner is composed of three parts: extreme point extraction to reduce the length of the original time series, forward mining, and reverse mining to discover COPPs. Forward mining contains three steps: group pattern fusion strategy to generate candidate patterns, the support rate calculation method to efficiently calculate the support of a pattern, and two pruning strategies to further prune candidate patterns. Reverse mining uses one pruning strategy to prune candidate patterns and consists of applying the same process as forward mining. Experimental results validate the efficiency of the proposed algorithm and show that top-kCOPPs can be used as features to obtain a better classification performance.
Youxi Wu, Yufei Meng, Yan Li 0087, Lei Guo 0015, Xingquan Zhu 0001, Philippe Fournier-Viger, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.6
2023 Co-location Pattern Mining Under the Spatial Structure Constraint
Rodrigue Govan, Nazha Selmaoui-Folcher, Aristotelis Giannakos, Philippe Fournier-Viger
DEXA (1)4
2023 Mining Frequent Sequential Subgraph Evolutions in Dynamic Attributed Graphs
Zhi Cheng, Landy Andriamampianina, Franck Ravat, Jiefu Song, Nathalie Vallès-Parlangeau, Philippe Fournier-Viger, Nazha Selmaoui-Folcher
PAKDD (2)6
2023 Efficient mining of top-k high utility itemsets through genetic algorithms
José María Luna, R. Uday Kiran, Philippe Fournier-Viger, Sebastián Ventura
Inf. Sci.3
2023 A novel correlation Gaussian process regression-based extreme learning machine
Xuan Ye, Yu-Lin He, Manjing Zhang, Philippe Fournier-Viger, Joshua Zhexue Huang
Knowl. Inf. Syst.4
2023 Mining High Utility Itemsets Using Prefix Trees and Utility Vectors
abstract
High utility itemsets can reveal combinations of items that have a high profit, expense, or importance. Mining high utility itemsets in a database with$n$items generally results in a huge search space, composed of$2^{n}$itemsets, and heavy utility calculations for the explored itemsets. Previous algorithms using prefix tree structures perform two phases, namely candidate generation and testing. To avoid generating candidate itemsets, one-phase algorithms use list or hyper-link structures and have been proven to be superior to two-phase algorithms. However, it should be noted that a prefix tree is still an efficient structure for itemset mining problems, and especially algorithms using prefix trees such as FP-Growth have shown excellent performance for mining frequent itemsets. This paper proposes Hamm, a High-performance AlgorithM for Mining high utility itemsets. Hamm employs a novel TV (prefix Tree and utility Vector) structure and mines high utility itemsets in one phase without candidate generation. We also develop an efficient optimization which is incorporated into Hamm as a component. Using prefix trees and utility vectors, Hamm outperforms state-of-the-art algorithms on various databases in experiments. Experimental results also show that the proposed optimization remarkably reduces the search space and speeds up Hamm.
Jun-Feng Qu, Philippe Fournier-Viger, Mengchi Liu, Bo Hang, Chunyang Hu
IEEE Trans. Knowl. Data Eng.2
2023 OPR-Miner: Order-Preserving Rule Mining for Time Series
abstract
Discovering frequent trends in time series is a critical task in data mining. Recently, order-preserving matching was proposed to find all occurrences of a pattern in a time series, where the pattern is a relative order (regarded as a trend) and an occurrence is a sub-time series whose relative order coincides with the pattern. Inspired by the order-preserving matching, the existing order-preserving pattern (OPP) mining algorithm employs order-preserving matching to calculate the support, which leads to low efficiency. To address this deficiency, this paper proposes an algorithm called efficient frequent OPP miner (EFO-Miner) to find all frequent OPPs. EFO-Miner is composed of four parts: a pattern fusion strategy to generate candidate patterns, a matching process for the results of sub-patterns to calculate the support of super-patterns, a screening strategy to dynamically reduce the size of prefix and suffix arrays, and a pruning strategy to further dynamically prune candidate patterns. Moreover, this paper explores the order-preserving rule (OPR) mining and proposes an algorithm called OPR-Miner to discover strong rules from all frequent OPPs using EFO-Miner. Experimental results verify that OPR-Miner gives better performance than other competitive algorithms. More importantly, clustering and classification experiments further validate that OPR-Miner achieves good performance.
Youxi Wu, Xiaoqian Zhao, Yan Li 0087, Lei Guo 0015, Xingquan Zhu 0001, Philippe Fournier-Viger, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.6
2022 Constraint-based Sequential Rule Mining
abstract
Sequential rule mining (SRM) is an alternative to sequential pattern mining (SPM) when dealing with sequence data. SRM has a wide range of applications in numerous data analysis scenarios. Existing SRM algorithms usually discover the entire set of rules in the databases, which makes it not only difficult to analyze results because the discovered set is too large, but also does not consider the user’s expectations and background knowledge. To tackle this problem, researchers have explored related algorithms with different constraints according to their requirements. In this paper, we propose a flexible constraint-based SRM algorithm called ConSRM for discovering only the sequential rules within user-specified time bounds in a sequence database. This algorithm uses an efficient rule-growth method and develops corresponding constraints and pruning strategies to reduce the search space and speed up calculation. Comprehensive experiments were carried out on four real datasets to evaluate the performance (both effectiveness and efficiency) of ConSRM.
Zhaowen Yin, Wensheng Gan, Gengsen Huang, Yongdong Wu, Philippe Fournier-Viger
DSAA5
2022 Discovering Representative Attribute-stars via Minimum Description Length
abstract
Graphs are a popular data type found in many domains. Numerous techniques have been proposed to find interesting patterns in graphs to help understand the data and support decision-making. However, there are generally two limitations that hinder their practical use: (1) they have multiple parameters that are hard to set but greatly influence results, (2) and they generally focus on identifying complex subgraphs while ignoring relationships between attributes of nodes. Graphs are a popular data type found in many domains. Numerous techniques have been proposed to find interesting patterns in graphs to help understand the data and support decision-making. However, there are generally two limitations that hinder their practical use: (1) they have multiple parameters that are hard to set but greatly influence results, (2) and they generally focus on identifying complex subgraphs while ignoring relationships between attributes of nodes. To address these problems, we propose a parameter-free algorithm named CSPM (Compressing Star Pattern Miner) which identifies star-shaped patterns that indicate strong correlations among attributes via the concept of conditional entropy and the minimum description length principle. Experiments performed on several benchmark datasets show that CSPM reveals insightful and interpretable patterns and is efficient in runtime. Moreover, quantitative evaluations on two real-world applications show that CSPM has broad applications as it successfully boosts the accuracy of graph attribute completion models by up to 30.68% and uncovers important patterns in telecommunication alarm data.
Jiahong Liu 0001, Min Zhou 0006, Philippe Fournier-Viger, Menglin Yang 0001, Lujia Pan, Mourad Nouioua
ICDE3
2022 A Dynamic Variational Framework for Open-World Node Classification in Structured Sequences
abstract
Structured sequences are a popular data representation, used to model complex data such as traffic networks. A key machine learning task for structured sequences is node classification, that is predicting the class labels of unlabeled nodes. Though many node classification models were proposed, they assume a closed world setting, that all class labels appear in the training data. But in the real-world, the presence of never-before-seen class labels in testing data can considerably degrade a classifier’s accuracy. A promising solution to this issue is to build classifiers for an open-world setting, where samples with unknown class labels are continuously observed such that training and testing data may have different class label spaces. Several approaches have been proposed for open-world learning problems in computer vision and natural language processing, but they cannot be applied directly to structured sequences due to the complexity of their non-Euclidean properties and their dynamic nature. This paper addresses this important research gap by proposing a novel Open-world Structured Sequence node Classification (OSSC) model, to learn from structured sequences in an open-world setting. OSSC captures the structural and temporal information via a GCN-based dynamic variational framework. A latent distribution sequence is learned for each node using both stochastic states and deterministic states, to capture the evolution of node attributes and topology, followed by a sampling process to generate node representations. An open-world classification loss is further adopted to ensure that node representations are sensitive to unknown classes. And a combination of Openmax and Softmax is utilized to recognize nodes from unknown classes and to classify others to one of the known classes. Experiments on real-world datasets show that the proposed OSSC method is capable of learning accurate open-world node classifiers from structured sequence data.
Qin Zhang 0011, Qincai Li, Xiaojun Chen 0006, Peng Zhang 0001, Shirui Pan, Philippe Fournier-Viger, Joshua Zhexue Huang
ICDM6
2022 An efficient parallel algorithm for mining weighted clickstream patterns
Huy Minh Huynh, Loan T. T. Nguyen, Bay Vo, Zuzana Komínková Oplatková, Philippe Fournier-Viger, Unil Yun
Inf. Sci.5
2022 H-FHAUI: Hiding frequent high average utility itemsets
Bac Le, Tin Truong 0001, Hai Duong 0001, Philippe Fournier-Viger, Hamido Fujita
Inf. Sci.4
2022 CSPM: Discovering compressing stars in attributed graphs
Jiahong Liu 0001, Philippe Fournier-Viger, Min Zhou 0006, Ganghuan He, Mourad Nouioua
Inf. Sci.2
2022 Efficient mining of cross-level high-utility itemsets in taxonomy quantitative databases
N. T. Tung, Loan T. T. Nguyen, Trinh D. D. Nguyen, Philippe Fournier-Viger, Ngoc Thanh Nguyen 0001, Bay Vo
Inf. Sci.4
2022 A graph based approach for mining significant places in trajectory data
abstract
Significant place mining in spatiotemporal trajectory data is a key task for mobile pattern mining, useful for supporting location-aware services. State-of-the-art trajectory clustering algorithms utilize a density-based distance measure. However, some major problems with this approach are that (1) results are often inaccurate, especially on data of varying density, (2) the user must fine-tune many thresholds that are unintuitive to set, and (3) boundary points between clusters are often assigned to the wrong locations. Performance is also a major issue as many state-of-the-art algorithms have a very high time complexity. Motivated by these issues, this paper proposes an approach inspired by the data field theory and community detection. It is a graph-based significant place mining algorithm, called GB-SPM, for capturing and characterizing the essence of similarity between nodes. GB-SPM first applies a novel low index neighborhood velocity point filtration method to extract characteristic points. Then, a characteristic point index neighborhood is used to map them to graph nodes. In this way, the original problem is transformed into a community detection problem in complex community networks. Finally, a new edge weight metric is proposed to capture and characterize the nature of similarity between nodes. To evaluate clustering quality, we used the Silhouette (SI) for unannotated data to value inter-cluster separation and intra-cluster homogeneity. To evaluate mining effectiveness, we used Matthew’s correlation coefficient (MCC) for annotated data. Numerous experiments were carried out on real world datasets, and the accuracy and performance of the designed algorithm was compared with the state-of-the-art algorithms. Results show that GB-SPM improves on average SI by 13.9%, MCC by 20.7%, and runtime by 5.15 times.
Shimin Wang, Xinzheng Niu, Philippe Fournier-Viger, Dongmei Zhou, Fan Min 0001
Inf. Sci.3
2022 NWP-Miner: Nonoverlapping weak-gap sequential pattern mining
Youxi Wu, Yan Li 0087, Lei Guo 0015, Philippe Fournier-Viger, Xindong Wu 0001
Inf. Sci.5
2022 Fuzzy-driven periodic frequent pattern mining
Yanlin Qi, Guoting Chen, Wensheng Gan, Philippe Fournier-Viger
Inf. Sci.5
2022 Bayesian Attribute Bagging-Based Extreme Learning Machine for High-Dimensional Classification and Regression
abstract
This article presents a Bayesian attributebagging-based extreme learning machine (BAB-ELM)to handle high-dimensional classification and regression problems. First, thedecision-making degree (DMD)of a condition attribute is calculated based on the Bayesian decision theory, i.e., the conditional probability of the condition attribute given the decision attribute. Second, the condition attribute with the highest DMD is put into thecondition attribute group (CAG)corresponding to the specific decision attribute. Third, thebagging attribute groups (BAGs)are used to train an ensemble learning model ofextreme learning machines (ELMs).Each base ELM is trained on a BAG which is composed of condition attributes that are randomly selected from the CAGs. Fourth, the information amount ratios of bagging condition attributes to all condition attributes is used as the weights to fuse the predictions of base ELMs in BAB-ELM. Exhaustive experiments have been conducted to compare the feasibility and effectiveness of BAB-ELM with seven other ELM models, i.e., ELM, ensemble-based ELM (EN-ELM), voting-based ELM (V-ELM), ensemble ELM (E-ELM), ensemble ELM based on multi-activation functions (MAF-EELM), bagging ELM, and simple ensemble ELM. Experimental results show that BAB-ELM is convergent with the increase of base ELMs and also can yield higher classification accuracy and lower regression error for high-dimensional classification and regression problems.
Yu-Lin He, Xuan Ye, Joshua Zhexue Huang, Philippe Fournier-Viger
ACM Trans. Intell. Syst. Technol.4
2022 NTP-Miner: Nonoverlapping Three-Way Sequential Pattern Mining
abstract
Nonoverlapping sequential pattern mining is an important type of sequential pattern mining (SPM) with gap constraints, which not only can reveal interesting patterns to users but also can effectively reduce the search space using the Apriori (anti-monotonicity) property. However, the existing algorithms do not focus on attributes of interest to users, meaning that existing methods may discover many frequent patterns that are redundant. To solve this problem, this article proposes a task called nonoverlapping three-way sequential pattern (NTP) mining, where attributes are categorized according to three levels of interest: strong, medium, and weak interest. NTP mining can effectively avoid mining redundant patterns since the NTPs are composed of strong and medium interest items. Moreover, NTPs can avoid serious deviations (the occurrence is significantly different from its pattern) since gap constraints cannot match with strong interest patterns. To mine NTPs, an effective algorithm is put forward, called NTP-Miner, which applies two main steps: support (frequency occurrence) calculation and candidate pattern generation. To calculate the support of an NTP, depth-first and backtracking strategies are adopted, which do not require creating a whole Nettree structure, meaning that many redundant nodes and parent–child relationships do not need to be created. Hence, time and space efficiency is improved. To generate candidate patterns while reducing their number, NTP-Miner employs a pattern join strategy and only mines patterns of strong and medium interest. Experimental results on stock market and protein datasets show that NTP-Miner not only is more efficient than other competitive approaches but can also help users find more valuable patterns. More importantly, NTP mining has achieved better performance than other competitive methods in clustering tasks. Algorithms and data are available at: https://github.com/wuc567/Pattern-Mining/tree/master/NTP-Miner .
Youxi Wu, Lanfang Luo, Yan Li 0087, Lei Guo 0015, Philippe Fournier-Viger, Xingquan Zhu 0001, Xindong Wu 0001
ACM Trans. Knowl. Discov. Data5
2021 Mining Partially-Ordered Episode Rules in an Event Sequence
Philippe Fournier-Viger, Yangming Chen, Farid Nouioua, Jerry Chun-Wei Lin
ACIIDS1
2021 Investigating Crossover Operators in Genetic Algorithms for High-Utility Itemset Mining
M. Saqib Nawaz, Philippe Fournier-Viger, Wei Song 0004, Jerry Chun-Wei Lin, Bernd Noack
ACIIDS2
2021 TKQ: Top-K Quantitative High Utility Itemset Mining
Mourad Nouioua, Philippe Fournier-Viger, Wensheng Gan, Youxi Wu, Jerry Chun-Wei Lin, Farid Nouioua
ADMA2
2021 Discovering Relative High Utility Itemsets in Very Large Transactional Databases Using Null-Invariant Measure
abstract
High utility itemset mining is an important model in data mining. It involves discovering all itemsets in a quantitative transactional database that satisfy a user-specified minimum utility (minUtil) constraint. MinUtil controls the minimum value that an itemset must maintain in a database. Since the model evaluates an itemset’s interestingness using only the minUtil constraint, it implicitly assumes that all items in the database have similar utility values. However, some items have high utility, while others may have relatively low utility in a database. If minUtil is set too high, the user will miss all itemsets containing low utility items. To find itemsets that involve both high and low utility items, minUtil has to be set very low. However, this may cause a combinatorial explosion as the items with high utility may combine with others in all possible ways. This dilemma is called the low utility item problem. This paper proposes a flexible model of relative high utility itemset to address this problem. We introduce a new null-invariant measure, called utility ratio, to evaluate the interestingness of an itemset in the database. We also present a fast single scan algorithm to find all desired itemsets in the database. Experimental results demonstrate that the proposed algorithm is efficient. Finally, a case study on Yahoo! JAPAN retail data shows that the proposed model is useful.
R. Uday Kiran, Pradeep Pallikila, José María Luna, Philippe Fournier-Viger, Masashi Toyoda, P. Krishna Reddy
IEEE BigData4
2021 Mining Partially-Ordered Episode Rules with the Head Support
Yangming Chen, Philippe Fournier-Viger, Farid Nouioua, Youxi Wu
DaWaK2
2021 Stable High Utility Itemset Mining
abstract
High Utility Itemset Mining (HUIM) aims at finding all sets of items that have high importance in a database, as measured by a utility function. Although HUIM has many applications, a key limitation is that the discovered patterns often have an unstable utility over time. For example, while a set of products may yield a high utility (profit) over a year, that utility may fluctuate from weeks to weeks. To discover patterns that have a stable utility and hence that are more suitable for decision-making, this paper redefines HUIM as the task of discovering Stable High Utility Itemsets (StableHUI). An efficient tree-based and pattern-growth algorithm named Stable-Growth is proposed to extract all the StableHUI. Several experiments on two real-world datasets and two synthetic datasets show that Stable-Growth is up to 60% faster than a baseline and that it can filter out numerous unstable HUI.
Acquah Hackman, Yu Huang 0018, Philippe Fournier-Viger, Vincent S. Tseng
iiWAS3
2021 Average utility driven data analytics on damped windows for intelligent systems with data streams
abstract
In industrial areas, most of databases are dynamic databases, and the volume of the databases has grown with the passage of time. Especially, pattern mining for incremental database needs different approaches from static database because the profit or the accuracy of the previously inserted data can be reduced. Since data is time- sensitive, the recent data has a relatively higher value than the old data. In this paper, we suggest the damped window based average utility driven data analytics for intelligent systems, which the damped window reflects the importance according to the arrival time of the transactions. The proposed mining approach adopts novel data structure, which modify the importance of item as the passage of time, and it improves mining efficiency with several pruning strategies and without generating candidate patterns. To evaluate the performance of the proposed mining approach, we conducted various experiments using several real and synthetic data sets. The result of the experiments presented that the suggested method performs better in terms of runtime and memory usage than the other state-of-the-art mining techniques. Moreover, through the scalability experiments, which changed the number of different items or transactions, we verified that the proposed algorithm maintained a stable performance under various environmental changes.
Jongseong Kim, Unil Yun, Taewoong Ryu, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Witold Pedrycz
Int. J. Intell. Syst.6
2021 Mining local periodic patterns in a discrete sequence
Philippe Fournier-Viger, R. Uday Kiran, Sebastián Ventura, José María Luna
Inf. Sci.1
2021 A guided FP-Growth algorithm for mining multitude-targeted item-sets and class association rules in imbalanced data
Lior Shabtay, Philippe Fournier-Viger, Rami Yaari, Itai Dattner
Inf. Sci.2
2021 Efficient algorithms for mining frequent high utility sequences with constraints
Tin Truong 0001, Hai Duong 0001, Bac Le, Philippe Fournier-Viger, Unil Yun, Hamido Fujita
Inf. Sci.4
2021 Utility Mining Across Multi-Dimensional Sequences
abstract
Knowledge extraction from database is the fundamental task in database and data mining community, which has been applied to a wide range of real-world applications and situations. Different from the support-based mining models, the utility-oriented mining framework integrates the utility theory to provide more informative and useful patterns. Time-dependent sequence data are commonly seen in real life. Sequence data have been widely utilized in many applications, such as analyzing sequential user behavior on the Web, influence maximization, route planning, and targeted marketing. Unfortunately, all the existing algorithms lose sight of the fact that the processed data not only contain rich features (e.g., occur quantity, risk, and profit), but also may be associated with multi-dimensional auxiliary information, e.g., transaction sequence can be associated with purchaser profile information. In this article, we first formulate the problem of utility mining across multi-dimensional sequences, and propose a novel framework named MDUS to extract Multi-Dimensional Utility-oriented Sequential useful patterns. To the best of our knowledge, this is the first study that incorporates the time-dependent sequence-order, quantitative information, utility factor, and auxiliary dimension. Two algorithms respectively named MDUS EM and MDUS SD are presented to address the formulated problem. The former algorithm is based on database transformation, and the later one performs pattern joins and a searching method to identify desired patterns across multi-dimensional sequences. Extensive experiments are carried on six real-life datasets and one synthetic dataset to show that the proposed algorithms can effectively and efficiently discover the useful knowledge from multi-dimensional sequential databases. Moreover, the MDUS framework can provide better insight, and it is more adaptable to real-life situations than the current existing models.
Wensheng Gan, Jerry Chun-Wei Lin, Jiexiong Zhang, Hongzhi Yin, Philippe Fournier-Viger, Han-Chieh Chao, Philip S. Yu
ACM Trans. Knowl. Discov. Data5
2021 A Survey of Utility-Oriented Pattern Mining
abstract
The main purpose of data mining and analytics is to find novel, potentially useful patterns that can be utilized in real-world applications to derive beneficial knowledge. For identifying and evaluating the usefulness of different kinds of patterns, many techniques and constraints have been proposed, such as support, confidence, sequence order, and utility parameters (e.g., weight, price, profit, quantity, satisfaction, etc.). In recent years, there has been an increasing demand for utility-oriented pattern mining (UPM, or called utility mining). UPM is a vital task, with numerous high-impact applications, including cross-marketing, e-commerce, finance, medical, and biomedical applications. This survey aims to provide a general, comprehensive, and structured overview of the state-of-the-art methods of UPM. First, we introduce an in-depth understanding of UPM, including concepts, examples, and comparisons with related concepts. A taxonomy of the most common and state-of-the-art approaches for mining different kinds of high-utility patterns is presented in detail, including Apriori-based, tree-based, projection-based, vertical-/horizontal-data-format-based, and other hybrid approaches. A comprehensive review of advanced topics of existing high-utility pattern mining techniques is offered, with a discussion of their pros and cons. Finally, we present several well-known open-source software packages for UPM. We conclude our survey with a discussion on open and practical challenges in this field.
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao, Vincent S. Tseng, Philip S. Yu
IEEE Trans. Knowl. Data Eng.3
2020 Mining Attribute Evolution Rules in Dynamic Attributed Graphs
Philippe Fournier-Viger, Ganghuan He, Jerry Chun-Wei Lin, Heitor Murilo Gomes
DaWaK1
2020 Discovering Frequent Spatial Patterns in Very Large Spatiotemporal Databases
abstract
Frequent pattern mining is an important model in data mining. It involves finding all patterns in a transactional database that satisfy the user-specified minimum support (minSup) constraint. The minSup controls the minimum number of transactions that a pattern must cover in a transactional database. Since only minSup is used to evaluate a pattern's interestingness, the frequent pattern model implicitly assumes that spatial information of the items will not impact the interestingness of a pattern in the database. This assumption limits the applicability of the frequent pattern model in many real-world applications. It is because patterns whose items are close to each other are typically more attractive to the user than the patterns whose items are far from each other in a coordinate system. With this motivation, this paper proposes a novel model of frequent spatial pattern that may exist in a spatiotemporal database. An efficient pattern-growth algorithm, called Frequent Spatial Pattern-growth (FSP-growth), has also been presented to mine all desired patterns in a database. Experimental results demonstrate that our algorithm is efficient. The usefulness of the proposed patterns has also been shown with a real-world application.
R. Uday Kiran, Sourabh Shrivastava, Philippe Fournier-Viger, Koji Zettsu, Masashi Toyoda, Masaru Kitsuregawa
SIGSPATIAL/GIS3
2020 Mining Locally Trending High Utility Itemsets
Philippe Fournier-Viger, Jerry Chun-Wei Lin, Jaroslav Frnda
PAKDD (2)1
2020 Discovering rare correlated periodic patterns in multiple sequences
Philippe Fournier-Viger, Zhitian Li, Jerry Chun-Wei Lin, R. Uday Kiran
Data Knowl. Eng.1
2020 EHAUSM: An efficient algorithm for high average utility sequence mining
Tin Truong 0001, Hai Duong 0001, Bac Le, Philippe Fournier-Viger
Inf. Sci.4
2020 High average-utility sequential pattern mining based on uncertain databases
Jerry Chun-Wei Lin, Ting Li 0011, Matin Pirouz, Ji Zhang 0001, Philippe Fournier-Viger
Knowl. Inf. Syst.5
2019 HUE-Span: Fast High Utility Episode Mining
Philippe Fournier-Viger, Jerry Chun-Wei Lin, Unil Yun
ADMA1
2019 Mining High-Utility Sequential Patterns from Big Datasets
abstract
High-Utility Sequential Pattern Mining (HUSPM) has become an emerging issue in recent decades since it reveals more information such as the utility and sequence factors for knowledge discovery. For the previous works, many algorithms were presented to speed up the mining performance regarding a single machine with small datasets. In real-world applications, the size of dataset can be collected from many places or devices, such as PC, Internet of Things (IoT), mobile devices, and shopping malls, among others. It is necessary to build an efficient model to handle the big dataset for HUSPM. In this paper, we present a four-stages MapReduce framework based on the Spark platform for mining the high-utility sequential patterns from a very large database. From the experimental results, we then can observe that the designed model outperforms the state-of-the-art approaches for handling the very big dataset.
Jerry Chun-Wei Lin, Yuanfa Li, Philippe Fournier-Viger, Youcef Djenouri, Shyue-Liang Wang
IEEE BigData3
2019 A GA-based Framework for Mining High Fuzzy Utility Itemsets
abstract
Comparing to frequent itemset mining (FIM), utility-pattern mining receives increasing attention in the field of data mining recently. With the flourishing development of utility-pattern mining, most studies focused on the efficiency problem by considering the efficient data structure to compress the original data and pruning strategies to reduce the search space for knowledge discovery. However, those approaches can only handle the binary situation, thus the discovered knowledge cannot be represented as the linguistic variables. Previous works have addressed this problem by introducing the generic approaches to find the high fuzzy utility itemsets in a small database. In real-world situations, the dataset may be very large, and it is costly to mine all the required information from a very large database. In this paper, we first present a HFUI-GA framework to discover the high fuzzy utility itemsets in a limited time. Several improvement strategies are also proposed to speed up the evolutionary progress. Experiments are then conducted to show the performance of the variants of the designed HFUI-GA framework in terms of number of the discovered high fuzzy utility itemsets (HFUIs) and the results are convincing to show that the designed GA-based HFUI-GA framework is a promising solution to mine for HFUIs.
Jimmy Ming-Tai Wu, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Tomasz Wiktorski, Tzung-Pei Hong, Matin Pirouz
IEEE BigData3
2019 Finding Strongly Correlated Trends in Dynamic Attributed Graphs
Philippe Fournier-Viger, Zhi Cheng, Jerry Chun-Wei Lin, Nazha Selmaoui-Folcher
DaWaK1
2019 Discovering and Visualizing Efficient Patterns in Cost/Utility Sequences
Philippe Fournier-Viger, Jerry Chun-Wei Lin, Tin Truong 0001
DaWaK1
2019 Succinct BWT-Based Sequence Prediction
Rafael Ktistakis, Philippe Fournier-Viger, Simon J. Puglisi, Rajeev Raman
DEXA (2)2
2019 Efficiently Finding High Utility-Frequent Itemsets Using Cutoff and Suffix Utility
R. Uday Kiran, T. Yashwanth Reddy, Philippe Fournier-Viger, Masashi Toyoda, P. Krishna Reddy, Masaru Kitsuregawa
PAKDD (2)3
2019 Discovering Spatial High Utility Itemsets in Spatiotemporal Databases
abstract
In real-world databases, high utility itemset (HUI) is an important class of regularities. Most previous studies have focused on mining HUIs in transactional databases and did not consider the spatiotemporal characteristics of items. In this study, a more flexible model of spatial HUIs (SHUIs) that exist in spatiotemporal databases is proposed. In a spatiotemporal database (STD), an itemset is said to be an SHUI if its utility is not less than a user-specified minimum utility and the distance between any two of its items is not more than a user-specified maximum distance. Identifying SHUIs is very challenging because the generated itemsets do not satisfy the anti-monotonic property. In this study, we present two novel pruning techniques for reducing computational costs. Moreover, a fast single scan algorithm is presented for effectively evaluating all SHUIs in a STD. Furthermore, two case studies are presented, in which the proposed model is used to identify useful information in traffic congestion data and air pollution data.
R. Uday Kiran, Koji Zettsu, Masashi Toyoda, Philippe Fournier-Viger, P. Krishna Reddy, Masaru Kitsuregawa
SSDBM4
2019 Exploiting GPU parallelism in improving bees swarm optimization for mining big transactional databases
Youcef Djenouri, Djamel Djenouri, Asma Belhadi, Philippe Fournier-Viger, Jerry Chun-Wei Lin, Ahcène Bendjoudi
Inf. Sci.4
2019 Efficient algorithms to identify periodic patterns in multiple sequences
Philippe Fournier-Viger, Zhitian Li, Jerry Chun-Wei Lin, R. Uday Kiran, Hamido Fujita
Inf. Sci.1
2019 Mining local and peak high utility itemsets
Philippe Fournier-Viger, Jerry Chun-Wei Lin, Hamido Fujita, Yun Sing Koh
Inf. Sci.1
2019 BILU-NEMH: A BILU neural-encoded mention hypergraph for mention extraction
abstract
The natural language processing (NLP) denotes a technique used to process data such as text and speech. Some of the fundamental research in NLP includes the named entity recognition, which recognizes the named entities (i.e., persons and companies) from texts, the semantic parsing, which converts a natural language utterance to a logical form, and the co-reference resolution, which extracts the nouns (including pronouns and noun phrases) pointing to the same reference body. In this paper, we focus on the mention extraction and classification, proposing a neural-encoded mention-hypergraph model named the BILU-NEMH to extract the mention entities from a content. The proposed BILU-NEMH model combines a mention hypergraph model with the encoding schema and neural network. The proposed model can effectively capture the overlapping mention entities of an unbounded length. The proposed model was verified by the experiments, and the obtained experimental results showed that the proposed model achieved better performance and greater effectiveness than the existing related models on most standard datasets.
Jerry Chun-Wei Lin, Yinan Shao, Philippe Fournier-Viger, Hamido Fujita
Inf. Sci.3
2019 A Survey of Parallel Sequential Pattern Mining
abstract
With the growing popularity of shared resources, large volumes of complex data of different types are collected automatically. Traditional data mining algorithms generally have problems and challenges including huge memory cost, low processing speed, and inadequate hard disk space. As a fundamental task of data mining, sequential pattern mining (SPM) is used in a wide variety of real-life applications. However, it is more complex and challenging than other pattern mining tasks, i.e., frequent itemset mining and association rule mining, and also suffers from the above challenges when handling the large-scale data. To solve these problems, mining sequential patterns in a parallel or distributed computing environment has emerged as an important issue with many applications. In this article, an in-depth survey of the current status of parallel SPM (PSPM) is investigated and provided, including detailed categorization of traditional serial SPM approaches, and state-of-the art PSPM. We review the related work of PSPM in details including partition-based algorithms for PSPM, apriori-based PSPM, pattern-growth-based PSPM, and hybrid algorithms for PSPM, and provide deep description (i.e., characteristics, advantages, disadvantages, and summarization) of these parallel approaches of PSPM. Some advanced topics for PSPM, including parallel quantitative/weighted/utility SPM, PSPM from uncertain data and stream data, hardware acceleration for PSPM, are further reviewed in details. Besides, we review and provide some well-known open-source software of PSPM. Finally, we summarize some challenges and opportunities of PSPM in the big data era.
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao, Philip S. Yu
ACM Trans. Knowl. Discov. Data3
2019 Efficient Vertical Mining of High Average-Utility Itemsets Based on Novel Upper-Bounds
abstract
Mining High Average-Utility Itemsets (HAUIs) in a quantitative database is an extension of the traditional problem of frequent itemset mining, having several practical applications. Discovering HAUIs is more challenging than mining frequent itemsets using the traditional support model since the average-utilities of itemsets do not satisfy the downward-closure property. To design algorithms for mining HAUIs that reduce the search space of itemsets, prior studies have proposed various upper-bounds on the average-utilities of itemsets. However, these algorithms can generate a huge amount of unpromising HAUI candidates, which result in high memory consumption and long runtimes. To address this problem, this paper proposes four tight average-utility upper-bounds, based on a vertical database representation, and three efficient pruning strategies. Furthermore, a novel generic framework for comparing average-utility upper-bounds is presented. Based on these theoretical results, an efficient algorithm named dHAUIM is introduced for mining the complete set of HAUIs. dHAUIM represents the search space and quickly compute upper-bounds using a novel IDUL structure. Extensive experiments show that dHAUIM outperforms four state-of-the-art algorithms for mining HAUIs in terms of runtime on both real-life and synthetic databases. Moreover, results show that the proposed pruning strategies dramatically reduce the number of candidate HAUIs.
Tin Truong 0001, Hai Duong 0001, Bac Le, Philippe Fournier-Viger
IEEE Trans. Knowl. Data Eng.4
2018 Discovering High Utility Change Points in Customer Transaction Data
Philippe Fournier-Viger, Jerry Chun-Wei Lin, Yun Sing Koh
ADMA1
2018 Discovering Periodic Patterns Common to Multiple Sequences
Philippe Fournier-Viger, Zhitian Li, Jerry Chun-Wei Lin, R. Uday Kiran, Hamido Fujita
DaWaK1
2018 Anonymization of Multiple and Personalized Sensitive Attributes
Jerry Chun-Wei Lin, Qiankun Liu 0002, Philippe Fournier-Viger, Youcef Djenouri, Ji Zhang 0001
DaWaK3
2018 Mining Local High Utility Itemsets
Philippe Fournier-Viger, Jerry Chun-Wei Lin, Hamido Fujita, Yun Sing Koh
DEXA (2)1
2018 A Metaheuristic Algorithm for Hiding Sensitive Itemsets
Jerry Chun-Wei Lin, Yuyu Zhang, Philippe Fournier-Viger, Youcef Djenouri, Ji Zhang 0001
DEXA (2)3
2018 Mining diversified association rules in big datasets: A cluster/GPU/genetic approach
Youcef Djenouri, Asma Belhadi, Philippe Fournier-Viger, Hamido Fujita
Inf. Sci.3
2018 Fast and effective cluster-based information retrieval using frequent closed itemsets
Youcef Djenouri, Asma Belhadi, Philippe Fournier-Viger, Jerry Chun-Wei Lin
Inf. Sci.3
2018 Exploiting highly qualified pattern with frequency and weight occupancy
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao, Justin Zhijun Zhan, Ji Zhang 0001
Knowl. Inf. Syst.3
2017 Extracting Non-redundant Correlated Purchase Behaviors by Utility Measure
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao
DaWaK3
2017 Mining High-Utility Itemsets with Both Positive and Negative Unit Profits from Uncertain Databases
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao, Vincent S. Tseng
PAKDD (1)3
2017 Discovering Periodic Patterns in Non-uniform Temporal Databases
R. Uday Kiran, J. N. Venkatesh, Philippe Fournier-Viger, Masashi Toyoda, P. Krishna Reddy, Masaru Kitsuregawa
PAKDD (2)3
2017 A two-phase approach to mine short-period high-utility itemsets in transactional databases
Jerry Chun-Wei Lin, Jiexiong Zhang, Philippe Fournier-Viger, Tzung-Pei Hong, Ji Zhang 0001
Adv. Eng. Informatics3
2017 An efficient algorithm for mining top-k on-shelf high utility itemsets
Thu-Lan Dam, Kenli Li 0001, Philippe Fournier-Viger, Quang-Huy Duong
Knowl. Inf. Syst.3
2017 FCloSM, FGenSM: two efficient algorithms for mining frequent closed and generator sequences using the local pruning strategy
Bac Le, Hai Duong 0001, Tin Truong 0001, Philippe Fournier-Viger
Knowl. Inf. Syst.4
2017 FDHUP: Fast algorithm for mining discriminative high utility patterns
Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong, Han-Chieh Chao
Knowl. Inf. Syst.3
2017 EFIM: a fast and memory efficient algorithm for high-utility itemset mining
Souleymane Zida, Philippe Fournier-Viger, Jerry Chun-Wei Lin, Cheng-Wei Wu, Vincent S. Tseng
Knowl. Inf. Syst.2
2016 Mining Discriminative High Utility Patterns
Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong
ACIIDS (2)3
2016 Efficient Mining of Fuzzy Frequent Itemsets with Type-2 Membership Functions
Jerry Chun-Wei Lin, Xianbiao Lv, Philippe Fournier-Viger, Tsu-Yang Wu, Tzung-Pei Hong
ACIIDS (2)3
2016 Mining Recent High Expected Weighted Itemsets from Uncertain Databases
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao
APWeb (1)3
2016 Mining Recent High-Utility Patterns from Temporal Databases with Time-Sensitive Constraint
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao
DaWaK3
2016 Mining Minimal High-Utility Itemsets
Philippe Fournier-Viger, Jerry Chun-Wei Lin, Cheng-Wei Wu, Vincent S. Tseng, Usef Faghihi
DEXA (1)1
2016 More Efficient Algorithms for Mining High-Utility Itemsets with Multiple Minimum Utility Thresholds
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao
DEXA (1)3
2016 The SPMF Open-Source Data Mining Library Version 2
Philippe Fournier-Viger, Jerry Chun-Wei Lin, Antonio Gomariz, Ted Gueniche, Azadeh Soltani, Zhi-Hong Deng 0001, Hoang Thanh Lam
ECML/PKDD (3)1
2016 More Efficient Algorithm for Mining Frequent Patterns with Multiple Minimum Supports
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao
WAIM (1)3
2016 Efficient Mining of Uncertain Data for High-Utility Itemsets
Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong, Vincent S. Tseng
WAIM (1)3
2016 Fast algorithms for mining high-utility itemsets with various discount strategies
Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong, Vincent S. Tseng
Adv. Eng. Informatics3
2016 An efficient algorithm to mine high average-utility itemsets
Jerry Chun-Wei Lin, Ting Li 0011, Philippe Fournier-Viger, Tzung-Pei Hong, Justin Zhijun Zhan, Miroslav Voznak
Adv. Eng. Informatics3
2016 Inferring social network user profiles using a partial social graph
Raïssa Yapan Dougnon, Philippe Fournier-Viger, Jerry Chun-Wei Lin, Roger Nkambou
J. Intell. Inf. Syst.2
2016 Efficient Algorithms for Mining Top-K High Utility Itemsets
abstract
High utility itemsets (HUIs) mining is an emerging topic in data mining, which refers to discovering all itemsets having a utility meeting a user-specified minimum utility threshold min_util. However, setting min_util appropriately is a difficult problem for users. Generally speaking, finding an appropriate minimum utility threshold by trial and error is a tedious process for users. If min_util is set too low, too many HUIs will be generated, which may cause the mining process to be very inefficient. On the other hand, if min_util is set too high, it is likely that no HUIs will be found. In this paper, we address the above issues by proposing a new framework for top-k high utility itemset mining, where k is the desired number of HUIs to be mined. Two types of efficient algorithms named TKU (mining Top-K Utility itemsets) and TKO (mining Top-K utility itemsets in One phase) are proposed for mining such itemsets without the need to set min_util. We provide a structural comparison of the two algorithms with discussions on their advantages and limitations. Empirical evaluations on both real and synthetic datasets show that the performance of the proposed algorithms is close to that of the optimal case of state-of-the-art utility mining algorithms.
Vincent S. Tseng, Cheng-Wei Wu, Philippe Fournier-Viger, Philip S. Yu
IEEE Trans. Knowl. Data Eng.3
2015 Mining Weighted Frequent Itemsets with the Recency Constraint
Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong
APWeb3
2015 Mining high-utility itemsets with various discount strategies
abstract
In recent years, mining high-utility itemsets (HUIs) has become as a key topic in data mining. However, most of the developed algorithms assume the unrealistic situations that unit profits of items remain unchanged over time. But in real-life situations, the profit of an item or itemset varies as a function of cost prices, sales prices and sales strategies. In this paper, a novel framework for mining HUIs with two algorithms under various Discount strategies (HUID) are introduced. HUID-tp is based on various discount strategies and a novel downward closure property to mine the complete set of HUIs. HUID-Miner is an algorithm relying on a compact data structure (Positive-and-Negative Utility-list, PNU-list) and new pruning strategies to efficiently discover HUIs without candidate generation, while considerably reducing the size of the search space. Furthermore, a strategy named Estimated Utility Co-occurrence Strategy which stores the relationships between 2-itemsets is also adopted in the proposed improvement HUID-EMiner algorithm to speed up computation. An extensive experimental study carried on several real-life datasets shows the performance of the proposed algorithms.
Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong, Vincent S. Tseng
DSAA3
2015 CPT+: Decreasing the Time/Space Complexity of the Compact Prediction Tree
Ted Gueniche, Philippe Fournier-Viger, Rajeev Raman, Vincent S. Tseng
PAKDD (2)2
2015 Mining Partially-Ordered Sequential Rules Common to Multiple Sequences
abstract
Sequential rule mining is an important data mining problem with multiple applications. An important limitation of algorithms for mining sequential rules common to multiple sequences is that rules are very specific and therefore many similar rules may represent the same situation. This can cause three major problems: (1) similar rules can be rated quite differently, (2) rules may not be found because they are individually considered uninteresting, and (3) rules that are too specific are less likely to be used for making predictions. To address these issues, we explore the idea of mining “partially-ordered sequential rules” (POSR), a more general form of sequential rules such that items in the antecedent and the consequent of each rule are unordered. To mine POSR, we propose the RuleGrowth algorithm, which is efficient and easily extendable. In particular, we present an extension (TRuleGrowth) that accepts a sliding-window constraint to find rules occurring within a maximum amount of time. A performance study with four real-life datasets show that RuleGrowth and TRuleGrowth have excellent performance and scalability compared to baseline algorithms and that the number of rules discovered can be several orders of magnitude smaller when the sliding-window constraint is applied. Furthermore, we also report results from a real application showing that POSR can provide a much higher prediction accuracy than regular sequential rules for sequence prediction.
Philippe Fournier-Viger, Cheng-Wei Wu, Vincent S. Tseng, Longbing Cao, Roger Nkambou
IEEE Trans. Knowl. Data Eng.1
2015 Efficient Algorithms for Mining the Concise and Lossless Representation of High Utility Itemsets
abstract
Mining high utility itemsets (HUIs) from databases is an important data mining task, which refers to the discovery of itemsets with high utilities (e.g. high profits). However, it may present too many HUIs to users, which also degrades the efficiency of the mining process. To achieve high efficiency for the mining task and provide a concise mining result to users, we propose a novel framework in this paper for mining closed+high utility itemsets(CHUIs), which serves as a compact and lossless representation of HUIs. We propose three efficient algorithms named AprioriCH (Apriori-based algorithm for mining High utility Closed+itemsets), AprioriHC-D (AprioriHC algorithm with Discarding unpromising and isolated items) and CHUD (Closed+High Utility Itemset Discovery) to find this representation. Further, a method called DAHU (Derive All High Utility Itemsets) is proposed to recover all HUIs from the set of CHUIs without accessing the original database. Results on real and synthetic datasets show that the proposed algorithms are very efficient and that our approaches achieve a massive reduction in the number of HUIs. In addition, when all HUIs can be recovered by DAHU, the combination of CHUD and DAHU outperforms the state-of-the-art algorithms for mining HUIs.
Vincent S. Tseng, Cheng-Wei Wu, Philippe Fournier-Viger, Philip S. Yu
IEEE Trans. Knowl. Data Eng.3
2014 FHN: Efficient Mining of High-Utility Itemsets with Negative Unit Profits
Philippe Fournier-Viger
ADMA1
2014 Novel Concise Representations of High Utility Itemsets Using Generator Patterns
Philippe Fournier-Viger, Cheng-Wei Wu, Vincent S. Tseng
ADMA1
2014 VGEN: Fast Vertical Mining of Sequential Generator Patterns
Philippe Fournier-Viger, Antonio Gomariz, Michal Sebek, Martin Hlosta
DaWaK1
2014 ERMiner: Sequential Rule Mining Using Equivalence Classes
Philippe Fournier-Viger, Ted Gueniche, Souleymane Zida, Vincent S. Tseng
IDA1
2014 Fast Vertical Mining of Sequential Patterns Using Co-occurrence Information
Philippe Fournier-Viger, Antonio Gomariz, Manuel Campos, Rincy Thomas
PAKDD (1)1
2013 TKS: Efficient Mining of Top-K Sequential Patterns
Philippe Fournier-Viger, Antonio Gomariz, Ted Gueniche, Espérance Mwamikazi, Rincy Thomas
ADMA (1)1
2013 MEIT: Memory Efficient Itemset Tree for Targeted Association Rule Mining
Philippe Fournier-Viger, Espérance Mwamikazi, Ted Gueniche, Usef Faghihi
ADMA (2)1
2013 Mining Maximal Sequential Patterns without Candidate Maintenance
Philippe Fournier-Viger, Cheng-Wei Wu, Vincent S. Tseng
ADMA (1)1
2013 Compact Prediction Tree: A Lossless Model for Accurate Sequence Prediction
Ted Gueniche, Philippe Fournier-Viger, Vincent S. Tseng
ADMA (2)2
2012 Using Partially-Ordered Sequential Rules to Generate More Accurate Sequence Prediction
Philippe Fournier-Viger, Ted Gueniche, Vincent S. Tseng
ADMA1
2011 Mining Top-K Sequential Rules
Philippe Fournier-Viger, Vincent S. Tseng
ADMA (2)1
2011 Efficient Mining of a Concise and Lossless Representation of High Utility Itemsets
abstract
Mining high utility item sets from transactional databases is an important data mining task, which refers to the discovery of item sets with high utilities (e.g. high profits). Although several studies have been carried out, current methods may present too many high utility item sets for users, which degrades the performance of the mining task in terms of execution and memory efficiency. To achieve high efficiency for the mining task and provide a concise mining result to users, we propose a novel framework in this paper for mining closed+ high utility item sets, which serves as a compact and loss less representation of high utility item sets. We present an efficient algorithm called CHUD (Closed+ High Utility item set Discovery) for mining closed+ high utility item sets. Further, a method called DAHU (Derive All High Utility item sets) is proposed to recover all high utility item sets from the set of closed+ high utility item sets without accessing the original database. Results of experiments on real and synthetic datasets show that CHUD and DAHU are very efficient with a massive reduction (up to 800 times in our experiments) in the number of high utility item sets. In addition, when all high utility item sets are recovered by DAHU, the approach combining CHUD and DAHU also outperforms the state-of-the-art algorithms in mining high utility item sets.
Cheng-Wei Wu, Philippe Fournier-Viger, Philip S. Yu, Vincent S. Tseng
ICDM2