VLDB 2026 Research / reviewers in the wild / expert
Wensheng Gan
dblp:145/5903
· DBLP profile ↗
83ranked-venue papers in the field
21as first author
59since 2021 · last 2026
0000-0002-5781-8116ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 28 (7 first)Data Mining & Knowledge Discovery · 23 (6 first)Knowledge Engineering, Semantic Web & Information Systems · 16 (2 first)Database Systems & Data Management · 8 (5 first)Information Retrieval & Web Search · 5 (1 first)Other / Interdisciplinary · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Targeted mining of non-overlapping high-utility sequential patterns
Wensheng Gan, Zhidong Lin, Zhenlian Qi, Jian Zhu 0001, Ruichu Cai, Zhifeng Hao 0004 |
Inf. Sci. | 2 |
| 2026 | SeqRFM: Fast RFM analysis in sequence data
Yanxin Zheng, Wensheng Gan, Pinlyu Zhou, Philippe Fournier-Viger |
Inf. Sci. | 2 |
| 2026 | High-Utility Sequential Rule Mining Utilizing Segmentation Guided by ConfidenceabstractWithin the domain of data mining, one critical objective is the discovery of sequential rules with high utility. The goal is to discover sequential rules that exhibit both high utility and strong confidence, which are valuable in real-world applications. However, existing high-utility sequential rule mining algorithms suffer from redundant utility computations, as different rules may consist of the same sequence of items. When these items can form multiple distinct rules, additional utility calculations are required. To address this issue, this study proposes a sequential rule mining algorithm that utilizes segmentation guided by confidence (RSC), which employs confidence-guided segmentation to reduce redundant utility computation. It adopts a method that precomputes the confidence of segmented rules by leveraging the support of candidate subsequences in advance. Once the segmentation point is determined, all rules with different antecedents and consequents are generated simultaneously. RSC uses a utility-linked table to accelerate candidate sequence generation and introduces a stricter utility upper bound, called the reduced remaining utility of a sequence, to address sequences with duplicate items. Finally, the proposed RSC method was evaluated on multiple datasets, and the results demonstrate improvements over state-of-the-art approaches. Chunkai Zhang, Jiarui Deng, Maohua Lyu, Wensheng Gan, Philip S. Yu |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | Large Language Models for Fault Diagnosis
Zhenlian Qi, Junyu Ren, Wensheng Gan, Philip S. Yu |
IEEE Big Data | 3 |
| 2025 | AI-Driven Log Analysis: Advances and Challenges
Yongheng Wang, Wensheng Gan, Philip S. Yu |
IEEE Big Data | 2 |
| 2025 | Large Language Models for Bioinformatics: Applications and Challenges
Wenxi Zhu, Wensheng Gan, Zhenlian Qi, Philip S. Yu |
IEEE Big Data | 2 |
| 2025 | A generic framework for mining sequences with various interestingness measures in dynamic attributed graphs
Jiayu Cai, Guoting Chen, Wensheng Gan |
Knowl. Inf. Syst. | 4 |
| 2025 | Mining high utility contrast patterns in sequences
Chunkai Zhang, Yuting Yang 0005, Ryan Han-Yuan Zhang, Wensheng Gan, Philip S. Yu |
Knowl. Inf. Syst. | 5 |
| 2025 | Graph Contrastive Learning on Multi-label Classification for RecommendationsabstractIn business analysis, providing effective recommendations is crucial for boosting company profits. Graph structures, especially bipartite graphs, are favored for analyzing complex data relationships. Link prediction is crucial for recommending specific items to users. Traditional methods have primarily focused on binary classification tasks. These methods, which identify patterns in graph structures or use representation techniques like graph neural networks (GNNs), face challenges with increasing data volume and label count. Data growth strains system performance and efficiency. More labels intensify data sparsity, as users and items focus on only a few labels, leading to sparse matrices that hamper recommendation algorithms. To tackle these issues, we introduce the Graph Contrastive Learning for Multi-label Classification (MCGCL) model. It uses contrastive learning to improve recommendations and has two training phases: a main task of holistic user–item graph learning to grasp user–item relationships, and a subtask of constructing homogeneous user–user (item–item) subgraphs to capture user–user and item–item relationships. Comparative experiments with state-of-the-art methods confirm the effectiveness of MCGCL, highlighting its potential for improving recommendation systems. Jiayang Wu 0001, Wensheng Gan, Huashen Lu, Philip S. Yu |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2025 | Towards Sequence Utility Maximization under Utility Occupancy MeasureabstractThe discovery of utility-driven patterns is a valuable and difficult research topic. It can extract significant and interesting information from specific and varied databases, increasing the value of the services provided. In practice, the utility measure is often used to reflect the importance, profit, or risk of an object or pattern. In the database, while utility is a flexible criterion for patterns, it is also a somewhat limited criterion due to the overlook of utility sharing. This leads to the derived patterns only exploring partial and local knowledge in the database. Utility occupancy considers the problem of mining with high utility but low occupancy. However, existing studies are focused on itemsets that cannot reveal the temporal relationship of object occurrences. Therefore, this article first defines the concept of utility occupancy of sequence data and raises the problem of High-Utility Occupancy Sequential Pattern Mining (HUOSPM). Three dimensions, including frequency, utility, and occupancy, are comprehensively evaluated in HUOSPM. An algorithm called Sequence Utility Maximization with Utility occupancy measure (SUMU) is proposed. Furthermore, two data structures for storing pattern-related information, including Utility-Occupancy-List-Chain (UOL-Chain) and Utility-Occupancy-Table (UO-Table), are designed, and six upper bounds are proposed to improve efficiency. Extensive experiments are conducted to evaluate the efficiency and effectiveness of the novel algorithm. A specific case study is provided, and the effects of different upper bounds and pruning strategies are analyzed. The comprehensive results suggest that the HUOSPM task is useful and efficient. Gengsen Huang, Wensheng Gan, Philip S. Yu |
ACM Trans. Knowl. Discov. Data | 2 |
| 2025 | Towards Target Sequential RulesabstractIn many real-world applications, sequential rule mining (SRM) can offer prediction and recommendation functions for a variety of services. It is an important technique of pattern mining to discover all valuable rules that can reveal the temporal relationship between objects. Although several algorithms of SRM are proposed to solve various practical problems, there are no studies on the problem of targeted mining. Targeted sequential rule mining aims to obtain those interesting sequential rules that users focus on, thus avoiding the generation of other invalid and unnecessary rules. It can further improve the efficiency of users in analyzing rules and reduce the consumption of computing resources. In this paper, we first present the relevant definitions of target sequential rules and formulate the problem of targeted sequential rule mining. Then, we propose an efficient algorithm called TaSRM. Several pruning strategies and an optimization are introduced to improve the efficiency of TaSRM. Finally, a large number of experiments are conducted on different benchmarks, and we analyze the results in terms of running time, memory consumption, and scalability, as well as query cases with different query rules. It is shown that the novel algorithm TaSRM and its variants can achieve better experimental performance compared to the baseline algorithm. Wensheng Gan, Gengsen Huang, Jian Weng 0001, Tianlong Gu, Philip S. Yu |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2025 | Data Scarcity in Recommendation Systems: A SurveyabstractThe prevalence of online content has led to the widespread adoption of recommendation systems (RSs), which serve diverse purposes such as news, advertisements, and e-commerce recommendations. Despite their significance, data scarcity issues have significantly impaired the effectiveness of existing RS models and hindered their progress. To address this challenge, the concept of knowledge transfer, particularly from external sources like pre-trained language models, emerges as a potential solution to alleviate data scarcity and enhance RS development. However, the practice of knowledge transfer in RSs is intricate. Transferring knowledge between domains introduces data disparities, and the application of knowledge transfer in complex RS scenarios can yield negative consequences if not carefully designed. Therefore, this article contributes to this discourse by addressing the implications of data scarcity on RSs and introducing various strategies, such as data augmentation, self-supervised learning, transfer learning, broad learning, and knowledge graph utilization, to mitigate this challenge. Furthermore, it delves into the challenges and future direction within the RS domain, offering insights that are poised to facilitate the development and implementation of robust RSs, particularly when confronted with data scarcity. We aim to provide valuable guidance and inspiration for researchers and practitioners, ultimately driving advancements in the field of RS. Wensheng Gan, Jiayang Wu 0001, Kaixia Hu |
Trans. Recomm. Syst. | 2 |
| 2024 | RFMI-based Customer Segmentation with K-meansabstractThe development of e-marketing over recent decades has led offline and online retail enterprises to adopt various data analysis technologies to enhance their understanding of consumer behavior and increase revenue. One common approach involves segmenting consumers into distinct groups based on designed metrics, targeting high-value segments for specialized services. To evaluate customer worthiness, the popular RFM model uses three dimensions: recency (the time since their last purchase), frequency (how often they make purchases), and monetary value (total spending). Higher scores under this model are indicative of greater potential profitability for businesses. While this approach provides valuable insights, it may not fully capture all profitable customer behaviors accurately. To address these limitations, this paper introduces a new model, namely the RFMI (i.e., recency, frequency, monetary, and interval) model, for comprehensively evaluating customer value. The new model employs an analytic hierarchy process to derive the RFMI values of customers. Subsequently, we employ K-means clustering customers to group customers into six segments. Moreover, the experimental dataset was sourced from a real UK e-commerce platform. The experimental results indicate that the new model effectively distinguishes between various consumption patterns among customers. This enhanced understanding can enable retailers to improve their marketing strategies more precisely, optimize customer service, and increase profitability. Wensheng Gan, Pinlyu Zhou, Shicheng Wan, Jiyuan Zeng, Zhenlian Qi |
IEEE Big Data | 1 |
| 2024 | FCSG-Miner: Frequent closed subgraph mining in multi-graphs
Jiayu Cai, Guoting Chen, Wensheng Gan, Amaël Broustet |
Inf. Sci. | 4 |
| 2024 | Towards episode rules with non-overlapping frequency and targeted mining
Wensheng Gan |
Inf. Sci. | 2 |
| 2024 | HUSM: High utility subgraph mining in single graph databases
Zhaoming Chen 0001, Guoting Chen, Wensheng Gan, Philippe Fournier-Viger |
Inf. Sci. | 4 |
| 2024 | Privacy preserving rare itemset mining
Yijie Gui, Wensheng Gan, Yongdong Wu, Philip S. Yu |
Inf. Sci. | 2 |
| 2024 | Targeted mining of contiguous sequential patterns
Kaixia Hu, Wensheng Gan, Shan Huang 0009, Philippe Fournier-Viger |
Inf. Sci. | 2 |
| 2024 | Mining frequent temporal duration-based patterns on time interval sequential database
Fuyin Lai, Guoting Chen, Wensheng Gan, Mengfeng Sun |
Inf. Sci. | 3 |
| 2024 | TaSPM: Targeted Sequential Pattern MiningabstractSequential pattern mining (SPM) is an important technique in the field of pattern mining, which has many applications in reality. Although many efficient SPM algorithms have been proposed, there are few studies that can focus on targeted tasks. Targeted querying of the concerned sequential patterns can not only reduce the number of patterns generated, but also increase the efficiency of users in performing related analysis. The current algorithms available for targeted sequence querying are based on specific scenarios and can not be extended to other applications. In this article, we formulate the problem of targeted sequential pattern mining and propose a generic algorithm, namely TaSPM. What is more, to improve the efficiency of TaSPM on large-scale datasets and multiple-item-based sequence datasets, we propose several pruning strategies to reduce meaningless operations in the mining process. Totally four pruning strategies are designed in TaSPM, and hence TaSPM can terminate unnecessary pattern extensions quickly and achieve better performance. Finally, we conducted extensive experiments on different datasets to compare the baseline SPM algorithm with TaSPM. Experiments show that the novel targeted mining algorithm TaSPM can achieve faster running time and less memory consumption. Gengsen Huang, Wensheng Gan, Philip S. Yu |
ACM Trans. Knowl. Discov. Data | 2 |
| 2024 | Totally-ordered Sequential Rules for Utility MaximizationabstractHigh-utility sequential pattern mining (HUSPM) is a significant and valuable activity in knowledge discovery and data analytics with many real-world applications. In some cases, HUSPM can not provide an excellent measure to predict what will happen. High-utility sequential rule mining (HUSRM) discovers high utility and high confidence sequential rules, so it can solve the issue in HUSPM. However, all existing HUSRM algorithms aim to find high-utility partially-ordered sequential rules (HUSRs), which are not consistent with reality and may generate fake HUSRs. Therefore, in this article, we formulate the problem of high-utility totally-ordered sequential rule mining and propose a novel algorithm, called TotalSR, which aims to identify all high-utility totally-ordered sequential rules (HTSRs). TotalSR introduces a left-first expansion strategy that can utilize the anti-monotonic property to use a confidence pruning strategy. TotalSR also designs a new utility upper bound: RSPEU , which is tighter than the existing upper bounds. TotalSR can drastically reduce the search space with the help of utility upper bounds pruning strategies, avoiding much more meaningless computation. To effectively compute the information, TotalSR proposes an auxiliary antecedent record table that can efficiently calculate the antecedent’s support and a utility prefix sum list that can compute the upper bound in O (1) time for a sequence. Finally, there are numerous experimental results on both real and synthetic datasets demonstrating that TotalSR is more efficient than the existing algorithms. Chunkai Zhang, Maohua Lyu, Wensheng Gan, Philip S. Yu |
ACM Trans. Knowl. Discov. Data | 3 |
| 2024 | HUSP-SP: Faster Utility Mining on Sequence DataabstractHigh-utility sequential pattern mining (HUSPM) has emerged as an important topic due to its wide application and considerable popularity. However, due to the combinatorial explosion of the search space when the HUSPM problem encounters a low-utility threshold or large-scale data, it may be time-consuming and memory-costly to address the HUSPM problem. Several algorithms have been proposed for addressing this problem, but they still cost a lot in terms of running time and memory usage. In this article, to further solve this problem efficiently, we design a compact structure called sequence projection (seqPro) and propose an efficient algorithm, namely, discovering high-utility sequential patterns with the seqPro structure (HUSP-SP). HUSP-SP utilizes the compact seq-array to store the necessary information in a sequence database. The seqPro structure is designed to efficiently calculate candidate patterns’ utilities and upper-bound values. Furthermore, a new upper bound on utility, namely, tighter reduced sequence utility and two pruning strategies in search space, are utilized to improve the mining performance of HUSP-SP. Experimental results on both synthetic and real-life datasets show that HUSP-SP can significantly outperform the state-of-the-art algorithms in terms of running time, memory usage, search space pruning efficiency, and scalability. Chunkai Zhang, Yuting Yang 0005, Zilin Du, Wensheng Gan, Philip S. Yu |
ACM Trans. Knowl. Discov. Data | 4 |
| 2023 | Frequent Subgraph Mining in Dynamic DatabasesabstractFrequent subgraph mining is fundamental in graph mining, with wide-ranging applications in domains such as biology, chemistry, and social network analysis. Most existing algorithms are tailored for static graph databases. Real-world databases often exhibit dynamic attributes, such as data that may change over time. Existing methods for mining frequent subgraphs in databases with dynamic attributes primarily cater to dynamic graph databases, in which graphs evolve over time. However, in practice, a category of graph databases allows for adding or removing graphs. We refer to these databases as dynamic ones, which can be incrementally or decrementally updated while the remaining graphs do not change. This paper introduces frequent subgraph mining in this type of database and proposes the corresponding algorithm called DyFSM. We design a set called Fringe, which comprises DMFSand DMIS. DMFSis a novel concise representation based on the DFS code and can efficiently recover all frequent subgraphs. DMISis a set of subgraphs from which all infrequent subgraphs can be extended. Fringefacilitates updating frequent subgraphs in the renewed database. In our experiments, we collect four real-world graph datasets and conduct experiments using DyFSM. The results validate the accuracy and show good performance of our algorithm. Zhaoming Chen 0001, Guoting Chen, Wensheng Gan |
IEEE Big Data | 4 |
| 2023 | Large Language Models in Education: Vision and OpportunitiesabstractWith the rapid development of artificial intelligence technology, large language models (LLMs) have become a hot research topic. Education plays an important role in human social development and progress. Traditional education faces challenges such as individual student differences, insufficient allocation of teaching resources, and assessment of teaching effectiveness. Therefore, the applications of LLMs in the field of digital/smart education have broad prospects. The research on educational large models (EduLLMs) is constantly evolving, providing new methods and approaches to achieve personalized learning, intelligent tutoring, and educational assessment goals, thereby improving the quality of education and the learning experience. This article aims to investigate and summarize the application of LLMs in smart education. It first introduces the research background and motivation of LLMs and explains the essence of LLMs. It then discusses the relationship between digital education and EduLLMs and summarizes the current research status of educational large models. The main contributions are the systematic summary and vision of the research background, motivation, and application of large models for education (LLM4Edu). By reviewing existing research, this article provides guidance and insights for educators, researchers, and policy-makers to gain a deep understanding of the potential and challenges of LLM4Edu. It further provides guidance for further advancing the development and application of LLM4Edu, while still facing technical, ethical, and practical challenges requiring further research and exploration. Wensheng Gan, Zhenlian Qi, Jiayang Wu 0001, Jerry Chun-Wei Lin |
IEEE Big Data | 1 |
| 2023 | Model-as-a-Service (MaaS): A SurveyabstractDue to the increased number of parameters and data in the pre-trained model exceeding a certain level, a foundation model (e.g., a large language model) can significantly improve downstream task performance and emerge with some novel special abilities (e.g., deep learning, complex reasoning, and human alignment) that were not present before. Foundation models are a form of generative artificial intelligence (GenAI), and Model-as-a-Service (MaaS) has emerged as a groundbreaking paradigm that revolutionizes the deployment and utilization of GenAI models. MaaS represents a paradigm shift in how we use AI technologies and provides a scalable and accessible solution for developers and users to leverage pre-trained AI models without the need for extensive infrastructure or expertise in model training. In this paper, the introduction aims to provide a comprehensive overview of MaaS, its significance, and its implications for various industries. We provide a brief review of the development history of “X-as-a-Service” based on cloud computing and present the key technologies involved in MaaS. The development of GenAI models will become more democratized and flourish. We also review recent application studies of MaaS. Finally, we highlight several challenges and future issues in this promising area. MaaS is a new deployment and service paradigm for different AI-based models. We hope this review will inspire future research in the field of MaaS. Wensheng Gan, Shicheng Wan, Philip S. Yu |
IEEE Big Data | 1 |
| 2023 | ODTT: Optimized Dynamic Taxonomy Tree with Differential PrivacyabstractFor cybersecurity, privacy protection in big data has received more and more attention and research. Differential privacy is one of the important privacy protection methods, and our work pays attention to differential privacy based on the dynamic taxonomy tree, which can protect the publishing of set-valued data effectively. We propose the optimized dynamic taxonomy tree (ODTT) algorithm as a better and more general way to protect privacy in set-valued datasets. It makes better use of data and reduces noise compared to other privacy-preserving algorithms that use taxonomy tree partitioning. The previous algorithm did not make full use of the characteristics of the dataset when constructing the taxonomy tree, so a 2-itemset’s matrix is used in the proposed algorithm to increase the pseudoempty nodes and reduce the addition of noise. More importantly, we apply the consistency constraint method to the construction of the ODTT algorithm. This retains more statistical characteristics of the original dataset by constraining the noise counts in the leaf partitions of the partition tree. Furthermore, ODTT is extended to deal with dynamic datasets. Finally, we compare the proposed ODTT algorithm with the state-of-the-art CDTT algorithm, by performing a series of experiments and using some evaluation metrics. Experimental results show that ODTT is more general and has higher usability while satisfying the security of differential privacy. Yijie Gui, Wensheng Gan, Yongdong Wu |
IEEE Big Data | 3 |
| 2023 | Targeted Querying of Closed High-Utility ItemsetsabstractIn the era of big data, targeted querying of interesting itemsets having concise expressions is promising for improving the efficiency and capability of data mining applications. As the full set of high-utility itemsets is no longer explored, mining becomes more efficient. Nonetheless, identifying whether the current itemset includes the target pattern and is closed remains a challenge for enhancing data analysis efficiency. At present, no single-phase algorithm reliably identifies the targeted closed high-utility itemsets. In this article, we propose an algorithm called TQCUI, Targeted Querying of Closed high-Utility Itemsets containing the target patterns in a transactional database. The algorithm employs a compact attribute-utility-list structure for maintaining the utility and attribute information of the itemsets. Additionally, TQCUI utilizes several efficient pruning strategies to filter out unpromising itemsets, substantially reducing the search space. Moreover, to quickly prune non-closed itemsets, TQCUI introduces forward-extension and backward-extension checking schemes for the closure checking of itemsets. Extensive experimentation on both real and synthetic datasets demonstrates the TQCUI algorithm has good performance in terms of runtime, memory consumption, and scalability. Shan Huang 0009, Wensheng Gan, Jinbao Miao |
IEEE Big Data | 2 |
| 2023 | Interaction in Metaverse: A SurveyabstractHuman-computer interaction (HCI) emerged with the birth of the computer and has been upgraded through decades of development. Metaverse has attracted a lot of interest with its immersive experience, and HCI is the entrance to the Metaverse for people. It is predictable that HCI will determine the immersion of the Metaverse. However, the technologies of HCI in Metaverse are not mature enough. There are many issues that we should address for HCI in the Metaverse. To this end, the purpose of this paper is to provide a systematic literature review on the key technologies and applications of HCI in the Metaverse. This paper is a comprehensive survey of HCI for the Metaverse, focusing on current technology, future directions, and challenges. First, we provide a brief overview of HCI in the Metaverse and their mutually exclusive relationships. Then, we summarize the evolution of HCI and its future characteristics in the Metaverse. Next, we envision and present the key technologies involved in HCI in the Metaverse. We also review recent case studies of HCI in the Metaverse. Finally, we highlight several challenges and future issues in this promising area. Zirun Gan, Wensheng Gan, Zhenlian Qi, Yuehua Wang, Philip S. Yu |
IEEE Big Data | 3 |
| 2023 | USER: Towards High-Utility Sequential Rules with Repetitive ItemsabstractDiscovering interesting sequential rules in the sequence database is quite important for a variety of fields, ranging from customer behavior analysis to intrusion detection. High utility sequential rule mining (HUSRM) was proposed to obtain more informative rules. Its goal is to find those sequential rules with high utility values and high confidence, i.e., HUSRs. As far as we know, a few algorithms are proposed to discover HUSRs. However, these algorithms do not fully consider the existence of repetitive items in the sequences of the database. In this paper, we propose an algorithm named USER to discover HUSRs in multi-sequences with the existence of repetitive items. A data structure called an occurrence information (OI)-list is designed to distinguish the different occurrences of items in a sequence. Moreover, the change in the upper bound value after the rule expansion is discussed in detail, which is complicated by the repetitive items. We also introduce two pruning strategies (ROOR and REIO-I) to optimize mining efficiency when there are too many repetitive items in the sequence. Finally, we conduct experiments on several datasets, and the results show that USER is able to discover HUSRs with more accurate utility values in an acceptable amount of time and memory consumption. Wensheng Gan, Gengsen Huang, Philip S. Yu |
IEEE Big Data | 2 |
| 2023 | Multimodal Large Language Models: A SurveyabstractThe exploration of multimodal language models integrates multiple data types, such as images, text, language, audio, and other heterogeneity. While the latest large language models excel in text-based tasks, they often struggle to understand and process other data types. Multimodal models address this limitation by combining various modalities, enabling a more comprehensive understanding of diverse data. This paper begins by defining the concept of multimodal and examining the historical development of multimodal algorithms. Furthermore, we introduce a range of multimodal products, focusing on the efforts of major technology companies. A practical guide is provided, offering insights into the technical aspects of multimodal models. Moreover, we present a compilation of the latest algorithms and commonly used datasets, providing researchers with valuable resources for experimentation and evaluation. Lastly, we explore the applications of multimodal models and discuss the challenges associated with their development. By addressing these aspects, this paper aims to facilitate a deeper understanding of multimodal models and their potentiality in various domains. Jiayang Wu 0001, Wensheng Gan, Shicheng Wan, Philip S. Yu |
IEEE Big Data | 2 |
| 2023 | Mining Rare Utility Patterns within Target ItemsabstractAs a crucial subfield of pattern discovery, high utility rare itemset mining (HURIM) is developed to discover abnormal but significant patterns. HURIM plays a vital role in various scenarios, such as network security, disease detection, and biomedicine. However, the traditional HURIM algorithms ignore the users’ demands, which generates massive needless patterns. In general, target-based HURIM algorithms can discover more useful information that meets the needs of users than traditional HURIM algorithms. To this end, we propose a targeted HURIM algorithm called Mining Rare Utility Patterns within Target Items (TIRUP). TIRUP adopts two techniques (projection and merging technologies), to diminish the consumption of database scanning. To effectively improve the performance of TIRUP, this paper utilizes several strategies based on frequency, utility, and target factors. Finally, a series of experiments are conducted to demonstrate the efficiency of the proposed TIRUP algorithm, and the experimental results indicate that TIRUP is suitable for processing large-scale and dense datasets. Cuiwei Peng, Jiahui Chen 0002, Wensheng Gan, Shicheng Wan |
IEEE Big Data | 4 |
| 2023 | Towards Contiguous Sequences in Uncertain DataabstractIn data mining, high-utility sequential pattern mining (HUSPM) focuses more on the specific values of items than on their frequency, making it more practical in real-life scenarios. HUSPM with the contiguous constraint can be used to solve some applications requiring the sequence elements to occur consecutively. Due to device, environment, privacy issues, and other factors, the data is often not accurate, and traditional algorithms for mining high utility continuous sequence patterns (HUCSPs) do not perform well in handling uncertain data. To address this challenge, this paper presents a new algorithm named uncertain utility-driven contiguous pattern mining (UUCPM), which can discover HUCSPs efficiently and correctly. The algorithm is designed to obtain results from sequence data with uncertain probabilities set on the item level. Two tighter upper bounds on utility and corresponding pruning strategies are also proposed, which can effectively process and reduce the number of candidate patterns generated during pattern mining, thereby improving the performance of the mining process. Through extensive experiments, the proposed UUCPM algorithm has been verified for accuracy and performance, demonstrating its advanced properties. Wensheng Gan, Gengsen Huang, Yanxin Zheng, Philip S. Yu |
DSAA | 2 |
| 2023 | Incremental Targeted Mining in SequencesabstractHigh utility sequential pattern mining (HUSPM) is a critical research topic in data analytics (e.g., smart-city technologies), which takes into consideration three pivotal factors of data: timestamp, internal quantization, and external utility. Recently, a query-enabled HUSPM approach has been proposed, which aims to discover patterns based on a query sequence. However, this approach only works on static data and does not solve the tasks well under dynamic data. When the data is updated, it needs to restart the mining process, which leads to a lot of duplicate calculations and resource consumption. In the paper, to address the mining task of increasing sequence data over time, we develop an Incremental Targeted HUSPM algorithm called ITUS. A tighter upper bound called tight extension sequence utility (TESU) is proposed to determine key candidates, which can avoid the generation of unpromising patterns. By using TESU, a target candidate pattern tree (TCP-tree) is utilized to record the sequence information, and several efficient strategies are implemented to incrementally update the tree. Finally, we extensively evaluate our proposed algorithm on both real-world and synthetic datasets. The experimental results clearly demonstrate that not only does the novel algorithm guarantee the accuracy of the results after multiple database updates, but it also achieves higher efficiency than the baseline approach. Kaixia Hu, Wensheng Gan, Gengsen Huang, Guoting Chen, Jerry Chun-Wei Lin |
DSAA | 2 |
| 2023 | Multi-Dimensional Graph Rule Learner
Jiayang Wu 0001, Zhenlian Qi, Wensheng Gan |
KSEM (1) | 3 |
| 2023 | Privacy-preserving federated mining of frequent itemsets
Wensheng Gan, Yongdong Wu, Philip S. Yu |
Inf. Sci. | 2 |
| 2023 | Mining high-utility sequences with positive and negative values
Fuyin Lai, Guoting Chen, Wensheng Gan |
Inf. Sci. | 4 |
| 2023 | US-Rule: Discovering Utility-driven Sequential RulesabstractUtility-driven mining is an important task in data science and has many applications in real life. High-utility sequential pattern mining (HUSPM) is one kind of utility-driven mining. It aims at discovering all sequential patterns with high utility. However, the existing algorithms of HUSPM can not provide a relatively accurate probability to deal with some scenarios for prediction or recommendation. High-utility sequential rule mining (HUSRM) is proposed to discover all sequential rules with high utility and high confidence. There is only one algorithm proposed for HUSRM, which is not efficient enough. In this article, we propose a faster algorithm called US-Rule, to efficiently mine high-utility sequential rules. It utilizes the rule estimated utility co-occurrence pruning strategy (REUCP) to avoid meaningless computations. Moreover, to improve its efficiency on dense and long sequence datasets, four tighter upper bounds (LEEU, REEU, LERSU, and RERSU) and corresponding pruning strategies (LEEUP, REEUP, LERSUP, and RERSUP) are designed. US-Rule also proposes the rule estimated utility recomputing pruning strategy (REURP) to deal with sparse datasets. Finally, a large number of experiments on different datasets compared to the state-of-the-art algorithm demonstrate that US-Rule can achieve better performance in terms of execution time, memory consumption, and scalability. Gengsen Huang, Wensheng Gan, Jian Weng 0001, Philip S. Yu |
ACM Trans. Knowl. Discov. Data | 2 |
| 2023 | Anomaly Rule Detection in Sequence DataabstractAnalyzing sequence data usually leads to the discovery of interesting patterns and then anomaly detection. In recent years, numerous frameworks and methods have been proposed to discover interesting patterns in sequence data as well as detect anomalous behavior. However, existing algorithms mainly focus on frequency-driven analytics, and they are challenging to be applied in real-world settings. In this work, we present a new anomaly detection framework called DUOS that enables Discovery of Utility-aware Outlier Sequential rules from a set of sequences. In this pattern-based anomaly detection algorithm, we incorporate both the anomalousness and utility of a group, and then introduce the concept of utility-aware outlier sequential rule (UOSR). We show that this is a more meaningful way for detecting anomalies. Besides, we propose some efficient pruning strategies w.r.t. upper bounds for mining UOSR, as well as the outlier detection. An extensive experimental study conducted on several real-world datasets shows that the proposed DUOS algorithm has a better effectiveness and efficiency. Finally, DUOS outperforms the baseline algorithm and has a suitable scalability. Wensheng Gan, Shicheng Wan, Jiahui Chen 0002, Chien-Ming Chen 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Federated Learning Attacks and Defenses: A SurveyabstractIn terms of artificial intelligence, there are several security and privacy deficiencies in the traditional centralized training methods of machine learning models by a server. To address this limitation, federated learning (FL) has been proposed and is known for breaking down "data silos" and protecting the privacy of users. However, FL has not yet gained popularity in the industry, mainly due to its security, privacy, and high cost of communication. For the purpose of advancing the research in this field, building a robust FL system, and realizing the wide application of FL, this paper sorts out the possible attacks and corresponding defenses of the current FL system systematically. Firstly, this paper briefly introduces the basic workflow of FL and related knowledge of attacks and defenses. It reviews a great deal of research about privacy theft and malicious attacks that have been studied in recent years. Most importantly, in view of the current three classification criteria, namely the three stages of machine learning, the three different roles in federated learning, and the CIA (Confidentiality, Integrity, and Availability) guidelines on privacy protection, we divide attack approaches into two categories according to the training stage and the prediction stage in machine learning. Furthermore, we also identify the CIA property violated for each attack method and potential attack role. Various defense mechanisms are then analyzed separately from the level of privacy and security. Finally, we summarize the possible challenges in the application of FL from the aspect of attacks and defenses and discuss the future development direction of FL systems. In this way, the designed FL system has the ability to resist different attacks and is more secure and stable. Yijie Gui, Wensheng Gan, Yongdong Wu |
IEEE Big Data | 4 |
| 2022 | Metaverse Security and Privacy: An OverviewabstractMetaverse is a living space and cyberspace that realizes the process of virtualizing and digitizing the real world. It integrates a plethora of existing technologies with the goal of being able to map the real world, even beyond the real world. Metaverse has a bright future and is expected to have many applications in various scenarios. The support of the Metaverse is based on numerous related technologies becoming mature. Hence, there is no doubt that the security risks of the development of the Metaverse may be more prominent and more complex. We present some Metaverse-related technologies and some potential security and privacy issues in the Metaverse. We present current solutions for Metaverse security and privacy derived from these technologies. In addition, we also raise some unresolved questions about the potential Metaverse. To summarize, this survey provides an in-depth review of the security and privacy issues raised by key technologies in Metaverse applications. We hope that this survey will provide insightful research directions and prospects for the Metaverse's development, particularly in terms of security and privacy protection in the Metaverse. Jiayang Wu 0001, Wensheng Gan, Zhenlian Qi |
IEEE Big Data | 3 |
| 2022 | Flexibly Mining Better PatternsabstractCorrelated high-utility pattern mining (CoUPM) considers the correlation between items in a pattern and offers a more reliable analysis for users. In real-world applications, the discovered patterns from CoUPM can present more interpretable information, but not all of them are useful. Generally, users pay attention to the number of items a pattern contains, which allows them to make reasonable decisions. In this paper, we solve the problem of mining those correlated high-utility patterns whose length is specified. A utility-list-based algorithm called Flexible Correlated Utility-based Pattern (FCoUP) is proposed. Furthermore, we propose some pruning strategies with the designed upper bounds for the two evaluation metrics: correlation and utility, reducing unwanted patterns generated and nodes visited during the mining process. Experiments show that FCoUP variants can produce more intelligent and flexible correlated high-utility patterns on a variety of datasets. Gengsen Huang, Wensheng Gan, Long Li 0005, Tianlong Gu, Jiahui Chen 0002 |
IEEE Big Data | 2 |
| 2022 | Metaverse in Education: Vision, Opportunities, and ChallengesabstractTraditional education has been updated with the development of information technology in human history. Within big data and cyber-physical systems, the Metaverse has generated strong interest in various applications (e.g., entertainment, business, and cultural travel) over the last decade. As a novel social work idea, the Metaverse consists of many kinds of technologies, e.g., big data, interaction, artificial intelligence, game design, Internet computing, Internet of Things, and blockchain. It is foreseeable that the usage of Metaverse will contribute to educational development. However, the architectures of the Metaverse in education are not yet mature enough. There are many questions we should address for the Metaverse in education. To this end, this paper aims to provide a systematic literature review of Metaverse in education. This paper is a comprehensive survey of the Metaverse in education, with a focus on current technologies, challenges, opportunities, and future directions. First, we present a brief overview of the Metaverse in education, as well as the motivation behind its integration. Then, we survey some important characteristics for the Metaverse in education, including the personal teaching environment and the personal learning environment. Next, we envisage what variations of this combination will bring to education in the future and discuss their strengths and weaknesses. We also review the state-of-the-art case studies (including technical companies and educational institutions) for Metaverse in education. Finally, we point out several challenges and issues in this promising area. Shicheng Wan, Wensheng Gan, Jiahui Chen 0002, Han-Chieh Chao |
IEEE Big Data | 3 |
| 2022 | Pattern Discovery with Utility OccupancyabstractTo mine potential and helpful patterns, the majority of studies on pattern discovery from databases have been conducted in the last few decades. They have several obvious drawbacks: 1) Each thing stands out on its own and varies in significance based on factors including utility, risk, interest, and weight. 2) In specific application settings, an object has a favorable or unfavorable effect (e.g., products are often cross-sold and have positive or negative unit profits, which affect benefits). 3) The user could not have all the necessary information because frequent-based patterns typically only include a small percentage of the relevant patterns (for example, occupancy). To address this issue, we apply economic utility theory to the database and data mining fields. We provide a one-phase approach called pnHUO for discovering High Utility Occupancy patterns with positive and negative utility values that beyond frequency and usefulness. According to user interests, frequency, and utility occupancy, there are various utility occupancy patterns with positive and negative utility values. To hold the necessary data, a new frequency-utility tree and an indexed data structure called a positive-and-negative utility-occupancy list are created during the mining process. A number of pruning strategies are further developed using the determined upper bound of utility occupancy to reduce the search space. To evaluate the usefulness and efficiency of the suggested algorithm, five real datasets were tested in experiments, and the results were positive. Jiayi Sun 0002, Wensheng Gan, Jerry Chun-Wei Lin, Han-Chieh Chao |
IEEE Big Data | 2 |
| 2022 | Targeted Mining of Rare High-Utility PatternsabstractPattern discovery has been widely studied and applied as a classical problem in data mining. As a subfield of itemset mining, identifying high-utility rare itemsets (HURI) can find abnormal but vital patterns in transaction databases. It plays a unique role in real-world scenarios such as anomaly detection and disease detection. However, with large-scale databases, the final results are often massive according to the user-specified threshold. In other words, the mining algorithm ignores the user’s subjective interests and lacks interaction during the mining process. A pattern discovery algorithm may output many useless or uninteresting patterns. To this end, in this paper, we define the problem of mining targeted HURIs and propose a list-based algorithm called TaRP for effectively solving this issue. In addition, based on preliminary research, we propose several effective pruning strategies for improving the algorithm’s performance. TaRP makes the results more interactive and specific by incorporating the user’s prior knowledge during mining. It also has a natural performance advantage with the help of effective strategies. We also evaluated the proposed algorithm on several real-life datasets. The extensive experimental results demonstrate that TaRP not only correctly solves the problem but also has advantages in runtime and memory consumption, especially on dense datasets. Peifeng Zhang, Jiahui Chen 0002, Shicheng Wan, Wensheng Gan |
IEEE Big Data | 4 |
| 2022 | Frequent Itemset Mining with Local Differential PrivacyabstractWith the development of the Internet, a large amount of transaction data (e.g., shopping records, web browsing history), which represents user data, has been generated. By collecting user transaction data and learning specific patterns and association rules from it, service providers can provide better services. However, because of the increasing privacy awareness and the formulation of laws on data protection, collecting data directly from users will raise privacy concerns. The concept of local differential privacy (LDP), which provides strict data privacy protection on the user side and allows effective statistical analysis on the server side, is able to protect user privacy and perform statistics on sensitive issues at the same time. This paper adopts padding-and-sampling-based frequent oracle (PSFO), combined with an interactive query-response method satisfying local differential privacy, to identify frequent itemsets in an efficient and accurate way. Therefore, this paper proposes FIML, an improved algorithm for finding frequent itemsets in the LDP setting of transaction data. The data collector generates frequent candidate sets based on the results of the previous stage and uses them for querying, and users randomize their responses in a reduced domain to achieve local differential privacy. Extensive experiments on real-world and synthetic datasets show that the FIML algorithm can find frequent itemsets more efficiently with the same privacy protection and computational cost. Wensheng Gan, Yijie Gui, Yongdong Wu, Philip S. Yu |
CIKM | 2 |
| 2022 | Fast Mining RFM Patterns for Behavioral AnalyticsabstractIn recent years, the problem of high-utility itemset mining (HUIM) has been extensively studied. However, HUIM algorithms only reveal profitable but generalized itemsets from transaction databases. In the market analysis domain, these mining results just reflect the sales trend of all customers and are not sufficient for making market strategies. In other words, it is hard to maintain specific customers for a long time due to the limitations of HUIM analysis of customer behaviors. In this paper, a novel data mining algorithm called RFM-Miner is proposed to discover RFM-patterns that are highly recent, frequent, and profitable in transaction databases. The novel algorithm relies on the array-bin structure to fast calculate adopted upper-bounds (i.e., transaction-weighted utilization, subtree and local utility) in linear time and space. In addition, RFM-Miner always searches for extension items of an itemset in a small projected database. And the merging technique is utilized to reduce the size of the search space. An extensive experimental study on four datasets (including real-life and synthetic) shows that RFM-Miner performs very well in terms of runtime and memory consumption. The novel algorithm also achieves better performance than the state-of-the-art benchmarks, especially on dense datasets. Shicheng Wan, Jieying Deng, Wensheng Gan, Jiahui Chen 0002, Philip S. Yu |
DSAA | 3 |
| 2022 | Constraint-based Sequential Rule MiningabstractSequential rule mining (SRM) is an alternative to sequential pattern mining (SPM) when dealing with sequence data. SRM has a wide range of applications in numerous data analysis scenarios. Existing SRM algorithms usually discover the entire set of rules in the databases, which makes it not only difficult to analyze results because the discovered set is too large, but also does not consider the user’s expectations and background knowledge. To tackle this problem, researchers have explored related algorithms with different constraints according to their requirements. In this paper, we propose a flexible constraint-based SRM algorithm called ConSRM for discovering only the sequential rules within user-specified time bounds in a sequence database. This algorithm uses an efficient rule-growth method and develops corresponding constraints and pruning strategies to reduce the search space and speed up calculation. Comprehensive experiments were carried out on four real datasets to evaluate the performance (both effectiveness and efficiency) of ConSRM. Zhaowen Yin, Wensheng Gan, Gengsen Huang, Yongdong Wu, Philippe Fournier-Viger |
DSAA | 2 |
| 2022 | Fuzzy-driven periodic frequent pattern mining
Yanlin Qi, Guoting Chen, Wensheng Gan, Philippe Fournier-Viger |
Inf. Sci. | 4 |
| 2022 | On-Shelf Utility Mining of Sequence DataabstractUtility mining has emerged as an important and interesting topic owing to its wide application and considerable popularity. However, conventional utility mining methods have a bias toward items that have longer on-shelf time as they have a greater chance to generate a high utility. To eliminate the bias, the problem of on-shelf utility mining (OSUM) is introduced. In this article, we focus on the task of OSUM of sequence data, where the sequential database is divided into several partitions according to time periods and items are associated with utilities and several on-shelf time periods. To address the problem, we propose two methods, OSUM of sequence data (OSUMS) and OSUMS + , to extract on-shelf high-utility sequential patterns. For further efficiency, we also design several strategies to reduce the search space and avoid redundant calculation with two upper bounds time prefix extension utility ( TPEU ) and time reduced sequence utility ( TRSU ). In addition, two novel data structures are developed for facilitating the calculation of upper bounds and utilities. Substantial experimental results on certain real and synthetic datasets show that the two methods outperform the state-of-the-art algorithm. In conclusion, OSUMS may consume a large amount of memory and is unsuitable for cases with limited memory, while OSUMS + has wider real-life applications owing to its high efficiency. Chunkai Zhang, Zilin Du, Yuting Yang 0005, Wensheng Gan, Philip S. Yu |
ACM Trans. Knowl. Discov. Data | 4 |
| 2021 | TKQ: Top-K Quantitative High Utility Itemset Mining
Mourad Nouioua, Philippe Fournier-Viger, Wensheng Gan, Youxi Wu, Jerry Chun-Wei Lin, Farid Nouioua |
ADMA | 3 |
| 2021 | Mining On-shelf High-utility Quantitative ItemsetsabstractA recently emerged branch of utility-based research, called high-utility quantitative itemset mining (HUQIM), has been widely applied in real-life, and it considers not only the utility factor but also the quantity with ranges of itemsets. However, most existing utility-mining algorithms assume that pat-terns always appear regardless of the period. For instance, some products may sell well at certain times of the year. Considering the rich information in the database, such as quantity and time, we propose an effective and efficient approach for discovering on-shelf high-utility quantitative itemsets (OHUQIs). To avoid scanning the database multiple times, we adopt a data structure to maintain some necessary information, and thus, OHUQI only accesses the database twice. Several pruning strategies are also designed to prune a large number of unpromising itemsets in advance to shrink the search space. Finally, the subsequent experimental results show that OHUQI performs well on several real-world datasets. Wensheng Gan, Jinbao Miao, Chien-Ming Chen 0001 |
IEEE BigData | 2 |
| 2021 | NSPIS: Mining Negative Sequential Patterns with Individual SupportabstractNegative sequential pattern (NSP) mining is crucial and sometimes carries more enlightening information than positive sequential pattern (PSP) mining in data mining. Owing to its computational complexity and exponential search space, the task of discovering NSPs is often much more difficult and challenging than that for PSPs. To date, a few NSP mining algorithms have been proposed. However, most algorithms only consider a single support, thus can not present good results in many special real-world applications. To solve this problem and achieve better efficiency on a long sequence database or a large-scale database, we propose a novel algorithm called Negative Sequential Patterns with Individual Support (NSPIS) in this paper. The projection mechanism is adopted to NSPIS, which allows greatly reduce the search space and simultaneously improve the efficiency. Finally, detailed results of the experiments show that NSPIS can achieve better performance and it uses less memory on large datasets compared to the state-of-the-art algorithm. Gengsen Huang, Wensheng Gan, Shan Huang 0009, Jiahui Chen 0002, Chien-Ming Chen 0001 |
IEEE BigData | 2 |
| 2021 | Joint Utility and Frequency for Pattern ClassificationabstractHigh-frequency itemset mining (HFIM) and high-utility itemset mining (HUIM) aim to discover itemsets with high occurrence and high utility, respectively, in a transaction database. A number of efficient algorithms have been developed to identify these high-utility itemsets (HUIs) or high-frequency itemsets (HFIs). Such algorithms play an increasingly important role in many occasions especially for analysis in commercial enterprises. In this paper, we propose a new model called joint utility and frequency for pattern classification, and two new algorithms, namely UFCgenand UFCfast. Both algorithms are designed to categorize each itemset into different type of patterns by setting the minimum thresholds of utility and frequency. We compare these algorithms on two datasets. The experimental results show that both algorithms can successfully collect three different types of itemsets from all candidate itemsets based on frequency and utility, and the list-based UFCfastalgorithm outperforms the level-wise-based UFCgenalgorithm in terms of execution time. Wensheng Gan, Yongdong Wu, Jiahui Chen 0002, Chien-Ming Chen 0001 |
IEEE BigData | 2 |
| 2021 | Targeted High-Utility Itemset QueryingabstractTraditional high-utility itemset mining (HUIM) aims to determine all high-utility itemsets (HUIs) that satisfy the minimum utility threshold in transaction databases. However, in most applications, not all HUIs are interesting because only specific parts are required. Thus, targeted mining based on user preferences is more important than traditional mining tasks. This paper is the first to propose a targeted HUIM problem and to provide a clear formulation of the targeted utility mining task in a quantitative transaction database. A tree-based algorithm known as Target-based high-Utility iteMset querying using (TargetUM) is proposed. The algorithm uses a lexicographic querying tree and three effective pruning strategies to improve the mining efficiency. We implemented experimental validation on several real and synthetic databases, and the results demonstrate that the performance of TargetUM is satisfactory, complete, and correct. Finally, owing to the lexicographic querying tree, the database no longer needs to be scanned repeatedly for multiple queries. Jinbao Miao, Shicheng Wan, Wensheng Gan, Jiayi Sun 0002, Jiahui Chen 0002 |
IEEE BigData | 3 |
| 2021 | A Generic Knowledge Based Medical Diagnosis Expert SystemabstractIn this paper, we design and implement a generic medical knowledge based system (MKBS) for identifying diseases from several symptoms. In this system, some important aspects like knowledge bases system, knowledge representation, inference engine have been addressed. The system asks users different questions and inference engines will use the certainty factor to prune out low possible solutions. The proposed disease diagnosis system also uses a graphical user interface (GUI) to facilitate users to interact with the expert system. Our expert system is generic and flexible, which can be integrated with any rule bases system in disease diagnosis. Xin Huang 0005, Xuejiao Tang, Wenbin Zhang 0002, Ji Zhang 0001, Wensheng Gan, Shichao Pei, Zhen Liu 0017, Yiyi Huang |
iiWAS | 6 |
| 2021 | Discovering high utility-occupancy patterns from uncertain data
Chien-Ming Chen 0001, Wensheng Gan, Lina Qiu, Weiping Ding 0001 |
Inf. Sci. | 3 |
| 2021 | TKUS: Mining top-k high utility sequential patterns
Chunkai Zhang, Zilin Du, Wensheng Gan, Philip S. Yu |
Inf. Sci. | 3 |
| 2021 | Utility Mining Across Multi-Dimensional SequencesabstractKnowledge extraction from database is the fundamental task in database and data mining community, which has been applied to a wide range of real-world applications and situations. Different from the support-based mining models, the utility-oriented mining framework integrates the utility theory to provide more informative and useful patterns. Time-dependent sequence data are commonly seen in real life. Sequence data have been widely utilized in many applications, such as analyzing sequential user behavior on the Web, influence maximization, route planning, and targeted marketing. Unfortunately, all the existing algorithms lose sight of the fact that the processed data not only contain rich features (e.g., occur quantity, risk, and profit), but also may be associated with multi-dimensional auxiliary information, e.g., transaction sequence can be associated with purchaser profile information. In this article, we first formulate the problem of utility mining across multi-dimensional sequences, and propose a novel framework named MDUS to extract Multi-Dimensional Utility-oriented Sequential useful patterns. To the best of our knowledge, this is the first study that incorporates the time-dependent sequence-order, quantitative information, utility factor, and auxiliary dimension. Two algorithms respectively named MDUS EM and MDUS SD are presented to address the formulated problem. The former algorithm is based on database transformation, and the later one performs pattern joins and a searching method to identify desired patterns across multi-dimensional sequences. Extensive experiments are carried on six real-life datasets and one synthetic dataset to show that the proposed algorithms can effectively and efficiently discover the useful knowledge from multi-dimensional sequential databases. Moreover, the MDUS framework can provide better insight, and it is more adaptable to real-life situations than the current existing models. Wensheng Gan, Jerry Chun-Wei Lin, Jiexiong Zhang, Hongzhi Yin, Philippe Fournier-Viger, Han-Chieh Chao, Philip S. Yu |
ACM Trans. Knowl. Discov. Data | 1 |
| 2021 | A Survey of Utility-Oriented Pattern MiningabstractThe main purpose of data mining and analytics is to find novel, potentially useful patterns that can be utilized in real-world applications to derive beneficial knowledge. For identifying and evaluating the usefulness of different kinds of patterns, many techniques and constraints have been proposed, such as support, confidence, sequence order, and utility parameters (e.g., weight, price, profit, quantity, satisfaction, etc.). In recent years, there has been an increasing demand for utility-oriented pattern mining (UPM, or called utility mining). UPM is a vital task, with numerous high-impact applications, including cross-marketing, e-commerce, finance, medical, and biomedical applications. This survey aims to provide a general, comprehensive, and structured overview of the state-of-the-art methods of UPM. First, we introduce an in-depth understanding of UPM, including concepts, examples, and comparisons with related concepts. A taxonomy of the most common and state-of-the-art approaches for mining different kinds of high-utility patterns is presented in detail, including Apriori-based, tree-based, projection-based, vertical-/horizontal-data-format-based, and other hybrid approaches. A comprehensive review of advanced topics of existing high-utility pattern mining techniques is offered, with a discussion of their pros and cons. Finally, we present several well-known open-source software packages for UPM. We conclude our survey with a discussion on open and practical challenges in this field. Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao, Vincent S. Tseng, Philip S. Yu |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2020 | OSUMI: On-Shelf Utility Mining from Itemset-based DataabstractAs an important technique for dealing with transactional database in the field of data mining, high-utility itemset mining (HUIM) can be used to discover itemsets which have a high utility. However, it has a bias when towarding the item combinations which have more exhibition period since they have more opportunity to generate a high utility. To address this, the on-shelf time period of items need to be considered, thus on-shelf utility mining (OSUM) can be applied in the application which is more closer to the actual situation. Currently several models have been proposed to deal with the OSUM problem, but they still suffer from the requirement that it needs to maintain a massive candidates in memory and to scan database many times. In this paper, we propose an effective algorithm named OSUMI (On-Shelf Utility Mining from Itemset-based data) which can discover the on-shelf itemsets with high utility in a more practical way. More precisely, in order to avoid the problems of high memory consumption, OSUMI applies some properties of on-shelf utility. Besides, two upper-bounds named subtree utility and local utility are applied to prune the search space. Finally, an extensive experimental study on two real on-shelf datasets shows that our proposed algorithm can be significantly faster than the state-of-the-art algorithm for this mining task. Jiahui Chen 0002, Xu Guo 0003, Wensheng Gan, Chien-Ming Chen 0001, Weiping Ding 0001, Guoting Chen |
IEEE BigData | 3 |
| 2020 | TopHUI: Top-k high-utility itemset mining with negative utilityabstractIn the field of data science, utility-driven data mining has become an emergent intelligent technique with wide applications. The existing utility mining algorithms usually discover all the patterns satisfying a given minimum utility threshold. However, a huge number of return results is not intuitive, not interpretable, and not easy for users to understand. Besides, it is often difficult and time-consuming for users to set a proper minimum utility threshold that is quite sensitive to the mining results. To address these issues, the problem of top-k high-utility itemset mining has been studied. In this paper, we present an efficient algorithm (named TopHUI) for finding top-k high-utility itemsets from transactional database that contains both positive and negative utility. This algorithm utilizes the positive-and-negative utility-list (PNU-list) to store the compress information, including positive, negative, and remaining utility. Besides, several threshold raising strategies and pruning strategies are proposed to prune the search space. Finally, some extensive experiments were conducted to evaluate the performance of the proposed TopHUI algorithm on both real-life and synthetic datasets, particularly in terms of effectiveness and efficiency. Wensheng Gan, Shicheng Wan, Jiahui Chen 0002, Chien-Ming Chen 0001, Lina Qiu |
IEEE BigData | 1 |
| 2020 | ProUM: Projection-based utility mining on sequence data
Wensheng Gan, Jerry Chun-Wei Lin, Jiexiong Zhang, Han-Chieh Chao, Hamido Fujita, Philip S. Yu |
Inf. Sci. | 1 |
| 2019 | Utility-Driven Mining of High Utility EpisodesabstractSequence data, e.g., complex event sequence, is more commonly seen than other types of data (e.g., transaction data) in real-world applications. For the mining task from sequence data, several problems have been formulated, such as sequential pattern mining, episode mining, and sequential rule mining. As one of the fundamental problems, episode mining has often been studied. The common wisdom is that discovering frequent episodes is not useful enough. In this paper, we propose an efficient utility mining approach namely UMEpi: Utility Mining of high-utility Episodes from complex event sequence. We propose the concept of remaining utility of episode, and achieve a tighter upper bound, namely episode-weighted utilization (EWU), which will provide better pruning. Thus, the optimized EWU-based pruning strategy can achieve better improvements in mining efficiency. Finally, experiments on two real-life datasets demonstrate that UMEpi can discover the complete high-utility episodes from complex event sequence, while state-of-the-art algorithms fail to return the correct results. Besides, the improved variants of UMEpi outperforms the baseline. Wensheng Gan, Jerry Chun-Wei Lin, Han-Chieh Chao, Philip S. Yu |
IEEE BigData | 1 |
| 2019 | Correlated utility-based pattern mining
Wensheng Gan, Jerry Chun-Wei Lin, Han-Chieh Chao, Hamido Fujita, Philip S. Yu |
Inf. Sci. | 1 |
| 2019 | A Survey of Parallel Sequential Pattern MiningabstractWith the growing popularity of shared resources, large volumes of complex data of different types are collected automatically. Traditional data mining algorithms generally have problems and challenges including huge memory cost, low processing speed, and inadequate hard disk space. As a fundamental task of data mining, sequential pattern mining (SPM) is used in a wide variety of real-life applications. However, it is more complex and challenging than other pattern mining tasks, i.e., frequent itemset mining and association rule mining, and also suffers from the above challenges when handling the large-scale data. To solve these problems, mining sequential patterns in a parallel or distributed computing environment has emerged as an important issue with many applications. In this article, an in-depth survey of the current status of parallel SPM (PSPM) is investigated and provided, including detailed categorization of traditional serial SPM approaches, and state-of-the art PSPM. We review the related work of PSPM in details including partition-based algorithms for PSPM, apriori-based PSPM, pattern-growth-based PSPM, and hybrid algorithms for PSPM, and provide deep description (i.e., characteristics, advantages, disadvantages, and summarization) of these parallel approaches of PSPM. Some advanced topics for PSPM, including parallel quantitative/weighted/utility SPM, PSPM from uncertain data and stream data, hardware acceleration for PSPM, are further reviewed in details. Besides, we review and provide some well-known open-source software of PSPM. Finally, we summarize some challenges and opportunities of PSPM in the big data era. Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao, Philip S. Yu |
ACM Trans. Knowl. Discov. Data | 1 |
| 2018 | CoUPM: Correlated Utility-based Pattern MiningabstractIn the field of data mining, many utility-oriented mining approaches have been extensively studied. Previous studies have, however, the limitation that they rarely consider the inherent correlation of items among the discovered patterns. For example, from the purchase behavior, a high-utility group of products (w.r.t. multi-products) may contain the items with both high or low utility. This pattern is also considered as a valuable pattern even if they may not be highly correlated, or even happened together by the chance. In this paper, we propose an efficient utility mining approach namely non-redundant Correlated high-Utility Pattern Miner (CoUPM) by considering both strong positive correlation and profitable value of the products. The derived patterns with high utility and strong correlation can lead to more insightful availability than those patterns only have high utility values. The utility-list structure is maintained and applied to store necessary information of correlation and utility. Several pruning strategies are further developed to improve the efficiency for discovering the desired patterns. Experimental results show that the non-redundant correlated high-utility patterns have more effectiveness than some other kinds of patterns. Moreover, the proposed CoUPM algorithm significantly outperforms the state-of-the-art algorithm. Wensheng Gan, Jerry Chun-Wei Lin, Han-Chieh Chao, Tzung-Pei Hong, Philip S. Yu |
IEEE BigData | 1 |
| 2018 | Privacy Preserving Utility Mining: A SurveyabstractIn big data era, the collected data usually contains rich information and hidden knowledge. Utility-oriented pattern mining and analytics have shown a powerful ability to explore these ubiquitous data, which may be collected from various fields and applications, such as market basket analysis, retail, click-stream analysis, medical analysis, and bioinformatics. However, analysis of these data with sensitive private information raises privacy concerns. To achieve better trade-off between utility maximizing and privacy preserving, Privacy-Preserving Utility Mining (PPUM) has become a critical issue in recent years. In this paper, we provide a comprehensive overview of PPUM. We first present the background of utility mining, privacy-preserving data mining and PPUM, then introduce the related preliminaries and problem formulation of PPUM, as well as some key evaluation criteria for PPUM. In particular, we present and discuss the current state-of-the-art PPUM algorithms, as well as their advantages and deficiencies in detail. Finally, we highlight and discuss some technical challenges and open directions for future research on PPUM. Wensheng Gan, Jerry Chun-Wei Lin, Han-Chieh Chao, Shyue-Liang Wang, Philip S. Yu |
IEEE BigData | 1 |
| 2018 | Exploiting highly qualified pattern with frequency and weight occupancy
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao, Justin Zhijun Zhan, Ji Zhang 0001 |
Knowl. Inf. Syst. | 1 |
| 2017 | Extracting Non-redundant Correlated Purchase Behaviors by Utility Measure
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao |
DaWaK | 1 |
| 2017 | Mining High-Utility Itemsets with Both Positive and Negative Unit Profits from Uncertain Databases
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao, Vincent S. Tseng |
PAKDD (1) | 1 |
| 2017 | FDHUP: Fast algorithm for mining discriminative high utility patterns
Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong, Han-Chieh Chao |
Knowl. Inf. Syst. | 2 |
| 2016 | Mining Discriminative High Utility Patterns
Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong |
ACIIDS (2) | 2 |
| 2016 | Mining Recent High Expected Weighted Itemsets from Uncertain Databases
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao |
APWeb (1) | 1 |
| 2016 | Mining Recent High-Utility Patterns from Temporal Databases with Time-Sensitive Constraint
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao |
DaWaK | 1 |
| 2016 | More Efficient Algorithms for Mining High-Utility Itemsets with Multiple Minimum Utility Thresholds
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao |
DEXA (1) | 1 |
| 2016 | More Efficient Algorithm for Mining Frequent Patterns with Multiple Minimum Supports
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao |
WAIM (1) | 1 |
| 2016 | Efficient Mining of Uncertain Data for High-Utility Itemsets
Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong, Vincent S. Tseng |
WAIM (1) | 2 |
| 2016 | Fast algorithms for mining high-utility itemsets with various discount strategies
Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong, Vincent S. Tseng |
Adv. Eng. Informatics | 2 |
| 2015 | Mining Weighted Frequent Itemsets with the Recency Constraint
Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong |
APWeb | 2 |
| 2015 | Mining high-utility itemsets with various discount strategiesabstractIn recent years, mining high-utility itemsets (HUIs) has become as a key topic in data mining. However, most of the developed algorithms assume the unrealistic situations that unit profits of items remain unchanged over time. But in real-life situations, the profit of an item or itemset varies as a function of cost prices, sales prices and sales strategies. In this paper, a novel framework for mining HUIs with two algorithms under various Discount strategies (HUID) are introduced. HUID-tp is based on various discount strategies and a novel downward closure property to mine the complete set of HUIs. HUID-Miner is an algorithm relying on a compact data structure (Positive-and-Negative Utility-list, PNU-list) and new pruning strategies to efficiently discover HUIs without candidate generation, while considerably reducing the size of the search space. Furthermore, a strategy named Estimated Utility Co-occurrence Strategy which stores the relationships between 2-itemsets is also adopted in the proposed improvement HUID-EMiner algorithm to speed up computation. An extensive experimental study carried on several real-life datasets shows the performance of the proposed algorithms. Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong, Vincent S. Tseng |
DSAA | 2 |
| 2015 | A fast updated algorithm to maintain the discovered high-utility itemsets for transaction modification
Jerry Chun-Wei Lin, Wensheng Gan, Tzung-Pei Hong |
Adv. Eng. Informatics | 2 |
| 2015 | Efficient algorithms for mining up-to-date high-utility patterns
Jerry Chun-Wei Lin, Wensheng Gan, Tzung-Pei Hong, Vincent S. Tseng |
Adv. Eng. Informatics | 2 |
| 2014 | Incrementally Updating High-Utility Itemsets with Transaction Insertion
Jerry Chun-Wei Lin, Wensheng Gan, Tzung-Pei Hong, Jeng-Shyang Pan 0001 |
ADMA | 2 |