Jerry Chun-Wei Lin

dblp:41/912 · also Chun-Wei Lin 0001, Chunwei Lin 0001 · DBLP profile ↗
← Back
113ranked-venue papers in the field
24as first author
44since 2021 · last 2025
ORCID · conflict

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 33 (7 first)Big Data, Cloud & Distributed Data Systems · 27 (2 first)Database Systems & Data Management · 23 (5 first)Knowledge Engineering, Semantic Web & Information Systems · 16 (2 first)Other / Interdisciplinary · 10 (6 first)Information Retrieval & Web Search · 4 (2 first)
YearPublicationVenuePosition
2025 KG-SEA: A Self-Evolving Framework for Iterative Knowledge Graph Construction in Graph-RAG Systems
Pi-Wei Chen, Myroslav Mishchuk, Alexandre Niyomugaba, Jerry Chun-Wei Lin, Rafal Cupek
IEEE Big Data4
2025 Dynamic sequential neighbor processing: A liquid neural network-inspired framework for enhanced graph neural networks
Kuijie Zhang, Jerry Chun-Wei Lin
Inf. Sci.6
2024 RECALL: Towards Generalized Representations in Unsupervised Federated Learning Under Non-IID Conditions
Pi-Wei Chen, Jerry Chun-Wei Lin, Feng-Hao Yeh, Rafal Cupek, Chao-Chun Chen
ACIIDS (1)2
2024 Masked Face-Landmark Prediction with Mask-Coefficient
Tzung-Pei Hong, Jerry Chun-Wei Lin, Ja-Hwung Su, Tang-Kai Yin
ACIIDS (1)2
2024 FedCali: Mitigating Overgeneralization for Anomaly Detection in Distributed Sensor Environments
abstract
In distributed manufacturing environments, Auto-mated Guided Vehicles (AGVs) equied with visual camera play a crucial role in automating material handling and optimizing production efficiency. Detecting anomalies during AGV operation is crucial to prevent potential malfunctions that could disrupt industrial processes. However, anomaly detection is challenging due to privacy concerns and the heterogeneity of data collected by AGVs across different factories. While sharing data across factories can improve the generalization capabilities of models, this can lead to overgeneralization in reconstruction-based anomaly detection, where the model reconstructs both normal and anomalous data too well, reducing its ability to detect anomalies. To address this problem, we propose FedCali, a federated learning framework that balances generalization and specialization across AGVs monitoring different manufacturing processes. Our proposed Gradient Guiding Mechanism (GGM) selectively aligns local model gradients with global knowledge only when necessary. This allows local models to retain their unique characteristics while benefiting from shared insights. Experiments with the MVTec dataset show that FedCali improves both reconstruction quality and anomaly detection accuracy, achieving higher AUROC scores and lower losses compared to baseline methods. This shows that FedCali is able to effectively process various manufacturing data collected by AGVs while maintaining data privacy.
Pi-Wei Chen, Jerry Chun-Wei Lin, Rafal Cupek, Chao-Chun Chen
IEEE Big Data2
2024 SSRepL-ADHD: Adaptive Complex Representation Learning Framework for ADHD Detection from Visual Attention Tasks
abstract
Self Supervised Representation Learning (SSRepL) can capture meaningful and robust representations of the Attention Deficit Hyperactivity Disorder (ADHD) data and have the potential to improve the model’s performance on also downstream different types of Neurodevelopmental disorder (NDD) detection. In this paper, a novel SSRepL and Transfer Learning (TL)-based framework that incorporates a Long Short-Term Memory (LSTM) and a Gated Recurrent Units (GRU) model is proposed to detect children with potential symptoms of ADHD. This model uses Electroencephalogram (EEG) signals extracted during visual attention tasks to accurately detect ADHD by preprocessing EEG signal quality through normalization, filtering, and data balancing. For the experimental analysis, we use three different models: 1) SSRepL and TL-based LSTM-GRU model named as SSRepL-ADHD, which integrates LSTM and GRU layers to capture temporal dependencies in the data, 2) lightweight SSRepL-based DNN model (LSSRepL-DNN), and 3) Random Forest (RF). In the study, these models are thoroughly evaluated using well-known performance metrics (i.e., accuracy, precision, recall, and F1-score). The results show that the proposed SSRepL-ADHD model achieves the maximum accuracy of 81.11% while admitting the difficulties associated with dataset imbalance and feature selection.
Abdul Rehman 0006, Ilona Heldal, Jerry Chun-Wei Lin
IEEE Big Data3
2024 Towards a Supporting Framework for Neuro-Developmental Disorder: Considering Artificial Intelligence, Serious Games and Eye Tracking
abstract
This paper focuses on developing a framework for uncovering insights about NDD children’s performance (e.g., raw gaze cluster analysis, duration analysis & area of interest for sustained attention, stimuli expectancy, loss of focus/motivation, inhibitory control) and informing their teachers. The hypothesis behind this work is that self-adaptation of games can contribute to improving students’ well-being and performance by suggesting personalized activities (e.g., highlighting stimuli to increase attention or choosing a difficulty level that matches students’ abilities). The aim is to examine how AI can be used to help solve this problem. The results would not only contribute to a better understanding of the problems of NDD children and their teachers but also help psychologists to validate the results against their clinical knowledge, improve communication with patients and identify areas for further investigation, e.g., by explaining the decision made and preserving the children’s private data in the learning process.
Abdul Rehman 0006, Ilona Heldal, Diana L. Stilwell, Jerry Chun-Wei Lin
IEEE Big Data4
2024 Enhancing Customer Behavior Prediction and Interpretability
abstract
This paper leverages insights from my previous works to analyze and predict customer behavior in different areas using data mining and machine learning techniques. The research focuses on identifying and interpreting customer preferences, purchasing patterns and risk factors for customer churn to develop predictive models that support decision making. By analyzing various datasets such as tabular data, clickstream data, and digital interactions, this study aims to provide a comprehensive framework for extracting actionable insights and enhancing the accuracy and transparency of customer behavior predictions. The proposed approach enables businesses to anticipate customer behavior, personalize marketing strategies, and improve the explainability and effectiveness of customer behavior prediction models. Future work will incorporate advanced visualization techniques and various data structures to further improve the model’s transparency and predictive capability.
Danny Yu-Chung Wang, Lars Arne Jordanger, Jerry Chun-Wei Lin
IEEE Big Data3
2024 A Utility-Mining-Driven Active Learning Approach for Analyzing Clickstream Sequences
abstract
In the rapidly evolving e-commerce industry, the ability to select high-quality data for model training is essential. This study introduces the High-Utility Sequential Pattern Mining using SHAP values (HUSPM-SHAP) model, a utility mining based active learning strategy, to tackle this challenge. We found that the parameter settings for positive and negative SHAP values affect the mining results of the model, introducing a key consideration into the active learning framework. Unlike traditional SHAP, which evaluates individual elements, HUSPM-SHAP utilizes SHAP values in combination with HUSPM to identify the valuable of high-utility sequential patterns for improving prediction models. In experiments to predict behaviors that actually lead to purchases, the developed HUSPM-SHAP model shows its superiority in different scenarios. The model’s ability to reduce labeling requirements while maintaining high predictive performance is highlighted. Our results show that the model is able to refine the processing of e-commerce data and lead to optimized, cost-efficient prediction modeling.
Danny Yu-Chung Wang, Lars Arne Jordanger, Jerry Chun-Wei Lin
IEEE Big Data3
2024 GPU-Based Efficient Parallel Heuristic Algorithm for High-Utility Itemset Mining in Large Transaction Datasets (Extended Abstract)
abstract
Heuristic algorithms have been developed to find approximate solutions for high-utility itemset mining (HUIM) problems that compensate for the performance bottlenecks of exact algorithms. However, heuristic algorithms still face the problem of long runtime and insufficient mining quality, especially for large transaction datasets with thousands to tens of thousands of items and up to millions of transactions. To solve these problems, a novel GPU-based efficient parallel heuristic algorithm for HUIM (PHA-HUIM) is proposed in this paper. The iterative process of PHA-HUIM consists of three main steps: the search strategy, fitness evaluation, and ring topology communication. The search strategy and ring topology communication are designed to run in constant time on GPU. The parallelism of fitness evolution helps to substantially accelerate the algorithm. To improve the mining quality, a multi-start strategy with an unbalanced allocation strategy is employed in the search process. Ring topology communication is adopted to maintain population diversity. A load balancing strategy is introduced to reduce the thread divergence to improve the parallel efficiency. The experimental results on nine large datasets show that PHA-HUIM outperforms state-of-the-art HUIM algorithms in terms of speedup performance, runtime, and mining quality.
Wei Fang 0001, Haipeng Jiang, Hengyang Lu, Jun Sun 0008, Xiaojun Wu 0001, Jerry Chun-Wei Lin
ICDE6
2024 TripleS: A Subsidy-Supported Storage for Electricity with Self-financing Management System
Jia-Hao Syu, Rafal Cupek, Chao-Chun Chen, Jerry Chun-Wei Lin
PAKDD (5)4
2024 Efficient approach of high average utility pattern mining with indexed list-based structure in dynamic environments
Hyeonmo Kim, Hanju Kim, Myungha Cho, Bay Vo, Jerry Chun-Wei Lin, Hamido Fujita, Unil Yun
Inf. Sci.5
2024 An efficient approach for incremental erasable utility pattern mining from non-binary data
Yoonji Baek, Hanju Kim, Myungha Cho, Hyeonmo Kim, Chanhee Lee 0005, Taewoong Ryu, Heonho Kim, Bay Vo, Vincent W. Gan, Philippe Fournier-Viger, Jerry Chun-Wei Lin, Witold Pedrycz, Unil Yun
Knowl. Inf. Syst.11
2024 GPU-Based Efficient Parallel Heuristic Algorithm for High-Utility Itemset Mining in Large Transaction Datasets
abstract
Heuristic algorithms have been developed to find approximate solutions for high-utility itemset mining (HUIM) problems that compensate for the performance bottlenecks of exact algorithms. However, heuristic algorithms still face the problem of long runtime and insufficient mining quality, especially for large transaction datasets with thousands to tens of thousands of items and up to millions of transactions. To solve these problems, a novel GPU-based efficient parallel heuristic algorithm for HUIM (PHA-HUIM) is proposed in this paper. The iterative process of PHA-HUIM consists of three main steps: the search strategy, fitness evaluation, and ring topology communication. The search strategy and ring topology communication are designed to run in constant time on GPU. The parallelism of fitness evolution helps to substantially accelerate the algorithm. A new data structure with a sort-mapping strategy is proposed to enhance the search ability and reduce memory usage. To improve the mining quality, a multi-start strategy with an unbalanced allocation strategy is employed in the search process. Ring topology communication is adopted to maintain population diversity. A load balancing strategy is introduced to reduce the thread divergence to improve the parallel efficiency. The experimental results on nine large datasets show that PHA-HUIM outperforms state-of-the-art HUIM algorithms in terms of speedup performance, runtime, and mining quality.
Wei Fang 0001, Haipeng Jiang, Hengyang Lu, Jun Sun 0008, Xiaojun Wu 0001, Jerry Chun-Wei Lin
IEEE Trans. Knowl. Data Eng.6
2023 Effective Prediction of Energy Consumption in Automated Guided Vehicles with Recurrent and Convolutional Neural Networks
abstract
Detection and prediction of failures in Automated Guided Vehicles (AGV) are essential for the uninterrupted operation of production plants. Anomaly detection is usually achieved by comparing expected measurement values with actual observations. Thus, it is crucial to predict telemetry signals properly. In this paper, we research the prediction of energy consumption using state-of-the-art Artificial Neural Networks architectures (SCINet) compared with other Recurrent Neural Network (RNN) approaches on the data streams acquired from CoBotAGV. We especially focus on the possibility of applying feature weighting. We show that it can improve prediction capabilities. We also investigate resource utilization in terms of time to fit the embedded AGV environment.
Pawel Benecki, Daniel Kostrzewa, Piotr Grzesik, Bohdan Shubyn, Jia-Hao Syu, Jerry Chun-Wei Lin, Vaidy S. Sunderam, Dariusz Mrozek
IEEE Big Data6
2023 The Calibration of Single Beam Distance Sensors based on Machine Learning Methods
abstract
Smart cities require the use of many different types of sensors to make the communication, and distance sensors are one of the most commonly used elements in transportation systems and related infrastructures. The introduction of increasingly advanced autonomous systems in many areas of smart cities requires high measurement precision of the sensors used. High precision is essential for proper operation, long-term use, and safety in machine-to-machine or machine-to-human interactions. This paper presents a comparison of the accuracy of distance measurements for two commercially available single-beam LiDARs and two ultrasonic sensors. The aim of the research was to develop a calibration method in order to improve the accuracy of distance sensors. Based on the collected distance measurements, the sensors were calibrated using selected machine learning algorithms. The results of the experiments show the effectiveness of the proposed calibration methods, which yield an average mean absolute error (MAE) of 1.76 E-05 meters (m) and a root mean square error (RMSE) of 1.33 E-04 m for the tested sensors.
Piotr Biernacki, Adam Ziebinski, Jia-Hao Syu, Jerry Chun-Wei Lin
IEEE Big Data4
2023 Large Language Models in Education: Vision and Opportunities
abstract
With the rapid development of artificial intelligence technology, large language models (LLMs) have become a hot research topic. Education plays an important role in human social development and progress. Traditional education faces challenges such as individual student differences, insufficient allocation of teaching resources, and assessment of teaching effectiveness. Therefore, the applications of LLMs in the field of digital/smart education have broad prospects. The research on educational large models (EduLLMs) is constantly evolving, providing new methods and approaches to achieve personalized learning, intelligent tutoring, and educational assessment goals, thereby improving the quality of education and the learning experience. This article aims to investigate and summarize the application of LLMs in smart education. It first introduces the research background and motivation of LLMs and explains the essence of LLMs. It then discusses the relationship between digital education and EduLLMs and summarizes the current research status of educational large models. The main contributions are the systematic summary and vision of the research background, motivation, and application of large models for education (LLM4Edu). By reviewing existing research, this article provides guidance and insights for educators, researchers, and policy-makers to gain a deep understanding of the potential and challenges of LLM4Edu. It further provides guidance for further advancing the development and application of LLM4Edu, while still facing technical, ethical, and practical challenges requiring further research and exploration.
Wensheng Gan, Zhenlian Qi, Jiayang Wu 0001, Jerry Chun-Wei Lin
IEEE Big Data4
2023 Explainability of Leverage Points Exploration for Customer Churn Prediction
abstract
Customer churn is a most crucial challenge faced by various industries, especially for the telecommunications sector since it causes over 5 - 6 times cost to keep customers. Thus, customer churn prediction has experienced substantial expansion. While a lot of prediction models are becoming increasingly accurate, the capability to interpret these models and perform causal analysis has become a new challenge. Identifying the key reasons or patterns that truly influence customer churn is vital for enabling industries to make significant improvements. This study then introduced a comprehensive framework. That integrates causal graphs and explainable models such as SHAP and Shapley flow, and concepts from system dynamics to explore customer churn patterns and leverage points. Experimental results indicates that the designed model is more explainable and interpretable to explore the customer behaviors for customer churn prediction.
Danny Yu-Chung Wang, Lars Arne Jordanger, Jerry Chun-Wei Lin
IEEE Big Data3
2023 Incremental Targeted Mining in Sequences
abstract
High utility sequential pattern mining (HUSPM) is a critical research topic in data analytics (e.g., smart-city technologies), which takes into consideration three pivotal factors of data: timestamp, internal quantization, and external utility. Recently, a query-enabled HUSPM approach has been proposed, which aims to discover patterns based on a query sequence. However, this approach only works on static data and does not solve the tasks well under dynamic data. When the data is updated, it needs to restart the mining process, which leads to a lot of duplicate calculations and resource consumption. In the paper, to address the mining task of increasing sequence data over time, we develop an Incremental Targeted HUSPM algorithm called ITUS. A tighter upper bound called tight extension sequence utility (TESU) is proposed to determine key candidates, which can avoid the generation of unpromising patterns. By using TESU, a target candidate pattern tree (TCP-tree) is utilized to record the sequence information, and several efficient strategies are implemented to incrementally update the tree. Finally, we extensively evaluate our proposed algorithm on both real-world and synthetic datasets. The experimental results clearly demonstrate that not only does the novel algorithm guarantee the accuracy of the results after multiple database updates, but it also achieves higher efficiency than the baseline approach.
Kaixia Hu, Wensheng Gan, Gengsen Huang, Guoting Chen, Jerry Chun-Wei Lin
DSAA5
2023 Anomaly Detection Networks and Fuzzy Control Modules for Energy Grid Management with Q-Learning-Based Decision Making
abstract
Renewable energy generation has attracted the interest of researchers, but it is volatile, and management systems are vulnerable to malicious attacks. Therefore, security issues are of paramount importance for energy management systems. In this paper, we propose a secure Q-learning- based energy network management system (SQEMS), which consists of an anomaly detection module, a fuzzy control module to mitigate attacks, and a decision-making module to manage the energy grid. Experimental results show that the proposed anomaly detection module has excellent performance on malicious suppliers attacks (MS), and the fuzzy control module can further mitigate the negative effects of false predictions. The robustness analysis shows the effectiveness, robustness, and transferability in anomaly detection and energy management.
Jia-Hao Syu, Jerry Chun-Wei Lin, Philip S. Yu
SDM2
2023 AF-GCN: Completing various graph tasks efficiently via adaptive quadratic frequency response function in graph spectral domain
Kuijie Zhang, Gan Wang, Jerry Chun-Wei Lin, Fuyu Wang 0003, Yuanyuan Zhang 0008
Inf. Sci.4
2023 Graph Attention Network for Text Classification and Detection of Mental Disorder
abstract
A serious issue in today’s society is Depression, which can have a devastating impact on a person’s ability to cope in daily life. Numerous studies have examined the use of data generated directly from users using social media to diagnose and detect Depression as a mental illness. Therefore, this paper investigates the language used in individuals’ personal expressions to identify depressive symptoms via social media. Graph Attention Networks (GATs) are used in this study as a solution to the problems associated with text classification of depression. These GATs can be constructed using masked self-attention layers. Rather than requiring expensive matrix operations such as similarity or knowledge of network architecture, this study implicitly assigns weights to each node in a neighbourhood. This is possible because nodes and words can carry properties and sentiments of their neighbours. Another aspect of the study that contributed to the expansion of the emotion lexicon was the use of hypernyms. As a result, our method performs better when applied to data from the Reddit subreddit Depression. Our experiments show that the emotion lexicon constructed by using the Graph Attention Network ROC achieves 0.91 while remaining simple and interpretable.
Usman Ahmed, Jerry Chun-Wei Lin, Gautam Srivastava 0001
ACM Trans. Web2
2022 Automated Guided Vehicles Challenges for Artificial Intelligence
abstract
The use of Artificial Intelligence (AI) to support the Automated Guided Vehicles (AGV) that are used by industry poses a number of challenges that are specific to smart internal logistics systems that are necessary for agile manufacturing. On the one hand, it might seem that experience with the autonomous navigation system that are used in autonomous vehicles can be easily transferred to AGV. However, in this paper, the authors highlight specific problems that are associated with the navigation system of AGV, which has to reflect its operation in an industrial environment with high level of interaction with other production systems and human staff. On the other hand, it may seem that the wealth of experience from using AI in smart manufacturing can be easily transferred to the use of AGV. However, the authors show that although AGV are production tools, the challenges that are associated with the use of AI can significantly differ from other smart manufacturing areas. The number of challenges that are specific to use of AI for AGV is also discussed. This paper systematizes these challenges and discusses the most promising AI methods that can be used for the internal logistics systems that are based on AGV.
Rafal Cupek, Jerry Chun-Wei Lin, Jia-Hao Syu
IEEE Big Data2
2022 Pattern Discovery with Utility Occupancy
abstract
To mine potential and helpful patterns, the majority of studies on pattern discovery from databases have been conducted in the last few decades. They have several obvious drawbacks: 1) Each thing stands out on its own and varies in significance based on factors including utility, risk, interest, and weight. 2) In specific application settings, an object has a favorable or unfavorable effect (e.g., products are often cross-sold and have positive or negative unit profits, which affect benefits). 3) The user could not have all the necessary information because frequent-based patterns typically only include a small percentage of the relevant patterns (for example, occupancy). To address this issue, we apply economic utility theory to the database and data mining fields. We provide a one-phase approach called pnHUO for discovering High Utility Occupancy patterns with positive and negative utility values that beyond frequency and usefulness. According to user interests, frequency, and utility occupancy, there are various utility occupancy patterns with positive and negative utility values. To hold the necessary data, a new frequency-utility tree and an indexed data structure called a positive-and-negative utility-occupancy list are created during the mining process. A number of pruning strategies are further developed using the determined upper bound of utility occupancy to reduce the search space. To evaluate the usefulness and efficiency of the suggested algorithm, five real datasets were tested in experiments, and the results were positive.
Jiayi Sun 0002, Wensheng Gan, Jerry Chun-Wei Lin, Han-Chieh Chao
IEEE Big Data3
2022 An Efficient and Secured Energy Management System for Automated Guided Vehicles
abstract
In this paper, we propose a Secure Energy Management System (SEMS) with anomaly detection and Q-Learning decision modules for Automated Guided Vehicles (AGV). The anomaly detection module is a multi-task learning network to simultaneously classify suppliers and predict the real supply quantities. The Q-learning decision module can then determine operating reserve and subsidies to manage the energy grid. Experimental results illustrate that the proposed anomaly detection module has an excellent performance in classifying malicious suppliers, excels at shaping supply distribution, and outperforms the existing benchmark systems.
Jia-Hao Syu, Jerry Chun-Wei Lin, Dariusz Mrozek
IEEE Big Data2
2022 Double-Environmental Q-Learning for Energy Management System in Smart Grid
abstract
In this research, we present a Q-learning based energy management system (DEQEMS) that is able to make decisions by using unique states and intuitive actions while maintaining a high degree of interpretability. The results of the experiments show that the DEQEMS reduces the number of days required for convergence to 633, with a mean absolute error (MAE) of supply distribution of 6.7%. This is a 63% and 71% reduction, respectively, compared to the conventional system, and a 34% and 21% reduction, respectively, compared to a state-of-the-art system. The experimental results demonstrate not only the usefulness and feasibility of the DEQEMS, but also its resilience with outstanding and consistent performance under a wide range of conditions.
Jia-Hao Syu, Jerry Chun-Wei Lin, Philip S. Yu
IEEE Big Data2
2022 Mental Health Treatments Using an Explainable Adaptive Clustering Model
Usman Ahmed, Jerry Chun-Wei Lin, Gautam Srivastava 0001
PAKDD (3)2
2022 Graph embedding-based intelligent industrial decision for complex sewage treatment processes
abstract
Intelligent algorithms-driven industrial decision systems have been a general demand for modeling complex sewage treatment processes (STP). Existing researches modeled complex STP with the use of various neural network models, yet neglecting the fact that latent and occasional relations exist inside complex STP. To deal with the challenge, this paper proposes graph embedding-based intelligent industrial decision for complex STP (GE-STP). The graph embedding (GE) scheme is employed to enhance feature extraction and neural computing structure is utilized to simulate uncertain biochemical transformation inside STP. The introduction of GE can not only improves the fineness of feature spaces, but also improves the representative ability of models towards complex industrial processes. On this basis, the GE-STP is evaluated on a real-world data set collected from a realistic sewage treatment plant equipped with a set of Internet of Things devices. And some typical neural network models that have been utilized for modeling complex STP, are selected as baseline methods. Three groups of experiments show that efficiency of the GE-STP exceeds baselines about 6%–12%, and that the GE-STP is not susceptible to parameter changing.
Zhiwei Guo 0004, Yu Shen 0004, Ali Kashif Bashir, Keping Yu, Jerry Chun-Wei Lin
Int. J. Intell. Syst.5
2022 Occupancy-based utility pattern mining in dynamic environments of intelligent systems
abstract
Utility pattern mining is a branch of data mining that extracts valid patterns by considering the quantity and weight of the items. In addition, utility occupancy pattern mining, which considers the quantity, importance, and proportion of the pattern in the transaction, has been proposed. Despite this advantage, there is no utility seizing approach to handle the dynamically generated data flows. As electronics are interconnected and intelligent systems are constructed, data is generated in real-time and accumulated rapidly. Therefore, a method to read data immediately in a dynamic environment and efficiently analyze massive data is required. To overcome the limitations of the existing utility occupancy methods, we propose a novel mining approach, HUOMI, which performs quickly on an increasing database. The suggested algorithm has an optimized data structure and an improved pruning technique, which can respond to the dynamic environment promptly. To indicate the effectiveness of the proposed method, performance evaluations were conducted on real and synthetic data sets. In the experimental results, the suggested algorithm showed a better performance than the other state-of-the-art algorithms.
Taewoong Ryu, Unil Yun, Chanhee Lee 0005, Jerry Chun-Wei Lin, Witold Pedrycz
Int. J. Intell. Syst.4
2022 Optimized scheduling of resource-constraints in projects for smart construction
Jerry Chun-Wei Lin, Dehu Yu, Gautam Srivastava 0001, Chun-Hao Chen
Inf. Process. Manag.1
2022 Deep learning based hashtag recommendation system for multimedia data
abstract
This work aims to provide a novel hybrid architecture to suggest appropriate hashtags to a collection of orpheline tweets. The methodology starts with defining the collection of batches used in the convolutional neural network. This methodology is based on frequent pattern extraction methods. The hashtags of the tweets are then learned using the convolution neural network that was applied to the collection of batches of tweets. In addition, a pruning approach should ensure that the learning process proceeds properly by reducing the number of common patterns. Besides, the evolutionary algorithm is involved to extract the optimal parameters of the deep learning model used in the learning process. This is achieved by using a genetic algorithm that learns the hyper-parameters of the deep architecture. The effectiveness of our methodology has been demonstrated in a series of detailed experiments on a set of Twitter archives. From the results of the experiments, it is clear that the proposed method is superior to the baseline methods in terms of efficiency.
Youcef Djenouri, Asma Belhadi, Gautam Srivastava 0001, Jerry Chun-Wei Lin
Inf. Sci.4
2022 An efficient approach for mining maximized erasable utility patterns
Chanhee Lee 0005, Yoonji Baek, Taewoong Ryu, Hyeonmo Kim, Heonho Kim, Jerry Chun-Wei Lin, Bay Vo, Unil Yun
Inf. Sci.6
2022 Scalable Mining of High-Utility Sequential Patterns With Three-Tier MapReduce Model
abstract
High-utility sequential pattern mining (HUSPM) is a hot research topic in recent decades since it combines both sequential and utility properties to reveal more information and knowledge rather than the traditional frequent itemset mining or sequential pattern mining. Several works of HUSPM have been presented but most of them are based on main memory to speed up mining performance. However, this assumption is not realistic and not suitable in large-scale environments since in real industry, the size of the collected data is very huge and it is impossible to fit the data into the main memory of a single machine. In this article, we first develop a parallel and distributed three-stage MapReduce model for mining high-utility sequential patterns based on large-scale databases. Two properties are then developed to hold the correctness and completeness of the discovered patterns in the developed framework. In addition, two data structures called sidset and utility-linked list are utilized in the developed framework to accelerate the computation for mining the required patterns. From the results, we can observe that the designed model has good performance in large-scale datasets in terms of runtime, memory, efficiency of the number of distributed nodes, and scalability compared to the serial HUSP-Span approach.
Jerry Chun-Wei Lin, Youcef Djenouri, Gautam Srivastava 0001, Yuanfa Li, Philip S. Yu
ACM Trans. Knowl. Discov. Data1
2021 Mining Partially-Ordered Episode Rules in an Event Sequence
Philippe Fournier-Viger, Yangming Chen, Farid Nouioua, Jerry Chun-Wei Lin
ACIIDS4
2021 Investigating Crossover Operators in Genetic Algorithms for High-Utility Itemset Mining
M. Saqib Nawaz, Philippe Fournier-Viger, Wei Song 0004, Jerry Chun-Wei Lin, Bernd Noack
ACIIDS4
2021 TKQ: Top-K Quantitative High Utility Itemset Mining
Mourad Nouioua, Philippe Fournier-Viger, Wensheng Gan, Youxi Wu, Jerry Chun-Wei Lin, Farid Nouioua
ADMA5
2021 Detection of Trajectory Outliers in Intelligent Transportation Systems
abstract
In this paper, we provide a technique for identifying outliers based on embedding trajectory deviation points and deep clustering. We begin by constructing the network topology and the neighbors of the nodes to create a structural embedding while capturing the interactions of the nodes. We then develop a strategy to determine the hidden representation of distraction points in the road network topology. To create a collection of sequences from a hierarchical multilayer network, a biased random walk is used. This sequence is used to fine tune the embedding of the nodes. The trip embedding was then determined by averaging the node embedding values. Finally, the embeddings are clustered using an LSTM-based pairwise classification strategy based on similarity metrics. The experimental results show that compared to the generic techniques Node2Vec and Struct2Vec, the proposed embedding learning trajectory captures the structural identity and improves the F-measure by 5.06% and 2.4%, respectively.
Usman Ahmed, Jerry Chun-Wei Lin, Gautam Srivastava 0001, Youcef Djenouri, Jimmy Ming-Tai Wu
IEEE BigData2
2021 Learning Probabilistic Latent Structure for Outlier Detection from Multi-view Data
Zhen Wang 0037, Ji Zhang 0001, Yizheng Chen 0003, Chenhao Lu, Jerry Chun-Wei Lin, Jing Xiao 0005, R. Uday Kiran
PAKDD (1)5
2021 Average utility driven data analytics on damped windows for intelligent systems with data streams
abstract
In industrial areas, most of databases are dynamic databases, and the volume of the databases has grown with the passage of time. Especially, pattern mining for incremental database needs different approaches from static database because the profit or the accuracy of the previously inserted data can be reduced. Since data is time- sensitive, the recent data has a relatively higher value than the old data. In this paper, we suggest the damped window based average utility driven data analytics for intelligent systems, which the damped window reflects the importance according to the arrival time of the transactions. The proposed mining approach adopts novel data structure, which modify the importance of item as the passage of time, and it improves mining efficiency with several pruning strategies and without generating candidate patterns. To evaluate the performance of the proposed mining approach, we conducted various experiments using several real and synthetic data sets. The result of the experiments presented that the suggested method performs better in terms of runtime and memory usage than the other state-of-the-art mining techniques. Moreover, through the scalability experiments, which changed the number of different items or transactions, we verified that the proposed algorithm maintained a stable performance under various environmental changes.
Jongseong Kim, Unil Yun, Taewoong Ryu, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Witold Pedrycz
Int. J. Intell. Syst.5
2021 Fuzzy high-utility pattern mining in parallel and distributed Hadoop framework
abstract
Over the past decade, high-utility itemset mining (HUIM) has received widespread attention that can emphasize more critical information than was previously possible using frequent itemset mining (FIM). Unfortunately, HUIM is very similar to FIM since the methodology determines itemsets using a binary model based on a pre-defined minimum utility threshold. Additionally, most previous works only focused on single, small datasets in HUIM, which is not realistic to any real-world scenarios today containing big data environments. In this work, the fuzzy-set theory and a MapReduce framework are both utilized to design a novel high fuzzy utility pattern mining algorithm to resolve the above issues. Fuzzy-set theory is first involved and a new algorithm called efficient high fuzzy utility itemset mining (EFUPM) is designed to discover high fuzzy utility patterns from a single machine. Two upper-bounds are then estimated to allow early pruning of unpromising candidates in the search space. To handle the large-scale of big datasets, a Hadoop-based high fuzzy utility pattern mining (HFUPM) algorithm is then developed to discover high fuzzy utility patterns based on the Hadoop framework. Experimental results clearly show that the proposed algorithms perform strongly to mine the required high fuzzy utility patterns whether in a single machine or a large-scale environment compared to the current state-of-the-art approaches.
Jimmy Ming-Tai Wu, Gautam Srivastava 0001, Unil Yun, Jerry Chun-Wei Lin
Inf. Sci.5
2021 RHUPS: Mining Recent High Utility Patterns with Sliding Window-based Arrival Time Control over Data Streams
abstract
Databases that deal with the real world have various characteristics. New data is continuously inserted over time without limiting the length of the database, and a variety of information about the items constituting the database is contained. Recently generated data has a greater influence than the previously generated data. These are called the time-sensitive non-binary stream databases, and they include databases such as web-server click data, market sales data, data from sensor networks, and network traffic measurement. Many high utility pattern mining and stream pattern mining methods have been proposed so far. However, they have a limitation that they are not suitable to analyze these databases, because they find valid patterns by analyzing a database with only some of the features described above. Therefore, knowledge-based software about how to find meaningful information efficiently by analyzing databases with these characteristics is required. In this article, we propose an intelligent information system that calculates the influence of the insertion time of each batch in a large-scale stream database by applying the sliding window model and mines recent high utility patterns without generating candidate patterns. In addition, a novel list-based data structure is suggested for a fast and efficient management of the time-sensitive stream databases. Moreover, our technique is compared with state-of-the-art algorithms through various experiments using real datasets and synthetic datasets. The experimental results show that our approach outperforms the previously proposed methods in terms of runtime, memory usage, and scalability.
Yoonji Baek, Unil Yun, Heonho Kim, Hyoju Nam, Jerry Chun-Wei Lin, Bay Vo, Witold Pedrycz
ACM Trans. Intell. Syst. Technol.6
2021 Trajectory Outlier Detection: New Problems and Solutions for Smart Cities
abstract
This article introduces two new problems related to trajectory outlier detection: (1) group trajectory outlier (GTO) detection and (2) deviation point detection for both individual and group of trajectory outliers. Five algorithms are proposed for the first problem by adapting DBSCAN , k nearest neighbors (kNN) , and feature selection (FS) . DBSCAN-GTO first applies DBSCAN to derive the micro clusters , which are considered as potential candidates. A pruning strategy based on density computation measure is then suggested to find the group of trajectory outliers. kNN-GTO recursively derives the trajectory candidates from the individual trajectory outliers and prunes them based on their density. The overall process is repeated for all individual trajectory outliers. FS-GTO considers the set of individual trajectory outliers as the set of all features, while the FS process is used to retrieve the group of trajectory outliers. The proposed algorithms are improved by incorporating ensemble learning and high-performance computing during the detection process. Moreover, we propose a general two-phase-based algorithm for detecting the deviation points, as well as a version for graphic processing units implementation using sliding windows. Experiments on a real trajectory dataset have been carried out to demonstrate the performance of the proposed approaches. The results show that they can efficiently identify useful patterns represented by group of trajectory outliers, deviation points, and that they outperform the baseline group detection algorithms.
Youcef Djenouri, Djamel Djenouri, Jerry Chun-Wei Lin
ACM Trans. Knowl. Discov. Data3
2021 Utility Mining Across Multi-Dimensional Sequences
abstract
Knowledge extraction from database is the fundamental task in database and data mining community, which has been applied to a wide range of real-world applications and situations. Different from the support-based mining models, the utility-oriented mining framework integrates the utility theory to provide more informative and useful patterns. Time-dependent sequence data are commonly seen in real life. Sequence data have been widely utilized in many applications, such as analyzing sequential user behavior on the Web, influence maximization, route planning, and targeted marketing. Unfortunately, all the existing algorithms lose sight of the fact that the processed data not only contain rich features (e.g., occur quantity, risk, and profit), but also may be associated with multi-dimensional auxiliary information, e.g., transaction sequence can be associated with purchaser profile information. In this article, we first formulate the problem of utility mining across multi-dimensional sequences, and propose a novel framework named MDUS to extract Multi-Dimensional Utility-oriented Sequential useful patterns. To the best of our knowledge, this is the first study that incorporates the time-dependent sequence-order, quantitative information, utility factor, and auxiliary dimension. Two algorithms respectively named MDUS EM and MDUS SD are presented to address the formulated problem. The former algorithm is based on database transformation, and the later one performs pattern joins and a searching method to identify desired patterns across multi-dimensional sequences. Extensive experiments are carried on six real-life datasets and one synthetic dataset to show that the proposed algorithms can effectively and efficiently discover the useful knowledge from multi-dimensional sequential databases. Moreover, the MDUS framework can provide better insight, and it is more adaptable to real-life situations than the current existing models.
Wensheng Gan, Jerry Chun-Wei Lin, Jiexiong Zhang, Hongzhi Yin, Philippe Fournier-Viger, Han-Chieh Chao, Philip S. Yu
ACM Trans. Knowl. Discov. Data2
2021 A Survey of Utility-Oriented Pattern Mining
abstract
The main purpose of data mining and analytics is to find novel, potentially useful patterns that can be utilized in real-world applications to derive beneficial knowledge. For identifying and evaluating the usefulness of different kinds of patterns, many techniques and constraints have been proposed, such as support, confidence, sequence order, and utility parameters (e.g., weight, price, profit, quantity, satisfaction, etc.). In recent years, there has been an increasing demand for utility-oriented pattern mining (UPM, or called utility mining). UPM is a vital task, with numerous high-impact applications, including cross-marketing, e-commerce, finance, medical, and biomedical applications. This survey aims to provide a general, comprehensive, and structured overview of the state-of-the-art methods of UPM. First, we introduce an in-depth understanding of UPM, including concepts, examples, and comparisons with related concepts. A taxonomy of the most common and state-of-the-art approaches for mining different kinds of high-utility patterns is presented in detail, including Apriori-based, tree-based, projection-based, vertical-/horizontal-data-format-based, and other hybrid approaches. A comprehensive review of advanced topics of existing high-utility pattern mining techniques is offered, with a discussion of their pros and cons. Finally, we present several well-known open-source software packages for UPM. We conclude our survey with a discussion on open and practical challenges in this field.
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao, Vincent S. Tseng, Philip S. Yu
IEEE Trans. Knowl. Data Eng.2
2020 Context-aware Adaptive Outlier Detection in Trajectory Data
abstract
With the advent of data mining and business processes automation, outlier detection has evolved into a major problem attracting significant research in relation to several application domains. Further advances in Global Positioning system, tracking of anomalous events based on data enhances effective decision making and pro-active measures to overcome risks and avoid unwarranted outputs. Significant work has been done in trajectory outlier detection although no singular approach fits all the domains. By including position and collective outliers on the same visualizations will enhance understanding of an outlier behavior. As such, we have leveraged Hidden Markov Method for prediction-based point outlier detection and pattern mining to identify points or segments of outliers in trajectory data.
Srinivas Danda, Ji Zhang 0001, Xiaohui Tao 0001, Jerry Chun-Wei Lin, Wenbin Zhang 0002
IEEE BigData4
2020 Mining High-Utility Sequential Patterns in Uncertain Databases
abstract
During our research conducted in this paper, we demonstrate a successful mining progress to mine the sequential high-utility patterns of uncertain databases. A PUL-Chain structure is developed and built in this paper with several pruning methods to decrease the search space of required patterns for mining efficiency improvement. In contrast to the standard HUS-Span, our experimental results show clearly that both in runtime as well as in the number of candidates discovered, the developed algorithms showed the effectiveness of the discovered patterns and its mining efficiency compared to the elder HUS-Span model. We present the details of our research here in this paper and also focus our attention to future directions that this research may take in the years to come.
Jerry Chun-Wei Lin, Gautam Srivastava 0001, Yuanfa Li, Tzung-Pei Hong, Shyue-Liang Wang
IEEE BigData1
2020 Fuzzy High-Utility Pattern Mining based on the Hadoop Framework
abstract
In this paper, fuzzy-set theory is first used and a new algorithm called efficient fuzzy high-utility itemset mining (EFUPM) algorithm is designed to discover the fuzzy high-utility patterns from a single machine. Two upper-bounds are then estimated to early prune the unpromising candidates in the search space. To handle the large-scale of big datasets, the Hadoop-based fuzzy high-utility pattern mining (HFUPM) algorithm is then developed to discover the fuzzy high-utility patterns based on the Hadoop framework. Experimental results show that the proposed algorithms can perform well to mine the required fuzzy high-utility patterns whether in a single machine or a large-scale environment compared to the state-of-the-art approaches.
Jimmy Ming-Tai Wu, Gautam Srivastava 0001, Jerry Chun-Wei Lin
IEEE BigData4
2020 High-Utility Pattern Mining in Hadoop Environments
abstract
In this article, we present an Efficient High Utility Pattern Mining framework to mine high-utility patterns with a reasonable pruning strategy to speed up the mining performance. Concurrently, for solving the problem of excessive data volume in the current era, we applied the developed framework to the MapReduce architecture used for improving the feasibility in practical applications. Our in-depth work in this paper culminates with some experimental results that clearly show that our proposed framework can perform well to mine the required pattern in a big-data dataset and shows great performance in a Hadoop computing cluster.
Jimmy Ming-Tai Wu, Gautam Srivastava 0001, Jerry Chun-Wei Lin
IEEE BigData4
2020 Mining Attribute Evolution Rules in Dynamic Attributed Graphs
Philippe Fournier-Viger, Ganghuan He, Jerry Chun-Wei Lin, Heitor Murilo Gomes
DaWaK3
2020 Mining Locally Trending High Utility Itemsets
Philippe Fournier-Viger, Jerry Chun-Wei Lin, Jaroslav Frnda
PAKDD (2)3
2020 Discovering rare correlated periodic patterns in multiple sequences
Philippe Fournier-Viger, Zhitian Li, Jerry Chun-Wei Lin, R. Uday Kiran
Data Knowl. Eng.4
2020 Efficiently mining erasable stream patterns for intelligent systems over uncertain data
abstract
Data mining is a method for extracting useful information that is necessary for a system from a database. As the types of data processed by the system are diversified, the transformed pattern mining techniques for processing these type of data have been proposed. Unlike the traditional pattern mining methods, erasable pattern mining is a technique for finding the patterns that can be removed by coming with a small profit. Erasable pattern mining should be able to process data by considering both the environment that the data are generated from and the characteristics of the data. An uncertain database is a database that is composed of uncertain data. Since erasable patterns discovered from uncertain data contain significant information, these patterns need to be extracted. In addition, databases gradually increase, because the data from various fields is generated and accumulated over data streams. Data streams should be processed as intelligently as possible to provide the useful data to the system in real time. In this paper, we propose an efficient erasable pattern mining algorithm that processes uncertain data that is generated over data streams. The uncertain erasable patterns discovered through the suggested technique are more meaningful information by considering the probability of the item and the profit. Moreover, the proposed method can perform efficient mining operations by using both tree and list structures. The performance of the suggested algorithm is verified through the performance tests compared with state-of-the-art algorithms using real data sets and synthetic data sets.
Yoonji Baek, Unil Yun, Jerry Chun-Wei Lin, Eunchul Yoon, Hamido Fujita
Int. J. Intell. Syst.3
2020 ProUM: Projection-based utility mining on sequence data
Wensheng Gan, Jerry Chun-Wei Lin, Jiexiong Zhang, Han-Chieh Chao, Hamido Fujita, Philip S. Yu
Inf. Sci.2
2020 Efficient approach of recent high utility stream pattern mining with indexed list structure and pruning strategy considering arrival times of transactions
Hyoju Nam, Unil Yun, Eunchul Yoon, Jerry Chun-Wei Lin
Inf. Sci.4
2020 High average-utility sequential pattern mining based on uncertain databases
Jerry Chun-Wei Lin, Ting Li 0011, Matin Pirouz, Ji Zhang 0001, Philippe Fournier-Viger
Knowl. Inf. Syst.1
2019 HUE-Span: Fast High Utility Episode Mining
Philippe Fournier-Viger, Jerry Chun-Wei Lin, Unil Yun
ADMA3
2019 Utility-Driven Mining of High Utility Episodes
abstract
Sequence data, e.g., complex event sequence, is more commonly seen than other types of data (e.g., transaction data) in real-world applications. For the mining task from sequence data, several problems have been formulated, such as sequential pattern mining, episode mining, and sequential rule mining. As one of the fundamental problems, episode mining has often been studied. The common wisdom is that discovering frequent episodes is not useful enough. In this paper, we propose an efficient utility mining approach namely UMEpi: Utility Mining of high-utility Episodes from complex event sequence. We propose the concept of remaining utility of episode, and achieve a tighter upper bound, namely episode-weighted utilization (EWU), which will provide better pruning. Thus, the optimized EWU-based pruning strategy can achieve better improvements in mining efficiency. Finally, experiments on two real-life datasets demonstrate that UMEpi can discover the complete high-utility episodes from complex event sequence, while state-of-the-art algorithms fail to return the correct results. Besides, the improved variants of UMEpi outperforms the baseline.
Wensheng Gan, Jerry Chun-Wei Lin, Han-Chieh Chao, Philip S. Yu
IEEE BigData2
2019 Mining Temporal Fuzzy Utility Itemsets by Tree Structure
abstract
More complicated than fuzzy data mining, temporal fuzzy utility data mining takes into account the temporal factor of transactions, purchased quantities, item profits, and linguistic terms. In this paper, a tree structure modified from the frequent-pattern tree is designed and a mining algorithm based on it was proposed to extract high temporal fuzzy utility patterns from transactional datasets with the temporal property. The method requires two-phase processing to find all high temporal fuzzy utility itemsets. Experimental results show that the proposed algorithm performs better than the Apriori-based mining algorithm.
Tzung-Pei Hong, Wei-Ming Huang, Shu-Min Li, Shyue-Liang Wang, Jerry Chun-Wei Lin
IEEE BigData6
2019 Mining High-Utility Sequential Patterns from Big Datasets
abstract
High-Utility Sequential Pattern Mining (HUSPM) has become an emerging issue in recent decades since it reveals more information such as the utility and sequence factors for knowledge discovery. For the previous works, many algorithms were presented to speed up the mining performance regarding a single machine with small datasets. In real-world applications, the size of dataset can be collected from many places or devices, such as PC, Internet of Things (IoT), mobile devices, and shopping malls, among others. It is necessary to build an efficient model to handle the big dataset for HUSPM. In this paper, we present a four-stages MapReduce framework based on the Spark platform for mining the high-utility sequential patterns from a very large database. From the experimental results, we then can observe that the designed model outperforms the state-of-the-art approaches for handling the very big dataset.
Jerry Chun-Wei Lin, Yuanfa Li, Philippe Fournier-Viger, Youcef Djenouri, Shyue-Liang Wang
IEEE BigData1
2019 A GA-based Framework for Mining High Fuzzy Utility Itemsets
abstract
Comparing to frequent itemset mining (FIM), utility-pattern mining receives increasing attention in the field of data mining recently. With the flourishing development of utility-pattern mining, most studies focused on the efficiency problem by considering the efficient data structure to compress the original data and pruning strategies to reduce the search space for knowledge discovery. However, those approaches can only handle the binary situation, thus the discovered knowledge cannot be represented as the linguistic variables. Previous works have addressed this problem by introducing the generic approaches to find the high fuzzy utility itemsets in a small database. In real-world situations, the dataset may be very large, and it is costly to mine all the required information from a very large database. In this paper, we first present a HFUI-GA framework to discover the high fuzzy utility itemsets in a limited time. Several improvement strategies are also proposed to speed up the evolutionary progress. Experiments are then conducted to show the performance of the variants of the designed HFUI-GA framework in terms of number of the discovered high fuzzy utility itemsets (HFUIs) and the results are convincing to show that the designed GA-based HFUI-GA framework is a promising solution to mine for HFUIs.
Jimmy Ming-Tai Wu, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Tomasz Wiktorski, Tzung-Pei Hong, Matin Pirouz
IEEE BigData2
2019 Finding Strongly Correlated Trends in Dynamic Attributed Graphs
Philippe Fournier-Viger, Zhi Cheng, Jerry Chun-Wei Lin, Nazha Selmaoui-Folcher
DaWaK4
2019 Discovering and Visualizing Efficient Patterns in Cost/Utility Sequences
Philippe Fournier-Viger, Jerry Chun-Wei Lin, Tin Truong 0001
DaWaK3
2019 Highly Efficient Pattern Mining Based on Transaction Decomposition
abstract
This paper introduces a highly efficient pattern mining technique called Clustering-Based Pattern Mining (CBPM). This technique discovers relevant patterns by studying the correlation between transactions in transaction databases using clustering techniques. The set of transactions are first clus-tered using the k-means algorithm, where highly correlated transactions are grouped together. Next, the relevant patterns are derived by applying a pattern mining algorithm to each cluster. We present two different pattern mining algorithms, one approximate and one exact. We demonstrate the efficiency and effectiveness of CBPM through a thorough experimental evaluation.
Youcef Djenouri, Jerry Chun-Wei Lin, Kjetil Nørvåg, Heri Ramampiaro
ICDE2
2019 Exploiting GPU parallelism in improving bees swarm optimization for mining big transactional databases
Youcef Djenouri, Djamel Djenouri, Asma Belhadi, Philippe Fournier-Viger, Jerry Chun-Wei Lin, Ahcène Bendjoudi
Inf. Sci.5
2019 Efficient algorithms to identify periodic patterns in multiple sequences
Philippe Fournier-Viger, Zhitian Li, Jerry Chun-Wei Lin, R. Uday Kiran, Hamido Fujita
Inf. Sci.3
2019 Mining local and peak high utility itemsets
Philippe Fournier-Viger, Jerry Chun-Wei Lin, Hamido Fujita, Yun Sing Koh
Inf. Sci.3
2019 Correlated utility-based pattern mining
Wensheng Gan, Jerry Chun-Wei Lin, Han-Chieh Chao, Hamido Fujita, Philip S. Yu
Inf. Sci.2
2019 BILU-NEMH: A BILU neural-encoded mention hypergraph for mention extraction
abstract
The natural language processing (NLP) denotes a technique used to process data such as text and speech. Some of the fundamental research in NLP includes the named entity recognition, which recognizes the named entities (i.e., persons and companies) from texts, the semantic parsing, which converts a natural language utterance to a logical form, and the co-reference resolution, which extracts the nouns (including pronouns and noun phrases) pointing to the same reference body. In this paper, we focus on the mention extraction and classification, proposing a neural-encoded mention-hypergraph model named the BILU-NEMH to extract the mention entities from a content. The proposed BILU-NEMH model combines a mention hypergraph model with the encoding schema and neural network. The proposed model can effectively capture the overlapping mention entities of an unbounded length. The proposed model was verified by the experiments, and the obtained experimental results showed that the proposed model achieved better performance and greater effectiveness than the existing related models on most standard datasets.
Jerry Chun-Wei Lin, Yinan Shao, Philippe Fournier-Viger, Hamido Fujita
Inf. Sci.1
2019 A Survey of Parallel Sequential Pattern Mining
abstract
With the growing popularity of shared resources, large volumes of complex data of different types are collected automatically. Traditional data mining algorithms generally have problems and challenges including huge memory cost, low processing speed, and inadequate hard disk space. As a fundamental task of data mining, sequential pattern mining (SPM) is used in a wide variety of real-life applications. However, it is more complex and challenging than other pattern mining tasks, i.e., frequent itemset mining and association rule mining, and also suffers from the above challenges when handling the large-scale data. To solve these problems, mining sequential patterns in a parallel or distributed computing environment has emerged as an important issue with many applications. In this article, an in-depth survey of the current status of parallel SPM (PSPM) is investigated and provided, including detailed categorization of traditional serial SPM approaches, and state-of-the art PSPM. We review the related work of PSPM in details including partition-based algorithms for PSPM, apriori-based PSPM, pattern-growth-based PSPM, and hybrid algorithms for PSPM, and provide deep description (i.e., characteristics, advantages, disadvantages, and summarization) of these parallel approaches of PSPM. Some advanced topics for PSPM, including parallel quantitative/weighted/utility SPM, PSPM from uncertain data and stream data, hardware acceleration for PSPM, are further reviewed in details. Besides, we review and provide some well-known open-source software of PSPM. Finally, we summarize some challenges and opportunities of PSPM in the big data era.
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao, Philip S. Yu
ACM Trans. Knowl. Discov. Data2
2019 High-Utility Itemset Mining with Effective Pruning Strategies
abstract
High-utility itemset mining is a popular data mining problem that considers utility factors, such as quantity and unit profit of items besides frequency measure from the transactional database. It helps to find the most valuable and profitable products/items that are difficult to track by using only the frequent itemsets. An item might have a high-profit value which is rare in the transactional database and has a tremendous importance. While there are many existing algorithms to find high-utility itemsets (HUIs) that generate comparatively large candidate sets, our main focus is on significantly reducing the computation time with the introduction of new pruning strategies. The designed pruning strategies help to reduce the visitation of unnecessary nodes in the search space, which reduces the time required by the algorithm. In this article, two new stricter upper bounds are designed to reduce the computation time by refraining from visiting unnecessary nodes of an itemset. Thus, the search space of the potential HUIs can be greatly reduced, and the mining procedure of the execution time can be improved. The proposed strategies can also significantly minimize the transaction database generated on each node. Experimental results showed that the designed algorithm with two pruning strategies outperform the state-of-the-art algorithms for mining the required HUIs in terms of runtime and number of revised candidates. The memory usage of the designed algorithm also outperforms the state-of-the-art approach. Moreover, a multi-thread concept is also discussed to further handle the problem of big datasets.
Jimmy Ming-Tai Wu, Jerry Chun-Wei Lin, Ashish Tamrakar
ACM Trans. Knowl. Discov. Data2
2018 An Approach for Diverse Group Stock Portfolio Optimization Using the Fuzzy Grouping Genetic Algorithm
Chun-Hao Chen, Bing-Yang Chiang, Tzung-Pei Hong, Ding-Chau Wang, Jerry Chun-Wei Lin
ACIIDS (1)5
2018 Discovering High Utility Change Points in Customer Transaction Data
Philippe Fournier-Viger, Jerry Chun-Wei Lin, Yun Sing Koh
ADMA3
2018 A Genetic Algorithm Based Technique for Outlier Detection with Fast Convergence
Ji Zhang 0001, Zewen Hu, Hongzhou Li, Liang Chang 0003, Youwen Zhu, Jerry Chun-Wei Lin, Yongrui Qin
ADMA7
2018 CoUPM: Correlated Utility-based Pattern Mining
abstract
In the field of data mining, many utility-oriented mining approaches have been extensively studied. Previous studies have, however, the limitation that they rarely consider the inherent correlation of items among the discovered patterns. For example, from the purchase behavior, a high-utility group of products (w.r.t. multi-products) may contain the items with both high or low utility. This pattern is also considered as a valuable pattern even if they may not be highly correlated, or even happened together by the chance. In this paper, we propose an efficient utility mining approach namely non-redundant Correlated high-Utility Pattern Miner (CoUPM) by considering both strong positive correlation and profitable value of the products. The derived patterns with high utility and strong correlation can lead to more insightful availability than those patterns only have high utility values. The utility-list structure is maintained and applied to store necessary information of correlation and utility. Several pruning strategies are further developed to improve the efficiency for discovering the desired patterns. Experimental results show that the non-redundant correlated high-utility patterns have more effectiveness than some other kinds of patterns. Moreover, the proposed CoUPM algorithm significantly outperforms the state-of-the-art algorithm.
Wensheng Gan, Jerry Chun-Wei Lin, Han-Chieh Chao, Tzung-Pei Hong, Philip S. Yu
IEEE BigData2
2018 Privacy Preserving Utility Mining: A Survey
abstract
In big data era, the collected data usually contains rich information and hidden knowledge. Utility-oriented pattern mining and analytics have shown a powerful ability to explore these ubiquitous data, which may be collected from various fields and applications, such as market basket analysis, retail, click-stream analysis, medical analysis, and bioinformatics. However, analysis of these data with sensitive private information raises privacy concerns. To achieve better trade-off between utility maximizing and privacy preserving, Privacy-Preserving Utility Mining (PPUM) has become a critical issue in recent years. In this paper, we provide a comprehensive overview of PPUM. We first present the background of utility mining, privacy-preserving data mining and PPUM, then introduce the related preliminaries and problem formulation of PPUM, as well as some key evaluation criteria for PPUM. In particular, we present and discuss the current state-of-the-art PPUM algorithms, as well as their advantages and deficiencies in detail. Finally, we highlight and discuss some technical challenges and open directions for future research on PPUM.
Wensheng Gan, Jerry Chun-Wei Lin, Han-Chieh Chao, Shyue-Liang Wang, Philip S. Yu
IEEE BigData2
2018 Reducing Database Scan in Maintaining Erasable Itemsets from Product Deletion
abstract
Mining erasable itemsets is a problem derived from the production planning of the manufacturing industry. In the past, the erasable-itemset mining with product insertion has been designed. In this paper, we further consider the maintenance problem from product deletion. We propose an efficient method to solve it. The method is based on the concept of pre-large itemsets to maintain the correct results for product deletion. It improves the efficiency of the mining process by further reducing the number of times required for rescanning the database. When the ratio of the number of deleted products over the total number of products in the original database is less than a certain degree, there will be no need to rescan the original product database for maintaining the correct mining results. Finally, the experiments are made to evaluate the performance of the proposed approach.
Tzung-Pei Hong, Chia-Che Li, Shyue-Liang Wang, Jerry Chun-Wei Lin
IEEE BigData4
2018 SLIND: Identifying Stable Links in Online Social Networks
Ji Zhang 0001, Leonard Tan, Xiaohui Tao 0001, Xiaoyao Zheng, Yonglong Luo, Jerry Chun-Wei Lin
DASFAA (2)6
2018 Discovering Periodic Patterns Common to Multiple Sequences
Philippe Fournier-Viger, Zhitian Li, Jerry Chun-Wei Lin, R. Uday Kiran, Hamido Fujita
DaWaK3
2018 Anonymization of Multiple and Personalized Sensitive Attributes
Jerry Chun-Wei Lin, Qiankun Liu 0002, Philippe Fournier-Viger, Youcef Djenouri, Ji Zhang 0001
DaWaK1
2018 Mining Local High Utility Itemsets
Philippe Fournier-Viger, Jerry Chun-Wei Lin, Hamido Fujita, Yun Sing Koh
DEXA (2)3
2018 A Recommender System with Advanced Time Series Medical Data Analysis for Diabetes Patients in a Telehealth Environment
Raid Lafta, Ji Zhang 0001, Xiaohui Tao 0001, Jerry Chun-Wei Lin, Fulong Chen 0002, Yonglong Luo, Xiaoyao Zheng
DEXA (2)4
2018 A Metaheuristic Algorithm for Hiding Sensitive Itemsets
Jerry Chun-Wei Lin, Yuyu Zhang, Philippe Fournier-Viger, Youcef Djenouri, Ji Zhang 0001
DEXA (2)1
2018 On Link Stability Detection for Online Social Networks
Ji Zhang 0001, Xiaohui Tao 0001, Leonard Tan, Jerry Chun-Wei Lin, Hongzhou Li, Liang Chang 0003
DEXA (1)4
2018 Fast and effective cluster-based information retrieval using frequent closed itemsets
Youcef Djenouri, Asma Belhadi, Philippe Fournier-Viger, Jerry Chun-Wei Lin
Inf. Sci.4
2018 Exploiting highly qualified pattern with frequency and weight occupancy
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao, Justin Zhijun Zhan, Ji Zhang 0001
Knowl. Inf. Syst.2
2017 Quasi-erasable itemset mining
abstract
Erasable-itemset mining used in production planning identifies itemsets (or components) that, if removed, would not affect profits. Formally, an itemset is erasable if its gain ratio is equal to or smaller than a given maximum gain-ratio threshold r. Since new products with different components may be added, the original batch algorithm will waste time in gathering up-to-date erasable itemsets. In this paper, we propose the concept of the ε-quasi-erasable itemsets and use it to improve mining performance. The itemsets in both the original database and the new product can then be divided into erasable, ε-quasi-erasable, and nonerasable. Thus, there are nine combinations that are then processed in different ways. Experiments are finally made to verify the performance.
Tzung-Pei Hong, Lu-Hung Chen, Shyue-Liang Wang, Jerry Chun-Wei Lin, Bay Vo
IEEE BigData4
2017 Extracting Non-redundant Correlated Purchase Behaviors by Utility Measure
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao
DaWaK2
2017 Mining High-Utility Itemsets with Both Positive and Negative Unit Profits from Uncertain Databases
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao, Vincent S. Tseng
PAKDD (1)2
2017 A two-phase approach to mine short-period high-utility itemsets in transactional databases
Jerry Chun-Wei Lin, Jiexiong Zhang, Philippe Fournier-Viger, Tzung-Pei Hong, Ji Zhang 0001
Adv. Eng. Informatics1
2017 FDHUP: Fast algorithm for mining discriminative high utility patterns
Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong, Han-Chieh Chao
Knowl. Inf. Syst.1
2017 EFIM: a fast and memory efficient algorithm for high-utility itemset mining
Souleymane Zida, Philippe Fournier-Viger, Jerry Chun-Wei Lin, Cheng-Wei Wu, Vincent S. Tseng
Knowl. Inf. Syst.3
2016 Mining Discriminative High Utility Patterns
Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong
ACIIDS (2)1
2016 Efficient Mining of Fuzzy Frequent Itemsets with Type-2 Membership Functions
Jerry Chun-Wei Lin, Xianbiao Lv, Philippe Fournier-Viger, Tsu-Yang Wu, Tzung-Pei Hong
ACIIDS (2)1
2016 Mining Recent High Expected Weighted Itemsets from Uncertain Databases
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao
APWeb (1)2
2016 Mining Recent High-Utility Patterns from Temporal Databases with Time-Sensitive Constraint
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao
DaWaK2
2016 Mining Minimal High-Utility Itemsets
Philippe Fournier-Viger, Jerry Chun-Wei Lin, Cheng-Wei Wu, Vincent S. Tseng, Usef Faghihi
DEXA (1)2
2016 More Efficient Algorithms for Mining High-Utility Itemsets with Multiple Minimum Utility Thresholds
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao
DEXA (1)2
2016 The SPMF Open-Source Data Mining Library Version 2
Philippe Fournier-Viger, Jerry Chun-Wei Lin, Antonio Gomariz, Ted Gueniche, Azadeh Soltani, Zhi-Hong Deng 0001, Hoang Thanh Lam
ECML/PKDD (3)2
2016 More Efficient Algorithm for Mining Frequent Patterns with Multiple Minimum Supports
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao
WAIM (1)2
2016 Efficient Mining of Uncertain Data for High-Utility Itemsets
Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong, Vincent S. Tseng
WAIM (1)1
2016 Fast algorithms for mining high-utility itemsets with various discount strategies
Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong, Vincent S. Tseng
Adv. Eng. Informatics1
2016 An efficient algorithm to mine high average-utility itemsets
Jerry Chun-Wei Lin, Ting Li 0011, Philippe Fournier-Viger, Tzung-Pei Hong, Justin Zhijun Zhan, Miroslav Voznak
Adv. Eng. Informatics1
2016 Inferring social network user profiles using a partial social graph
Raïssa Yapan Dougnon, Philippe Fournier-Viger, Jerry Chun-Wei Lin, Roger Nkambou
J. Intell. Inf. Syst.3
2015 Mining Weighted Frequent Itemsets with the Recency Constraint
Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong
APWeb1
2015 Mining high-utility itemsets with various discount strategies
abstract
In recent years, mining high-utility itemsets (HUIs) has become as a key topic in data mining. However, most of the developed algorithms assume the unrealistic situations that unit profits of items remain unchanged over time. But in real-life situations, the profit of an item or itemset varies as a function of cost prices, sales prices and sales strategies. In this paper, a novel framework for mining HUIs with two algorithms under various Discount strategies (HUID) are introduced. HUID-tp is based on various discount strategies and a novel downward closure property to mine the complete set of HUIs. HUID-Miner is an algorithm relying on a compact data structure (Positive-and-Negative Utility-list, PNU-list) and new pruning strategies to efficiently discover HUIs without candidate generation, while considerably reducing the size of the search space. Furthermore, a strategy named Estimated Utility Co-occurrence Strategy which stores the relationships between 2-itemsets is also adopted in the proposed improvement HUID-EMiner algorithm to speed up computation. An extensive experimental study carried on several real-life datasets shows the performance of the proposed algorithms.
Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong, Vincent S. Tseng
DSAA1
2015 A fast updated algorithm to maintain the discovered high-utility itemsets for transaction modification
Jerry Chun-Wei Lin, Wensheng Gan, Tzung-Pei Hong
Adv. Eng. Informatics1
2015 Efficient algorithms for mining up-to-date high-utility patterns
Jerry Chun-Wei Lin, Wensheng Gan, Tzung-Pei Hong, Vincent S. Tseng
Adv. Eng. Informatics1
2015 Efficient updating of discovered high-utility itemsets for transaction deletion in dynamic databases
Jerry Chun-Wei Lin, Tzung-Pei Hong, Guo-Cheng Lan, Jia-Wei Wong, Wen-Yang Lin
Adv. Eng. Informatics1
2014 Incrementally Updating High-Utility Itemsets with Transaction Insertion
Jerry Chun-Wei Lin, Wensheng Gan, Tzung-Pei Hong, Jeng-Shyang Pan 0001
ADMA1
2014 Maintenance of prelarge trees for data mining with modified records
Jerry Chun-Wei Lin, Tzung-Pei Hong
Inf. Sci.1
2012 Integration of Multiple Fuzzy FP-trees
Tzung-Pei Hong, Jerry Chun-Wei Lin, Tsung-Ching Lin, Shing-Tai Pan
ACIIDS (1)2
2010 Efficiently Mining High Average Utility Itemsets with a Tree Structure
Jerry Chun-Wei Lin, Tzung-Pei Hong, Wen-Hsiang Lu
ACIIDS (1)1
2007 Using the Pre-FUFP Algorithm for Handling New Transactions in Incremental Mining
abstract
In the past, we proposed a Fast Updated FP-tree (FUFP-tree) structure to efficiently handle new transactions and to make the tree update process become easier. In this paper, we attempt to modify the FUFP-tree construction based on the concept of pre-large itemsets. Pre-large itemsets are defined by a lower support threshold and an upper support threshold. The proposed approach can achieve a good execution time for tree construction especially when each time a small number of transactions are inserted. Experimental results also show that the proposed Pre-FUFP maintenance algorithm has a good performance for incrementally handling new transactions.
Jerry Chun-Wei Lin, Tzung-Pei Hong, Wen-Hsiang Lu
CIDM1