Jimmy Ming-Tai Wu

dblp:183/2637 · also Jimmy Ming-Thai Wu · DBLP profile ↗
← Back
38ranked-venue papers
25as first author
26since 2021 · last 2026
0000-0003-3740-2102ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 13 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 7 · 5 first-author · 3 since 2021Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Computer networks · 4 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing Opening Range Breakout Strategies with LSTM-Based True Range Prediction
Mu-En Wu, Sheng-Chi Luo, Wei-Xi Lin, Chien-Ping Chung, Jun-Yo Wu, Jimmy Ming-Tai Wu
ACIIDS (2)6
2026 A hybrid multi-model architecture for at-the-money options forecasting
Mu-En Wu, Kuan-Li Ko, Jimmy Ming-Tai Wu
Eng. Appl. Artif. Intell.3
2026 A study on signal filtering techniques in trend-following strategies with LSTM integration
Jimmy Ming-Tai Wu, Yi-Chun Cheng, Sheng-Chi Luo, Ju-Fang Yen, Mu-En Wu
Expert Syst. Appl.1
2025 Convert index trading to option strategies via LSTM architecture
abstract
Abstract In the past, most strategies were mainly designed to focus on stocks or futures as the trading target. However, due to the enormous number of companies in the market, it is not easy to select a set of stocks or futures for investment. By investigating each company’s financial situation and the trend of the overall financial market, people can invest precisely in the market and choose to go long or short. Moreover, how to determine the position size of the transaction is also a problematic issue. In the past, many money management theories were based on the Kelly criterion. And they put a certain percentage of their total funds into the market for trading. Nonetheless, three massive problems cannot be overcome. First, futures are leveraged transactions, and extra funds must be deposited as margin. It causes that the position size is hard to be estimated by the Kelly criterion. The second point is that the trading strategy is difficult to determine the winning rate in the financial market and cannot be brought into the Kelly criterion to calculate the optimal fraction. Last, the financial data are always massive. A big data technique should be applied to resolve this issue and enhance the performance of the framework to reveal knowledge in the financial data. Therefore, in this paper, a concept of converting the original futures trading strategy into options trading is proposed. An LSTM (long short-term memory)-based framework is proposed to predict the profit probability of the original futures strategy and convert the corresponding daily take-profit and stop-loss points according to the delta value of the options. Finally, the proposed framework brings the results into the Kelly criterion to get the optimal fraction of options trading. The final research results show that options trading is closer to the optimal fraction calculated by the Kelly criterion than futures trading. If the original futures trading strategy can profit, the benefits after converting to options trading can be further superior.
Jimmy Ming-Tai Wu, Mu-En Wu, Pang-Jen Hung, Mohammad Mehedi Hassan, Giancarlo Fortino
Neural Comput. Appl.1
2025 High-utility sequential pattern mining in incremental database
Huizhen Yan, Fengyang Li, Ming-Chia Hsieh, Jimmy Ming-Tai Wu
J. Supercomput.4
2024 SeqClin: Pattern-Based Analysis and Classification of Clinical Datasets
abstract
Accurate analysis and classification of clinical datasets are crucial for understanding disease patterns, identifying risk factors and devising targeted interventions that ultimately contribute towards effective healthcare systems and improved patient outcomes. However, existing analysis and classification methods often fall short of effectively capturing complex sequential relationships within patient data and have limited interpretability. To overcome these challenges, we introduce SeqClin, a novel approach that utilizes frequent pattern mining to obtain valuable sequential information from clinical datasets. SeqClin first transforms clinical datasets into an appropriate format. Then, it employs sequential pattern mining algorithms to find frequent sequential patterns as well as rules of patient features in the datasets. These identified feature patterns and their respective values are then used for classification/detection. The performance of SeqClin is evaluated on four clinical datasets, where six classification models and evaluation metrics are employed for a comprehensive assessment. The obtained results show that the proposed approach surpassed previous approaches, with the extracted patterns and rules providing valuable insights into the key patient features and their values in clinical datasets.
M. Saqib Nawaz, Philippe Fournier-Viger, Jimmy Ming-Tai Wu
BIBM3
2024 On the design of searching algorithm for parameter plateau in quantitative trading strategies using particle swarm optimization
abstract
Quantitative trading, relying on diverse parameter combinations, is becoming increasingly the norm for trading strategies in financial investments. The performance of these strategies is intricately linked to these parameters. However, the performance on the training set after backtesting does not ensure success on a test set and may lead to overfitting. This study emphasizes enhancing stability and robustness in trading-strategy parameters by introducing a ’parameter plateau.’ Traditional brute-force methods for exploring high-dimensional parameter spaces can be intricate and time-consuming. To address this challenge, we present an efficient alternative that identifies stable and robust parameters by configuring parameter plateaus to mitigate overfitting risks. A step-by-step search algorithm is proposed to determine the optimal parameters, leveraging the power of particle-swarm optimization. In continuous, multi-dimensional solution spaces, particle-swarm optimization is invaluable for the swift and effective discovery of the desired solutions. Experiments underscore the substantial influence of the parameter plateau concept on parameter selection, highlighting the pivotal role of particle-swarm optimization in efficiently navigating complex solution spaces and thereby enabling the discovery of stable and profitable trading strategies.
Jimmy Ming-Tai Wu, Wen-Yu Lin, Ko-Wei Huang, Mu-En Wu
Knowl. Based Syst.1
2023 Updating Top-k Dominate Individuals with Incomplete Data Addition
abstract
Top-k dominance (TKD) query is an extended query method of skyline query and top-k query, which reveals the top-k dominant individuals in an incomplete dataset by analyzing the dominance relationships between individuals and is a common decision tool in intelligent recommendation applications. This research proposes two parallel query algorithms based on Spark computing engine to address the shortcomings of the parallel top-k-dominated query algorithms for dynamic incomplete datasets. The designed model achieves good performance in terms of runtime performance compared to previous studies.
Jimmy Ming-Tai Wu, Ke Wang 0068, Huizhen Yan, Chao-Chun Chen, Pei-Wei Chen, Jerry Chun-Wei Lin
ISIT1
2023 A Federated Mining Framework for Complete Frequent Itemsets
abstract
In this paper, we address the common features of horizontal federated learning in data mining and propose a federated mining framework, which adopts a client-server model that cooperates with multiple data-source clients. The proposed algorithm handles client-side mining and server-side aggregation. For client-side mining, the algorithm uses prelarge itemsets to collect additional information for the server to integrate the clients' local mining results. For server-side aggregation, the algorithm considers the characteristics of large and prelarge itemsets sent from the clients and use a boundary strategy for integration. Experiments show that our method acquires the complete mined results while protecting data.
Tzung-Pei Hong, Ya-Ping Hsu, Chun-Hao Chen, Jimmy Ming-Tai Wu
SMC4
2023 Mining Skyline Patterns from Big Data Environments based on a Spark Framework
Jimmy Ming-Tai Wu, Huiying Zhou, Jerry Chun-Wei Lin, Gautam Srivastava 0001, Mohamed Baza
J. Grid Comput.1
2023 Mining skyline frequent-utility patterns from big data environment based on MapReduce framework
abstract
When the concentration focuses on data mining, frequent itemset mining (FIM) and high-utility itemset mining (HUIM) are commonly addressed and researched. Many related algorithms are proposed to reveal the general relationship between utility, frequency, and items in transaction databases. Although these algorithms can mine FIMs or HUIMs quickly, these algorithms merely take into account frequency or utility as a unilateral criterion for itemsets but the other factors (e.g., distance, price) could be also valuable for decision-making. A new skyline framework has been presented to mine frequent high utility patterns (SFUPs) to better support user decision-making. Several new algorithms have been proposed one after another. However, the Internet of Things (IoT), mobile Internet, and traditional Internet are generating massive amounts of data every day, and these cutting-edge standalone algorithms can not satisfy the new challenge of finding interesting patterns from this data. Big Data uses a distributed architecture in the form of cloud computing to filter and process this data to extract useful information. This paper proposes a novel parallel algorithm on Hadoop as a three-stage iterative algorithm based on MapReduce. MapReduce is used to divide the mining tasks of the whole large data set into multiple independent sub-tasks to find frequent and high utility patterns in parallel. Numerous experiments were done in this paper, and from the results, the algorithm can handle large datasets and show good performance on Hadoop clusters.
Jimmy Ming-Tai Wu, Mu-En Wu, Jerry Chun-Wei Lin
Intell. Data Anal.1
2023 A graph-based CNN-LSTM stock price prediction algorithm with leading indicators
abstract
Abstract In today’s society, investment wealth management has become a mainstream of the contemporary era. Investment wealth management refers to the use of funds by investors to arrange funds reasonably, for example, savings, bank financial products, bonds, stocks, commodity spots, real estate, gold, art, and many others. Wealth management tools manage and assign families, individuals, enterprises, and institutions to achieve the purpose of increasing and maintaining value to accelerate asset growth. Among them, in investment and financial management, people’s favorite product of investment often stocks, because the stock market has great advantages and charm, especially compared with other investment methods. More and more scholars have developed methods of prediction from multiple angles for the stock market. According to the feature of financial time series and the task of price prediction, this article proposes a new framework structure to achieve a more accurate prediction of the stock price, which combines Convolution Neural Network (CNN) and Long–Short-Term Memory Neural Network (LSTM). This new method is aptly named stock sequence array convolutional LSTM (SACLSTM). It constructs a sequence array of historical data and its leading indicators (options and futures), and uses the array as the input image of the CNN framework, and extracts certain feature vectors through the convolutional layer and the layer of pooling, and as the input vector of LSTM, and takes ten stocks in U.S.A and Taiwan as the experimental data. Compared with previous methods, the prediction performance of the proposed algorithm in this article leads to better results when compared directly.
Jimmy Ming-Tai Wu, Zhongcui Li, Norbert Herencsar, Bay Vo, Jerry Chun-Wei Lin
Multim. Syst.1
2023 A Privacy Frequent Itemsets Mining Framework for Collaboration in IoT Using Federated Learning
abstract
Rapid advancement of industrial internet of things (IoT) technology has changed the supply chain network to an open system to meet the high demand for individualized products and provide better customer experiences. However the open-system supply chain has forced many small and midsize enterprises (SMEs) to adopt vertical integration by being divided into smaller companies with a distinctive business for each SME but a central alliance to produce a range of products and gain competencies. Therefore, existing models do not guarantee the protection of data privacy of individual SMEs. Moreover, especially for the IoT environment, collecting data in a secure way and revealing valuable knowledge in an IoT network is difficult. How to share data in a secure framework is of paramount importance in the internet of behavior field. In this article, a privacy-preserving data-mining framework is proposed for joint-venture industrial collaborative activities by combining federated learning and a “pre-large concept” of data-mining techniques. The novelty of the proposed approach is that, while mining high-utility itemsets (HUIs) from multiple datasets, it does not require direct data sharing. In the proposed method, the federated-learning framework can learn from aggregated learning parameters without scanning all data from different sets. The pre-large concept in this approach reduces the amount of scanning into different datasets. Thus, the approach makes it possible to train federated learning more quickly while protecting the privacy of individual data owners. The approach has been tested on real industrial datasets in a collaborative environment. Extensive experimental results show that the approach achieves high accuracy compared with conventional data-mining techniques while preserving the privacy of datasets.
Jimmy Ming-Tai Wu, Qian Teng, Md. Shamsul Huda, Yeh-Cheng Chen, Chien-Ming Chen 0001
ACM Trans. Sens. Networks1
2022 Analytics of high average-utility patterns in the industrial internet of things
abstract
Abstract Recently, revealing more valuable information except for quantity value for a database is an essential research field. High utility itemset mining (HAUIM) was suggested to reveal useful patterns by average-utility measure for pattern analytics and evaluations. HAUIM provides a more fair assessment than generic high utility itemset mining and ignores the influence of the length of itemsets. There are several high-performance HAUIM algorithms proposed to gain knowledge from a disorganized database. However, most existing works do not concern the uncertainty factor, which is one of the characteristics of data gathered from IoT equipment. In this work, an efficient algorithm for HAUIM to handle the uncertainty databases in IoTs is presented. Two upper-bound values are estimated to early diminish the search space for discovering meaningful patterns that greatly solve the limitations of pattern mining in IoTs. Experimental results showed several evaluations of the proposed approach compared to the existing algorithms, and the results are acceptable to state that the designed approach efficiently reveals high average utility itemsets from an uncertain situation.
Jimmy Ming-Tai Wu, Zhongcui Li, Gautam Srivastava 0001, Unil Yun, Jerry Chun-Wei Lin
Appl. Intell.1
2022 Dynamic maintenance model for high average-utility pattern mining with deletion operation
abstract
Abstract The high average-utility itemset mining (HAUIM) was established to provide a fair measure instead of genetic high-utility itemset mining (HUIM) for revealing the satisfied and interesting patterns. In practical applications, the database is dynamically changed when insertion/deletion operations are performed on databases. Several works were designed to handle the insertion process but fewer studies focused on processing the deletion process for knowledge maintenance. In this paper, we then develop a PRE-HAUI-DEL algorithm that utilizes the pre-large concept on HAUIM for handling transaction deletion in the dynamic databases. The pre-large concept is served as the buffer on HAUIM that reduces the number of database scans while the database is updated particularly in transaction deletion. Two upper-bound values are also established here to reduce the unpromising candidates early which can speed up the computational cost. From the experimental results, the designed PRE-HAUI-DEL algorithm is well performed compared to the Apriori-like model in terms of runtime, memory, and scalability in dynamic databases.
Jimmy Ming-Tai Wu, Qian Teng, Shahab Tayeb, Jerry Chun-Wei Lin
Appl. Intell.1
2022 Top-k dominating queries on incomplete large dataset
Jimmy Ming-Tai Wu, Mu-En Wu, Shahab Tayeb
J. Supercomput.1
2021 Detection of Trajectory Outliers in Intelligent Transportation Systems
abstract
In this paper, we provide a technique for identifying outliers based on embedding trajectory deviation points and deep clustering. We begin by constructing the network topology and the neighbors of the nodes to create a structural embedding while capturing the interactions of the nodes. We then develop a strategy to determine the hidden representation of distraction points in the road network topology. To create a collection of sequences from a hierarchical multilayer network, a biased random walk is used. This sequence is used to fine tune the embedding of the nodes. The trip embedding was then determined by averaging the node embedding values. Finally, the embeddings are clustered using an LSTM-based pairwise classification strategy based on similarity metrics. The experimental results show that compared to the generic techniques Node2Vec and Struct2Vec, the proposed embedding learning trajectory captures the structural identity and improves the F-measure by 5.06% and 2.4%, respectively.
Usman Ahmed, Jerry Chun-Wei Lin, Gautam Srivastava 0001, Youcef Djenouri, Jimmy Ming-Tai Wu
IEEE BigData5
2021 A ML-Based Stock Trading Model for Profit Predication
Jimmy Ming-Tai Wu, Lingyun Sun, Gautam Srivastava 0001, Jerry Chun-Wei Lin
IEA/AIE (2)1
2021 Linguistic frequent pattern mining using a compressed structure
Jerry Chun-Wei Lin, Usman Ahmed, Gautam Srivastava 0001, Jimmy Ming-Tai Wu, Tzung-Pei Hong, Youcef Djenouri
Appl. Intell.4
2021 Hiding sensitive information in eHealth datasets
Jimmy Ming-Tai Wu, Gautam Srivastava 0001, Alireza Jolfaei, Philippe Fournier-Viger, Jerry Chun-Wei Lin
Future Gener. Comput. Syst.1
2021 Fuzzy high-utility pattern mining in parallel and distributed Hadoop framework
abstract
Over the past decade, high-utility itemset mining (HUIM) has received widespread attention that can emphasize more critical information than was previously possible using frequent itemset mining (FIM). Unfortunately, HUIM is very similar to FIM since the methodology determines itemsets using a binary model based on a pre-defined minimum utility threshold. Additionally, most previous works only focused on single, small datasets in HUIM, which is not realistic to any real-world scenarios today containing big data environments. In this work, the fuzzy-set theory and a MapReduce framework are both utilized to design a novel high fuzzy utility pattern mining algorithm to resolve the above issues. Fuzzy-set theory is first involved and a new algorithm called efficient high fuzzy utility itemset mining (EFUPM) is designed to discover high fuzzy utility patterns from a single machine. Two upper-bounds are then estimated to allow early pruning of unpromising candidates in the search space. To handle the large-scale of big datasets, a Hadoop-based high fuzzy utility pattern mining (HFUPM) algorithm is then developed to discover high fuzzy utility patterns based on the Hadoop framework. Experimental results clearly show that the proposed algorithms perform strongly to mine the required high fuzzy utility patterns whether in a single machine or a large-scale environment compared to the current state-of-the-art approaches.
Jimmy Ming-Tai Wu, Gautam Srivastava 0001, Unil Yun, Jerry Chun-Wei Lin
Inf. Sci.1
2021 Mining of High-Utility Patterns in Big IoT-based Databases
Jimmy Ming-Tai Wu, Gautam Srivastava 0001, Jerry Chun-Wei Lin, Youcef Djenouri, Reza M. Parizi, Mohammad S. Khan
Mob. Networks Appl.1
2021 A Provably Secure Authentication and Key Agreement Protocol in Cloud-Based Smart Healthcare Environments
abstract
The wide applications of the Internet of Things and cloud computing technologies have driven the development of many industries. With the improvement of living standards, health has become the top priority of people’s attention. The emergence of the wireless body area network (WBAN) enables people to master their physical condition all the time and make it more convenient for patients and doctors to communicate with each other. Doctors can provide real-time online treatment with cloud-based smart healthcare environments for patients. In this process, patients, health records, and doctors need to maintain security and privacy. Recently, Kumari et al. proposed a secure framework for the smart medical system. However, we found that their framework cannot provide the anonymity of patients and doctors, data confidentiality, and patient unlinkability and also is subject to impersonation attacks and desynchronization attacks. In order to ensure the security and privacy of patients and doctors, we propose an authentication and key exchange protocol in cloud-based smart healthcare environments. Formal and informal security analyses, as well as performance analysis, demonstrated that our protocol is suitable for these environments.
Tsu-Yang Wu, Lei Yang 0055, Jia-Ning Luo, Jimmy Ming-Tai Wu
Secur. Commun. Networks4
2021 A graph-based convolutional neural network stock price prediction with leading indicators
abstract
Abstract The stock market is a capitalistic haven where the issued shares are transferred, traded, and circulated. It bases stock prices on the issue market, however, the structure and trading activities of the stock market are much more complicated than the issue market itself. Therefore, making an accurate prediction becomes an intricate as well as highly difficult task. On the other hand, because of the potential benefits of stock prediction, it attracts generation after generation of scholars as well as investors to continuously develop various prediction methods from different perspectives, a myriad of theories, a multitude of investment strategies, and different practical experiences. In this article, aiming at the task of time series (financial) feature extraction and prediction of price movements, a new convolutional novel neural network that can be called a framework to improve the prediction accuracy of stock trading is proposed. The method that is proposed is called SSACNN, a short form of stock sequence array convolutional neural network. SSACNN collects data including historical data of prices and its leading indicators (options/futures) for a stock to take an array as the input graph of the convolutional neural network framework. In our experimental results, five Taiwanese and American stocks were used as a benchmark to compare with the previous algorithms and proposed algorithm, the motion prediction performance of SSACNN has been improved significantly and proved that it has the potential to be applied in the real financial market.
Jimmy Ming-Tai Wu, Zhongcui Li, Gautam Srivastava 0001, Meng-Hsiun Tsai, Jerry Chun-Wei Lin
Softw. Pract. Exp.1
2021 A Multi-Threshold Ant Colony System-based Sanitization Model in Shared Medical Environments
abstract
During the past several years, revealing some useful knowledge or protecting individual’s private information in an identifiable health dataset (i.e., within an Electronic Health Record) has become a tradeoff issue. Especially in this era of a global pandemic, security and privacy are often overlooked in lieu of usability. Privacy preserving data mining (PPDM) is definitely going to be have an important role to resolve this problem. Nevertheless, the scenario of mining information in an identifiable health dataset holds high complexity compared to traditional PPDM problems. Leaking individual private information in an identifiable health dataset has becomes a serious legal issue. In this article, the proposed Ant Colony System to Data Mining algorithm takes the multi-threshold constraint to secure and sanitize patents’ records in different lengths, which is applicable in a real medical situation. The experimental results show the proposed algorithm not only has the ability to hide all sensitive information but also to keep useful knowledge for mining usage in the sanitized database.
Jimmy Ming-Tai Wu, Gautam Srivastava 0001, Jerry Chun-Wei Lin, Qian Teng
ACM Trans. Internet Techn.1
2021 The Efficient Mining of Skyline Patterns from a Volunteer Computing Network
abstract
In the ever-growing world, the concepts of High-utility Itemset Mining (HUIM) as well as Frequent Itemset Mining (FIM) are fundamental works in knowledge discovery. Several algorithms have been designed successfully. However, these algorithms only used one factor to estimate an itemset. In the past, skyline pattern mining by considering both aspects of frequency and utility has been extensively discussed. In most cases, however, people tend to focus on purchase quantities of itemsets rather than frequencies. In this article, we propose a new knowledge called skyline quantity-utility pattern (SQUP) to provide better estimations in the decision-making process by considering quantity and utility together. Two algorithms, respectively, called SQU-Miner and SKYQUP are presented to efficiently mine the set of SQUPs. Moreover, the usage of volunteer computing is proposed to show the potential in real supermarket applications. Two new efficient utility-max structures are also mentioned for the reduction of the candidate itemsets, respectively, utilized in SQU-Miner and SKYQUP. These two new utility-max structures are used to store the upper-bound of utility for itemsets under the quantity constraint instead of frequency constraint, and the second proposed utility-max structure moreover applies a recursive updated process to further obtain strict upper-bound of utility. Our in-depth experimental results prove that SKYQUP has stronger performance when a comparison is made to SQU-Miner in terms of memory usage, runtime, and the number of candidates.
Jimmy Ming-Tai Wu, Qian Teng, Gautam Srivastava 0001, Matin Pirouz, Jerry Chun-Wei Lin
ACM Trans. Internet Techn.1
2020 Fuzzy High-Utility Pattern Mining based on the Hadoop Framework
abstract
In this paper, fuzzy-set theory is first used and a new algorithm called efficient fuzzy high-utility itemset mining (EFUPM) algorithm is designed to discover the fuzzy high-utility patterns from a single machine. Two upper-bounds are then estimated to early prune the unpromising candidates in the search space. To handle the large-scale of big datasets, the Hadoop-based fuzzy high-utility pattern mining (HFUPM) algorithm is then developed to discover the fuzzy high-utility patterns based on the Hadoop framework. Experimental results show that the proposed algorithms can perform well to mine the required fuzzy high-utility patterns whether in a single machine or a large-scale environment compared to the state-of-the-art approaches.
Jimmy Ming-Tai Wu, Gautam Srivastava 0001, Jerry Chun-Wei Lin
IEEE BigData1
2020 High-Utility Pattern Mining in Hadoop Environments
abstract
In this article, we present an Efficient High Utility Pattern Mining framework to mine high-utility patterns with a reasonable pruning strategy to speed up the mining performance. Concurrently, for solving the problem of excessive data volume in the current era, we applied the developed framework to the MapReduce architecture used for improving the feasibility in practical applications. Our in-depth work in this paper culminates with some experimental results that clearly show that our proposed framework can perform well to mine the required pattern in a big-data dataset and shows great performance in a Hadoop computing cluster.
Jimmy Ming-Tai Wu, Gautam Srivastava 0001, Jerry Chun-Wei Lin
IEEE BigData1
2020 Mining Multiple Fuzzy Frequent Patterns with Compressed List Structures
abstract
Fuzzy-set theory was invented to represent more meaningful representations of knowledge for human reasoning, which can also be applied and utilized for handling the quantitative database. In this paper, an efficient fuzzy mining (EFM) algorithm is presented to fast discover the multiple fuzzy frequent patterns from quantitative databases under type-2 fuzzy-set theory. A compressed fuzzy-list (CFL)-structure is developed to maintain complete information for rule generation. Two pruning techniques are developed to reduce the search space and speed up mining progress. Several experiments are carried out for the purpose of verifying the efficiency and effectiveness of the designed approach in terms of runtime and the number of examined nodes under different minimum support thresholds and the results indicated the designed EFM achieves the best performance compared to the existing models.
Jerry Chun-Wei Lin, Jimmy Ming-Tai Wu, Youcef Djenouri, Gautam Srivastava 0001, Tzung-Pei Hong
FUZZ-IEEE2
2020 Efficient Mining of Pareto-Front High Expected Utility Patterns
Usman Ahmed, Jerry Chun-Wei Lin, Jimmy Ming-Tai Wu, Youcef Djenouri, Gautam Srivastava 0001, Suresh Kumar Mukhiya
IEA/AIE3
2020 Maintenance of Prelarge High Average-Utility Patterns in Incremental Databases
Jimmy Ming-Tai Wu, Qian Teng, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Chien-Fu Cheng
IEA/AIE1
2019 A GA-based Framework for Mining High Fuzzy Utility Itemsets
abstract
Comparing to frequent itemset mining (FIM), utility-pattern mining receives increasing attention in the field of data mining recently. With the flourishing development of utility-pattern mining, most studies focused on the efficiency problem by considering the efficient data structure to compress the original data and pruning strategies to reduce the search space for knowledge discovery. However, those approaches can only handle the binary situation, thus the discovered knowledge cannot be represented as the linguistic variables. Previous works have addressed this problem by introducing the generic approaches to find the high fuzzy utility itemsets in a small database. In real-world situations, the dataset may be very large, and it is costly to mine all the required information from a very large database. In this paper, we first present a HFUI-GA framework to discover the high fuzzy utility itemsets in a limited time. Several improvement strategies are also proposed to speed up the evolutionary progress. Experiments are then conducted to show the performance of the variants of the designed HFUI-GA framework in terms of number of the discovered high fuzzy utility itemsets (HFUIs) and the results are convincing to show that the designed GA-based HFUI-GA framework is a promising solution to mine for HFUIs.
Jimmy Ming-Tai Wu, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Tomasz Wiktorski, Tzung-Pei Hong, Matin Pirouz
IEEE BigData1
2019 A Swarm-based Data Sanitization Algorithm in Privacy-Preserving Data Mining
abstract
In recent decades, data protection (PPDM), which not only hides information, but also provides information that is useful to make decisions, has become a critical concern. We present a sanitization algorithm with the consideration of four side effects based on multi-objective PSO and hierarchical clustering methods to find optimized solutions for PPDM. Experiments showed that compared to existing approaches, the designed sanitization algorithm based on the hierarchical clustering method achieves satisfactory performance in terms of hiding failure, missing cost, and artificial cost.
Jimmy Ming-Tai Wu, Jerry Chun-Wei Lin, Youcef Djenouri, Philippe Fournier-Viger, Yuyu Zhang
CEC1
2019 Efficient Mining of High Average-Utility Sequential Patterns from Uncertain Databases
abstract
In this paper, we address the limitation for mining of high-utility sequential-pattern mining from uncertain databases and present a probabilistic high average-utility sequential pattern mining framework for discovering the set of probabilistic high average-utility sequential patterns from uncertain databases. A level-wise algorithm and three pruning strategies are introduced to mine the set of the desired patterns. Several experiments are then evaluated to show that the proposed algorithm achieves promising performance.
Jerry Chun-Wei Lin, Jimmy Ming-Tai Wu, Philippe Fournier-Viger, Tzung-Pei Hong, Ting Li 0011
SMC2
2019 High-Utility Itemset Mining with Effective Pruning Strategies
abstract
High-utility itemset mining is a popular data mining problem that considers utility factors, such as quantity and unit profit of items besides frequency measure from the transactional database. It helps to find the most valuable and profitable products/items that are difficult to track by using only the frequent itemsets. An item might have a high-profit value which is rare in the transactional database and has a tremendous importance. While there are many existing algorithms to find high-utility itemsets (HUIs) that generate comparatively large candidate sets, our main focus is on significantly reducing the computation time with the introduction of new pruning strategies. The designed pruning strategies help to reduce the visitation of unnecessary nodes in the search space, which reduces the time required by the algorithm. In this article, two new stricter upper bounds are designed to reduce the computation time by refraining from visiting unnecessary nodes of an itemset. Thus, the search space of the potential HUIs can be greatly reduced, and the mining procedure of the execution time can be improved. The proposed strategies can also significantly minimize the transaction database generated on each node. Experimental results showed that the designed algorithm with two pruning strategies outperform the state-of-the-art algorithms for mining the required HUIs in terms of runtime and number of revised candidates. The memory usage of the designed algorithm also outperforms the state-of-the-art approach. Moreover, a multi-thread concept is also discussed to further handle the problem of big datasets.
Jimmy Ming-Tai Wu, Jerry Chun-Wei Lin, Ashish Tamrakar
ACM Trans. Knowl. Discov. Data1
2017 Extracting recent weighted-based patterns from uncertain temporal databases
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao, Jimmy Ming-Tai Wu, Justin Zhijun Zhan
Eng. Appl. Artif. Intell.5
2017 An ACO-based approach to mine high-utility itemsets
Jimmy Ming-Tai Wu, Justin Zhijun Zhan, Jerry Chun-Wei Lin
Knowl. Based Syst.1
2016 Mining high-utility itemsets based on particle swarm optimization
Jerry Chun-Wei Lin, Philippe Fournier-Viger, Jimmy Ming-Tai Wu, Tzung-Pei Hong, Shyue-Liang Wang, Justin Zhijun Zhan
Eng. Appl. Artif. Intell.4