EDBT 2026 Demo / reviewers in the wild / expert
Jiahui Chen 0002
dblp:122/5100-2
· DBLP profile ↗
11ranked-venue papers in the field
1as first author
9since 2021 · last 2023
0000-0001-7128-9778ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 9 (1 first)Database Systems & Data Management · 1Data Mining & Knowledge Discovery · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Mining Rare Utility Patterns within Target ItemsabstractAs a crucial subfield of pattern discovery, high utility rare itemset mining (HURIM) is developed to discover abnormal but significant patterns. HURIM plays a vital role in various scenarios, such as network security, disease detection, and biomedicine. However, the traditional HURIM algorithms ignore the users’ demands, which generates massive needless patterns. In general, target-based HURIM algorithms can discover more useful information that meets the needs of users than traditional HURIM algorithms. To this end, we propose a targeted HURIM algorithm called Mining Rare Utility Patterns within Target Items (TIRUP). TIRUP adopts two techniques (projection and merging technologies), to diminish the consumption of database scanning. To effectively improve the performance of TIRUP, this paper utilizes several strategies based on frequency, utility, and target factors. Finally, a series of experiments are conducted to demonstrate the efficiency of the proposed TIRUP algorithm, and the experimental results indicate that TIRUP is suitable for processing large-scale and dense datasets. Cuiwei Peng, Jiahui Chen 0002, Wensheng Gan, Shicheng Wan |
IEEE Big Data | 3 |
| 2023 | Anomaly Rule Detection in Sequence DataabstractAnalyzing sequence data usually leads to the discovery of interesting patterns and then anomaly detection. In recent years, numerous frameworks and methods have been proposed to discover interesting patterns in sequence data as well as detect anomalous behavior. However, existing algorithms mainly focus on frequency-driven analytics, and they are challenging to be applied in real-world settings. In this work, we present a new anomaly detection framework called DUOS that enables Discovery of Utility-aware Outlier Sequential rules from a set of sequences. In this pattern-based anomaly detection algorithm, we incorporate both the anomalousness and utility of a group, and then introduce the concept of utility-aware outlier sequential rule (UOSR). We show that this is a more meaningful way for detecting anomalies. Besides, we propose some efficient pruning strategies w.r.t. upper bounds for mining UOSR, as well as the outlier detection. An extensive experimental study conducted on several real-world datasets shows that the proposed DUOS algorithm has a better effectiveness and efficiency. Finally, DUOS outperforms the baseline algorithm and has a suitable scalability. Wensheng Gan, Shicheng Wan, Jiahui Chen 0002, Chien-Ming Chen 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2022 | Flexibly Mining Better PatternsabstractCorrelated high-utility pattern mining (CoUPM) considers the correlation between items in a pattern and offers a more reliable analysis for users. In real-world applications, the discovered patterns from CoUPM can present more interpretable information, but not all of them are useful. Generally, users pay attention to the number of items a pattern contains, which allows them to make reasonable decisions. In this paper, we solve the problem of mining those correlated high-utility patterns whose length is specified. A utility-list-based algorithm called Flexible Correlated Utility-based Pattern (FCoUP) is proposed. Furthermore, we propose some pruning strategies with the designed upper bounds for the two evaluation metrics: correlation and utility, reducing unwanted patterns generated and nodes visited during the mining process. Experiments show that FCoUP variants can produce more intelligent and flexible correlated high-utility patterns on a variety of datasets. Gengsen Huang, Wensheng Gan, Long Li 0005, Tianlong Gu, Jiahui Chen 0002 |
IEEE Big Data | 5 |
| 2022 | Metaverse in Education: Vision, Opportunities, and ChallengesabstractTraditional education has been updated with the development of information technology in human history. Within big data and cyber-physical systems, the Metaverse has generated strong interest in various applications (e.g., entertainment, business, and cultural travel) over the last decade. As a novel social work idea, the Metaverse consists of many kinds of technologies, e.g., big data, interaction, artificial intelligence, game design, Internet computing, Internet of Things, and blockchain. It is foreseeable that the usage of Metaverse will contribute to educational development. However, the architectures of the Metaverse in education are not yet mature enough. There are many questions we should address for the Metaverse in education. To this end, this paper aims to provide a systematic literature review of Metaverse in education. This paper is a comprehensive survey of the Metaverse in education, with a focus on current technologies, challenges, opportunities, and future directions. First, we present a brief overview of the Metaverse in education, as well as the motivation behind its integration. Then, we survey some important characteristics for the Metaverse in education, including the personal teaching environment and the personal learning environment. Next, we envisage what variations of this combination will bring to education in the future and discuss their strengths and weaknesses. We also review the state-of-the-art case studies (including technical companies and educational institutions) for Metaverse in education. Finally, we point out several challenges and issues in this promising area. Shicheng Wan, Wensheng Gan, Jiahui Chen 0002, Han-Chieh Chao |
IEEE Big Data | 4 |
| 2022 | Targeted Mining of Rare High-Utility PatternsabstractPattern discovery has been widely studied and applied as a classical problem in data mining. As a subfield of itemset mining, identifying high-utility rare itemsets (HURI) can find abnormal but vital patterns in transaction databases. It plays a unique role in real-world scenarios such as anomaly detection and disease detection. However, with large-scale databases, the final results are often massive according to the user-specified threshold. In other words, the mining algorithm ignores the user’s subjective interests and lacks interaction during the mining process. A pattern discovery algorithm may output many useless or uninteresting patterns. To this end, in this paper, we define the problem of mining targeted HURIs and propose a list-based algorithm called TaRP for effectively solving this issue. In addition, based on preliminary research, we propose several effective pruning strategies for improving the algorithm’s performance. TaRP makes the results more interactive and specific by incorporating the user’s prior knowledge during mining. It also has a natural performance advantage with the help of effective strategies. We also evaluated the proposed algorithm on several real-life datasets. The extensive experimental results demonstrate that TaRP not only correctly solves the problem but also has advantages in runtime and memory consumption, especially on dense datasets. Peifeng Zhang, Jiahui Chen 0002, Shicheng Wan, Wensheng Gan |
IEEE Big Data | 2 |
| 2022 | Fast Mining RFM Patterns for Behavioral AnalyticsabstractIn recent years, the problem of high-utility itemset mining (HUIM) has been extensively studied. However, HUIM algorithms only reveal profitable but generalized itemsets from transaction databases. In the market analysis domain, these mining results just reflect the sales trend of all customers and are not sufficient for making market strategies. In other words, it is hard to maintain specific customers for a long time due to the limitations of HUIM analysis of customer behaviors. In this paper, a novel data mining algorithm called RFM-Miner is proposed to discover RFM-patterns that are highly recent, frequent, and profitable in transaction databases. The novel algorithm relies on the array-bin structure to fast calculate adopted upper-bounds (i.e., transaction-weighted utilization, subtree and local utility) in linear time and space. In addition, RFM-Miner always searches for extension items of an itemset in a small projected database. And the merging technique is utilized to reduce the size of the search space. An extensive experimental study on four datasets (including real-life and synthetic) shows that RFM-Miner performs very well in terms of runtime and memory consumption. The novel algorithm also achieves better performance than the state-of-the-art benchmarks, especially on dense datasets. Shicheng Wan, Jieying Deng, Wensheng Gan, Jiahui Chen 0002, Philip S. Yu |
DSAA | 4 |
| 2021 | NSPIS: Mining Negative Sequential Patterns with Individual SupportabstractNegative sequential pattern (NSP) mining is crucial and sometimes carries more enlightening information than positive sequential pattern (PSP) mining in data mining. Owing to its computational complexity and exponential search space, the task of discovering NSPs is often much more difficult and challenging than that for PSPs. To date, a few NSP mining algorithms have been proposed. However, most algorithms only consider a single support, thus can not present good results in many special real-world applications. To solve this problem and achieve better efficiency on a long sequence database or a large-scale database, we propose a novel algorithm called Negative Sequential Patterns with Individual Support (NSPIS) in this paper. The projection mechanism is adopted to NSPIS, which allows greatly reduce the search space and simultaneously improve the efficiency. Finally, detailed results of the experiments show that NSPIS can achieve better performance and it uses less memory on large datasets compared to the state-of-the-art algorithm. Gengsen Huang, Wensheng Gan, Shan Huang 0009, Jiahui Chen 0002, Chien-Ming Chen 0001 |
IEEE BigData | 4 |
| 2021 | Joint Utility and Frequency for Pattern ClassificationabstractHigh-frequency itemset mining (HFIM) and high-utility itemset mining (HUIM) aim to discover itemsets with high occurrence and high utility, respectively, in a transaction database. A number of efficient algorithms have been developed to identify these high-utility itemsets (HUIs) or high-frequency itemsets (HFIs). Such algorithms play an increasingly important role in many occasions especially for analysis in commercial enterprises. In this paper, we propose a new model called joint utility and frequency for pattern classification, and two new algorithms, namely UFCgenand UFCfast. Both algorithms are designed to categorize each itemset into different type of patterns by setting the minimum thresholds of utility and frequency. We compare these algorithms on two datasets. The experimental results show that both algorithms can successfully collect three different types of itemsets from all candidate itemsets based on frequency and utility, and the list-based UFCfastalgorithm outperforms the level-wise-based UFCgenalgorithm in terms of execution time. Wensheng Gan, Yongdong Wu, Jiahui Chen 0002, Chien-Ming Chen 0001 |
IEEE BigData | 4 |
| 2021 | Targeted High-Utility Itemset QueryingabstractTraditional high-utility itemset mining (HUIM) aims to determine all high-utility itemsets (HUIs) that satisfy the minimum utility threshold in transaction databases. However, in most applications, not all HUIs are interesting because only specific parts are required. Thus, targeted mining based on user preferences is more important than traditional mining tasks. This paper is the first to propose a targeted HUIM problem and to provide a clear formulation of the targeted utility mining task in a quantitative transaction database. A tree-based algorithm known as Target-based high-Utility iteMset querying using (TargetUM) is proposed. The algorithm uses a lexicographic querying tree and three effective pruning strategies to improve the mining efficiency. We implemented experimental validation on several real and synthetic databases, and the results demonstrate that the performance of TargetUM is satisfactory, complete, and correct. Finally, owing to the lexicographic querying tree, the database no longer needs to be scanned repeatedly for multiple queries. Jinbao Miao, Shicheng Wan, Wensheng Gan, Jiayi Sun 0002, Jiahui Chen 0002 |
IEEE BigData | 5 |
| 2020 | OSUMI: On-Shelf Utility Mining from Itemset-based DataabstractAs an important technique for dealing with transactional database in the field of data mining, high-utility itemset mining (HUIM) can be used to discover itemsets which have a high utility. However, it has a bias when towarding the item combinations which have more exhibition period since they have more opportunity to generate a high utility. To address this, the on-shelf time period of items need to be considered, thus on-shelf utility mining (OSUM) can be applied in the application which is more closer to the actual situation. Currently several models have been proposed to deal with the OSUM problem, but they still suffer from the requirement that it needs to maintain a massive candidates in memory and to scan database many times. In this paper, we propose an effective algorithm named OSUMI (On-Shelf Utility Mining from Itemset-based data) which can discover the on-shelf itemsets with high utility in a more practical way. More precisely, in order to avoid the problems of high memory consumption, OSUMI applies some properties of on-shelf utility. Besides, two upper-bounds named subtree utility and local utility are applied to prune the search space. Finally, an extensive experimental study on two real on-shelf datasets shows that our proposed algorithm can be significantly faster than the state-of-the-art algorithm for this mining task. Jiahui Chen 0002, Xu Guo 0003, Wensheng Gan, Chien-Ming Chen 0001, Weiping Ding 0001, Guoting Chen |
IEEE BigData | 1 |
| 2020 | TopHUI: Top-k high-utility itemset mining with negative utilityabstractIn the field of data science, utility-driven data mining has become an emergent intelligent technique with wide applications. The existing utility mining algorithms usually discover all the patterns satisfying a given minimum utility threshold. However, a huge number of return results is not intuitive, not interpretable, and not easy for users to understand. Besides, it is often difficult and time-consuming for users to set a proper minimum utility threshold that is quite sensitive to the mining results. To address these issues, the problem of top-k high-utility itemset mining has been studied. In this paper, we present an efficient algorithm (named TopHUI) for finding top-k high-utility itemsets from transactional database that contains both positive and negative utility. This algorithm utilizes the positive-and-negative utility-list (PNU-list) to store the compress information, including positive, negative, and remaining utility. Besides, several threshold raising strategies and pruning strategies are proposed to prune the search space. Finally, some extensive experiments were conducted to evaluate the performance of the proposed TopHUI algorithm on both real-life and synthetic datasets, particularly in terms of effectiveness and efficiency. Wensheng Gan, Shicheng Wan, Jiahui Chen 0002, Chien-Ming Chen 0001, Lina Qiu |
IEEE BigData | 3 |