EDBT 2026 Demo / reviewers in the wild / expert
Shicheng Wan
dblp:289/2836
· DBLP profile ↗
10ranked-venue papers in the field
1as first author
9since 2021 · last 2024
0000-0002-1051-9426ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 8Database Systems & Data Management · 1Data Mining & Knowledge Discovery · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | RFMI-based Customer Segmentation with K-meansabstractThe development of e-marketing over recent decades has led offline and online retail enterprises to adopt various data analysis technologies to enhance their understanding of consumer behavior and increase revenue. One common approach involves segmenting consumers into distinct groups based on designed metrics, targeting high-value segments for specialized services. To evaluate customer worthiness, the popular RFM model uses three dimensions: recency (the time since their last purchase), frequency (how often they make purchases), and monetary value (total spending). Higher scores under this model are indicative of greater potential profitability for businesses. While this approach provides valuable insights, it may not fully capture all profitable customer behaviors accurately. To address these limitations, this paper introduces a new model, namely the RFMI (i.e., recency, frequency, monetary, and interval) model, for comprehensively evaluating customer value. The new model employs an analytic hierarchy process to derive the RFMI values of customers. Subsequently, we employ K-means clustering customers to group customers into six segments. Moreover, the experimental dataset was sourced from a real UK e-commerce platform. The experimental results indicate that the new model effectively distinguishes between various consumption patterns among customers. This enhanced understanding can enable retailers to improve their marketing strategies more precisely, optimize customer service, and increase profitability. Wensheng Gan, Pinlyu Zhou, Shicheng Wan, Jiyuan Zeng, Zhenlian Qi |
IEEE Big Data | 3 |
| 2023 | Model-as-a-Service (MaaS): A SurveyabstractDue to the increased number of parameters and data in the pre-trained model exceeding a certain level, a foundation model (e.g., a large language model) can significantly improve downstream task performance and emerge with some novel special abilities (e.g., deep learning, complex reasoning, and human alignment) that were not present before. Foundation models are a form of generative artificial intelligence (GenAI), and Model-as-a-Service (MaaS) has emerged as a groundbreaking paradigm that revolutionizes the deployment and utilization of GenAI models. MaaS represents a paradigm shift in how we use AI technologies and provides a scalable and accessible solution for developers and users to leverage pre-trained AI models without the need for extensive infrastructure or expertise in model training. In this paper, the introduction aims to provide a comprehensive overview of MaaS, its significance, and its implications for various industries. We provide a brief review of the development history of “X-as-a-Service” based on cloud computing and present the key technologies involved in MaaS. The development of GenAI models will become more democratized and flourish. We also review recent application studies of MaaS. Finally, we highlight several challenges and future issues in this promising area. MaaS is a new deployment and service paradigm for different AI-based models. We hope this review will inspire future research in the field of MaaS. Wensheng Gan, Shicheng Wan, Philip S. Yu |
IEEE Big Data | 2 |
| 2023 | Multimodal Large Language Models: A SurveyabstractThe exploration of multimodal language models integrates multiple data types, such as images, text, language, audio, and other heterogeneity. While the latest large language models excel in text-based tasks, they often struggle to understand and process other data types. Multimodal models address this limitation by combining various modalities, enabling a more comprehensive understanding of diverse data. This paper begins by defining the concept of multimodal and examining the historical development of multimodal algorithms. Furthermore, we introduce a range of multimodal products, focusing on the efforts of major technology companies. A practical guide is provided, offering insights into the technical aspects of multimodal models. Moreover, we present a compilation of the latest algorithms and commonly used datasets, providing researchers with valuable resources for experimentation and evaluation. Lastly, we explore the applications of multimodal models and discuss the challenges associated with their development. By addressing these aspects, this paper aims to facilitate a deeper understanding of multimodal models and their potentiality in various domains. Jiayang Wu 0001, Wensheng Gan, Shicheng Wan, Philip S. Yu |
IEEE Big Data | 4 |
| 2023 | Mining Rare Utility Patterns within Target ItemsabstractAs a crucial subfield of pattern discovery, high utility rare itemset mining (HURIM) is developed to discover abnormal but significant patterns. HURIM plays a vital role in various scenarios, such as network security, disease detection, and biomedicine. However, the traditional HURIM algorithms ignore the users’ demands, which generates massive needless patterns. In general, target-based HURIM algorithms can discover more useful information that meets the needs of users than traditional HURIM algorithms. To this end, we propose a targeted HURIM algorithm called Mining Rare Utility Patterns within Target Items (TIRUP). TIRUP adopts two techniques (projection and merging technologies), to diminish the consumption of database scanning. To effectively improve the performance of TIRUP, this paper utilizes several strategies based on frequency, utility, and target factors. Finally, a series of experiments are conducted to demonstrate the efficiency of the proposed TIRUP algorithm, and the experimental results indicate that TIRUP is suitable for processing large-scale and dense datasets. Cuiwei Peng, Jiahui Chen 0002, Wensheng Gan, Shicheng Wan |
IEEE Big Data | 5 |
| 2023 | Anomaly Rule Detection in Sequence DataabstractAnalyzing sequence data usually leads to the discovery of interesting patterns and then anomaly detection. In recent years, numerous frameworks and methods have been proposed to discover interesting patterns in sequence data as well as detect anomalous behavior. However, existing algorithms mainly focus on frequency-driven analytics, and they are challenging to be applied in real-world settings. In this work, we present a new anomaly detection framework called DUOS that enables Discovery of Utility-aware Outlier Sequential rules from a set of sequences. In this pattern-based anomaly detection algorithm, we incorporate both the anomalousness and utility of a group, and then introduce the concept of utility-aware outlier sequential rule (UOSR). We show that this is a more meaningful way for detecting anomalies. Besides, we propose some efficient pruning strategies w.r.t. upper bounds for mining UOSR, as well as the outlier detection. An extensive experimental study conducted on several real-world datasets shows that the proposed DUOS algorithm has a better effectiveness and efficiency. Finally, DUOS outperforms the baseline algorithm and has a suitable scalability. Wensheng Gan, Shicheng Wan, Jiahui Chen 0002, Chien-Ming Chen 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Metaverse in Education: Vision, Opportunities, and ChallengesabstractTraditional education has been updated with the development of information technology in human history. Within big data and cyber-physical systems, the Metaverse has generated strong interest in various applications (e.g., entertainment, business, and cultural travel) over the last decade. As a novel social work idea, the Metaverse consists of many kinds of technologies, e.g., big data, interaction, artificial intelligence, game design, Internet computing, Internet of Things, and blockchain. It is foreseeable that the usage of Metaverse will contribute to educational development. However, the architectures of the Metaverse in education are not yet mature enough. There are many questions we should address for the Metaverse in education. To this end, this paper aims to provide a systematic literature review of Metaverse in education. This paper is a comprehensive survey of the Metaverse in education, with a focus on current technologies, challenges, opportunities, and future directions. First, we present a brief overview of the Metaverse in education, as well as the motivation behind its integration. Then, we survey some important characteristics for the Metaverse in education, including the personal teaching environment and the personal learning environment. Next, we envisage what variations of this combination will bring to education in the future and discuss their strengths and weaknesses. We also review the state-of-the-art case studies (including technical companies and educational institutions) for Metaverse in education. Finally, we point out several challenges and issues in this promising area. Shicheng Wan, Wensheng Gan, Jiahui Chen 0002, Han-Chieh Chao |
IEEE Big Data | 2 |
| 2022 | Targeted Mining of Rare High-Utility PatternsabstractPattern discovery has been widely studied and applied as a classical problem in data mining. As a subfield of itemset mining, identifying high-utility rare itemsets (HURI) can find abnormal but vital patterns in transaction databases. It plays a unique role in real-world scenarios such as anomaly detection and disease detection. However, with large-scale databases, the final results are often massive according to the user-specified threshold. In other words, the mining algorithm ignores the user’s subjective interests and lacks interaction during the mining process. A pattern discovery algorithm may output many useless or uninteresting patterns. To this end, in this paper, we define the problem of mining targeted HURIs and propose a list-based algorithm called TaRP for effectively solving this issue. In addition, based on preliminary research, we propose several effective pruning strategies for improving the algorithm’s performance. TaRP makes the results more interactive and specific by incorporating the user’s prior knowledge during mining. It also has a natural performance advantage with the help of effective strategies. We also evaluated the proposed algorithm on several real-life datasets. The extensive experimental results demonstrate that TaRP not only correctly solves the problem but also has advantages in runtime and memory consumption, especially on dense datasets. Peifeng Zhang, Jiahui Chen 0002, Shicheng Wan, Wensheng Gan |
IEEE Big Data | 3 |
| 2022 | Fast Mining RFM Patterns for Behavioral AnalyticsabstractIn recent years, the problem of high-utility itemset mining (HUIM) has been extensively studied. However, HUIM algorithms only reveal profitable but generalized itemsets from transaction databases. In the market analysis domain, these mining results just reflect the sales trend of all customers and are not sufficient for making market strategies. In other words, it is hard to maintain specific customers for a long time due to the limitations of HUIM analysis of customer behaviors. In this paper, a novel data mining algorithm called RFM-Miner is proposed to discover RFM-patterns that are highly recent, frequent, and profitable in transaction databases. The novel algorithm relies on the array-bin structure to fast calculate adopted upper-bounds (i.e., transaction-weighted utilization, subtree and local utility) in linear time and space. In addition, RFM-Miner always searches for extension items of an itemset in a small projected database. And the merging technique is utilized to reduce the size of the search space. An extensive experimental study on four datasets (including real-life and synthetic) shows that RFM-Miner performs very well in terms of runtime and memory consumption. The novel algorithm also achieves better performance than the state-of-the-art benchmarks, especially on dense datasets. Shicheng Wan, Jieying Deng, Wensheng Gan, Jiahui Chen 0002, Philip S. Yu |
DSAA | 1 |
| 2021 | Targeted High-Utility Itemset QueryingabstractTraditional high-utility itemset mining (HUIM) aims to determine all high-utility itemsets (HUIs) that satisfy the minimum utility threshold in transaction databases. However, in most applications, not all HUIs are interesting because only specific parts are required. Thus, targeted mining based on user preferences is more important than traditional mining tasks. This paper is the first to propose a targeted HUIM problem and to provide a clear formulation of the targeted utility mining task in a quantitative transaction database. A tree-based algorithm known as Target-based high-Utility iteMset querying using (TargetUM) is proposed. The algorithm uses a lexicographic querying tree and three effective pruning strategies to improve the mining efficiency. We implemented experimental validation on several real and synthetic databases, and the results demonstrate that the performance of TargetUM is satisfactory, complete, and correct. Finally, owing to the lexicographic querying tree, the database no longer needs to be scanned repeatedly for multiple queries. Jinbao Miao, Shicheng Wan, Wensheng Gan, Jiayi Sun 0002, Jiahui Chen 0002 |
IEEE BigData | 2 |
| 2020 | TopHUI: Top-k high-utility itemset mining with negative utilityabstractIn the field of data science, utility-driven data mining has become an emergent intelligent technique with wide applications. The existing utility mining algorithms usually discover all the patterns satisfying a given minimum utility threshold. However, a huge number of return results is not intuitive, not interpretable, and not easy for users to understand. Besides, it is often difficult and time-consuming for users to set a proper minimum utility threshold that is quite sensitive to the mining results. To address these issues, the problem of top-k high-utility itemset mining has been studied. In this paper, we present an efficient algorithm (named TopHUI) for finding top-k high-utility itemsets from transactional database that contains both positive and negative utility. This algorithm utilizes the positive-and-negative utility-list (PNU-list) to store the compress information, including positive, negative, and remaining utility. Besides, several threshold raising strategies and pruning strategies are proposed to prune the search space. Finally, some extensive experiments were conducted to evaluate the performance of the proposed TopHUI algorithm on both real-life and synthetic datasets, particularly in terms of effectiveness and efficiency. Wensheng Gan, Shicheng Wan, Jiahui Chen 0002, Chien-Ming Chen 0001, Lina Qiu |
IEEE BigData | 2 |