EDBT 2026 Demo / reviewers in the wild / expert
Carson K. Leung
dblp:29/654 · also Carson Kai-Sang Leung
· DBLP profile ↗
118ranked-venue papers in the field
54as first author
42since 2021 · last 2026
0000-0002-7541-9127ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 56 (27 first)Database Systems & Data Management · 29 (21 first)Big Data, Cloud & Distributed Data Systems · 17 (4 first)Information Retrieval & Web Search · 7Other / Interdisciplinary · 5 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ANNOYing HDBSCAN for Clustering
Ben M. Clark, Ginelle K. Elias, Carson K. Leung, Adam G. M. Pazdor, Thamira M. Randeniya, Xavier J. Schneider |
DaWaK | 3 |
| 2026 | Q-dEclat: A Vertical Quantitative Frequent Pattern Mining Algorithm
Carson K. Leung, Adam G. M. Pazdor |
DaWaK | 1 |
| 2025 | Machine Learning on Big Data for Crime Trend Analytics and Prediction
M. Akroma E. Ansah, Carson K. Leung, Omodesayo A. Owolabi, Alfredo Cuzzocrea |
IEEE Big Data | 2 |
| 2025 | Machine Learning on Big Economic Data for Predictive Financial Analytics
Carson K. Leung, Thanh Trung Jack Nguyen, Jing Xiang Ong, Alfredo Cuzzocrea |
IEEE Big Data | 2 |
| 2025 | An Enhanced FP-Growth Algorithm with Hybrid Adaptive Support Threshold for Association Rule Mining
Kanda Runapongsa Saikaew, Carson K. Leung, Kritbodin Phiwhorm |
DaWaK | 2 |
| 2025 | A Hybrid Data Model to Support Transportation Analytics of Emergency Service Vehicles
Carson K. Leung |
DEXA (1) | 1 |
| 2024 | A Real-Time Sentiment Feedback System: Binary Categorization and Context Understanding Based on Product Reviews
Arshpreet S. Buttar, Olukoye O. Fatoki, Roba Geleta, Carson K. Leung |
ASONAM (4) | 5 |
| 2024 | Privacy-Preserving Publishing with Generative Adversarial Network (GAN) for Supporting Contact Tracing of Infectious DiseasesabstractGenerative artificial intelligence (AI) has become popular. The combination of increasingly complex datasets beyond human comprehension and the widespread availability of advanced computing systems—such as graphics processing unit (GPU) and tensor processing unit (TPU)—has driven the rapid advancement of generative AI. This technology has found applications in areas such as voice recognition, recommendation systems and data privacy preservation, which foster more data sharing and reuse. While challenges related to bias, fairness and uncertainty in AI continue to evolve, emerging government regulations aim to ensure ethical use and maximize societal benefits. In this paper, we present a system that leverages generative adversarial network (GAN) to enable privacy-preserving data publishing. The system supports contact tracing for infectious diseases like coronavirus disease 2019 (COVID-19) and monkey-pox. Evaluation using COVID-19 data highlights the practicality and effectiveness of our system. Anifat M. Olawoyin, Carson K. Leung, Hoang Hai Nguyen, Alfredo Cuzzocrea |
IEEE Big Data | 2 |
| 2024 | Are Existing Large Language Models Robust Against Jailbreak Attacks?abstractThe safety and robustness of Large Language Models (LLMs) are major challenges in developing generative AI applications. One key issue is the vulnerability to prompt jailbreak attacks, which pose a significant threat to building secure and resilient LLM-based applications. In this work, we present a framework for understanding and evaluating the behaviors of popular LLMs by categorizing their responses into five distinct exposure levels. Additionally, we introduce a novel language attack that circumvents LLMs’ defenses by translating jailbreak prompts into languages such as Arabic, Chinese, and Greek. Despite ongoing efforts to enhance LLMs’ safety, we find that nearly all popular LLMs can be jailbroken. Our findings offer detailed insights into LLMs’ behavior, improve diagnostic capabilities, and support targeted safety improvements. Baha Rababah, S. Tommy Wu, Matthew Kwiatkowski, Carson K. Leung, Cuneyt Gurcan Akcora |
IEEE Big Data | 4 |
| 2024 | SoK: Prompt Hacking of Large Language ModelsabstractThe safety and robustness of large language models (LLMs) based applications remain critical challenges in artificial intelligence. Among the key threats to these applications are prompt hacking attacks, which can significantly undermine the security and reliability of LLM-based systems. In this work, we offer a comprehensive and systematic overview of three distinct types of prompt hacking: jailbreaking, leaking, and injection, addressing the nuances that differentiate them despite their overlapping characteristics. To enhance the evaluation of LLM-based applications, we propose a novel framework that categorizes LLM responses into five distinct classes, moving beyond the traditional binary classification. This approach provides more granular insights into the AI’s behavior, improving diagnostic precision and enabling more targeted enhancements to the system’s safety and robustness. Baha Rababah, S. Tommy Wu, Matthew Kwiatkowski, Carson K. Leung, Cuneyt Gurcan Akcora |
IEEE Big Data | 4 |
| 2024 | A Metaverse Platform for Air Pollution Analysis in Supporting Smart and Sustainable City DevelopmentabstractIn today's data-centric world, analyzing vast volumes of diverse and complex information has paved the way for uncovering valuable insights. These extensive datasets are often known as big data. Big data find application in various fields such as healthcare, gaming, financial markets, and business intelligence. Additionally, analyzing big data can contribute to enhancing environmental sustainability, as well as city planning and development. On the one hand, air pollution in urban areas is frequently identified as a major factor negatively affecting human health, with vehicle emissions being a significant contributor to poor air quality. On the other hand, increased greenspace and vegetation positively contribute to better air quality. In this paper, we present a data science and advanced analytics solution—specifically, a metaverse platform—to examine the relationship between urban factors and air quality. Our solution leverages data mining and visualization techniques in a metaverse platform to extract meaningful insights. Moreover, we analyze real traffic data from a mid-size Canadian city to guide our study. The findings from this data science research can inform practical strategies—such as promoting green infrastructure and implementing zoning policies—towards building and development of smart and sustainable cities. Juan C. Armijos, Garik Avagyan, Carson K. Leung, Jasmine J. Tabuzo, Aivee F. Teodocio |
DSAA | 3 |
| 2024 | A Database Engineered System for Big Data Analytics on Tornado Climatology
Fengfan Bian, Carson K. Leung, Piers Grenier, Harry Pu, Samuel Ning |
IDEAS | 2 |
| 2024 | Discovering Interesting Patterns from HypergraphsabstractA hypergraph is a complex data structure capable of expressing associations among any number of data entities. Overcoming the limitations of traditional graphs, hypergraphs are useful to model real-life problems. Frequent pattern mining is one of the most popular problems in data mining with a lot of applications. To the best of our knowledge, there exists no flexible pattern mining framework for hypergraph databases decomposing associations among data entities. In this article, we propose a flexible and complete framework for mining frequent patterns from a collection of hypergraphs. To discover more interesting patterns beyond the traditional frequent patterns, we propose frameworks for weighted and uncertain hypergraph mining also. We develop three algorithms for mining frequent, weighted, and uncertain hypergraph patterns efficiently by introducing a canonical labeling technique for isomorphic hypergraphs. Extensive experiments have been conducted on real-life hypergraph databases to show both the effectiveness and efficiency of our proposed frameworks and algorithms. Md. Tanvir Alam, Chowdhury Farhan Ahmed, Mohammad Samiullah 0001, Carson K. Leung |
ACM Trans. Knowl. Discov. Data | 4 |
| 2023 | Social network mining and analytics for quantitative patternsabstractFrequent pattern mining has gained popularity in the realm of knowledge discovery and big data analytics as it identifies sets of items that frequently co-occur (e.g., popular merchandise items or social events). In general, frequent pattern mining can be broadly classified into two categories: (i) transaction-centric algorithms that mines frequent patterns horizontally and (ii) item-centric mining algorithms that mines frequent patterns vertically. Irrespective of their categories, traditional frequent pattern mining algorithms aim to find Boolean frequent patterns, revealing whether some specific items are present in (or absent from) the discovered patterns. In the context of social network mining and analytics, Boolean frequent pattern algorithms can help reveal whether a social entity follows another in a network or on a social networking site. However, in numerous real-life applications, quantities of items within patterns become essential. For example, the quantity of followed items (e.g., like posts) can significantly influence the social interactions between entities in a network. In this paper, we present a social network mining and analytics algorithm---called QSN---for discovering quantitative frequent patterns from social networks. The algorithm represents the big data as a collection of item-centric bitmaps, each capturing the absence or presence of a transaction containing the item, along with the quantity of that item in each transaction. Subsequently, it vertically mines quantitative frequent patterns, strategically avoiding the generation of an excessive number of redundant candidate patterns, thereby accelerating the mining process. Results of our evaluation demonstrate the superiority of our QSN algorithm over the existing horizontal quantitative frequent pattern algorithm called MQA-M, highlighting its efficacy in social network mining and analytics. Connor C. J. Hryhoruk, Carson K. Leung, Adam G. M. Pazdor |
ASONAM | 2 |
| 2023 | Personalized privacy-preserving semi-centralized recommendation system in a social networkabstractIn the contemporary era of big data, recommendation systems play a crucial role in guiding our daily decision-making amidst an overwhelming array of choices. Personalized recommendations have become increasingly popular by tailoring suggestions to user profiles, preferences, and interests. While many existing systems rely on centralizing data for making recommendations, the revelation of sensitive information poses a significant privacy concern. Studies indicate the potential to de-identify anonymous users, exposing details such as political views or sexual orientations through seemingly innocuous data, like movie ratings. In this paper, we introduce a personalized privacy-preserving semi-centralized recommendation system in a social network known as trust-based social network (TSN) to address these privacy challenges. TSN addresses privacy concerns by semi-centralizing data, treating each node in the network as an independent social entities. Data are distributed to social entities within trusted social networks, and the recommendation service provider only collects obfuscated data from social entities through the adoption of a differential-privacy mechanism. Consequently, data within TSN are either protected within local trusted social networks or obfuscated outside of these networks. The final recommendation is generated by combining local suggestions from the trusted social network with obfuscated global suggestions from the service provider. The emphasis on local suggestions ensures highly personalized recommendations. Evaluation results demonstrate that TSN achieves high accuracy in recommendations while effectively safeguarding user privacy. Carson K. Leung, Qi Wen 0001 |
ASONAM | 1 |
| 2023 | Machine-Learning-Based Multidimensional Big Data Analytics over Clouds via Multi-Columnar Big OLAP Data Cube CompressionabstractThis paper proposes a new theory on combining innovative Multidimensional Big Data Analytics with well-known Machine Learning (ML) in order to magnify the expressive power and the accuracy of knowledge insights discovery from massive big datasets. At the level of enabling technology, with the goal of fully supporting this novel paradigm, the issue of managing and mining big OLAP data cubes over Clouds arises. Due to computational complexity requirements, the latter challenge is addressed by proposing an innovative solution for (1) representing big OLAP data cubes over Clouds via a multi-column-based representation, and (2) compressing the deriving multi-column representations for achieving the desired effectiveness and efficiency. This paper introduces the fundamental model of Machine-Learning-Based Multidimensional Big Data Analytics, along with a reference architecture implementing it. Alfredo Cuzzocrea, Abderraouf Hafsaoui, Carson K. Leung |
IEEE Big Data | 3 |
| 2023 | A Privacy-Preserving Semi-Decentralized Personalized Recommendation SystemabstractIn the present era of big data, recommendation systems play a crucial role in our daily lives by assisting us in making quicker and more informed decisions from a vast array of choices. The concept of personalized recommendations has gained widespread popularity, offering suggestions based on user profiles, preferences, and/or interests. Although many existing systems centralize data for making recommendations, the revelation of sensitive data poses a privacy concern, as research indicates the potential to de-identify anonymous users. For instance, sensitive information such as political views or sexual orientations can be inferred from seemingly non-sensitive data like product review and ratings. In this paper, we present a privacy-preserving personalized recommendation system named P2RecSys to address these privacy issues. Our system takes a semi-decentralized approach by treating each node in the network as an agent. Data are distributed to each agent within trusted networks, and the recommendation service provider only collects obfuscated data from agents using a differential-privacy mechanism. Consequently, data in P2RecSys are either safeguarded within local trusted networks or obfuscated outside of these networks. The final recommendation is then generated by combining local suggestions from the trusted network with obfuscated global suggestions from the service provider. The emphasis on local suggestions allows for highly personalized recommendations. Evaluation results demonstrate that P2RecSys achieves high accuracy in recommendations while effectively safeguarding user privacy. Carson K. Leung, Evan Madill, Qi Wen 0001 |
IEEE Big Data | 1 |
| 2023 | Bitwise Vertical Mining of Minimal Rare Patterns
Elieser Capillar, Chowdhury Abdul Mumin Ishmam, Carson K. Leung, Adam G. M. Pazdor, Prabhanshu Shrivastava, Ngoc Bao Chau Truong |
DaWaK | 3 |
| 2023 | Privacy-Preserving Learning via Data and Knowledge DistillationabstractIn the current era of data science, deep learning, computer vision and image analysis have become ubiquitous across various sectors, ranging from government agencies and large corporations to small end devices, due to their ability to simplify people’s lives. However, the widespread use of sensitive image data and the high memorization capacity of deep learning present significant privacy risks. Now, a simple Google search can yield numerous images of a person, and the knowledge that a specific patient’s record was utilized for training a specific model associated with a disease may reveal the patient’s ailment, potentially leading to membership privacy leakage and other advanced attacks in the future. Furthermore, these unprotected models may also suffer from poor generalization due to this overfitting to train data. Previous state-of-the-art methods like differential privacy (DP) and regularizer-based defenses compromised functionality, i.e., task accuracy, to preserve privacy. Such an imbalanced trade-off raises concerns about the practicability of such defenses. Other existing knowledge-transfer-based methods either reuse private data or require more public data, which could compromise privacy and may not be viable in certain domains. To address these challenges, where membership privacy is of utmost importance and utility cannot be compromised, we propose a novel collaborative distillation approach that transfers the private model’s knowledge based on a minimal amount of distilled synthetic data, leading to a compact private model in an end-to-end fashion. Empirically, our proposed method guarantees superior performance compared to most advanced models currently in use, increasing utility by almost 8%, 34%, and 6% for CIFAR-10, CIFAR-100, and MNIST, respectively. The utility resembles non-private counterparts almost closely while maintaining a respectable level of membership privacy leakage of 50-53.5%, despite employing a smaller model with 50% fewer parameters. Fahim Faisal, Carson K. Leung, Noman Mohammed, Yang Wang 0003 |
DSAA | 2 |
| 2023 | Enhanced Mining of High Utility Patterns from Streams of Dynamic ProfitabstractFrequent pattern mining has been extended to the mining of other useful patterns. These include high-utility patterns. Many traditional high-utility mining algorithms focus on algorithmic efficiency when mining high-utility patterns from static databases. These algorithms rely on an assumption that the unit utility for a given item is a constant. However, as we are living in dynamic world where the unit utility (external unit profit) may change over time, such an assumption may not truly reflect reality in the real world. However, to the best of our knowledge, not a lot of works were done on mining dynamic profit from data streams yet. The emergence of big data has led to some performance challenges such that proper big data management techniques are needed for knowledge discovery from dynamic data streams. Traditional static data mining algorithms cannot directly apply to dynamic data. Furthermore, information in the data stream might not be uniformly distributed so it introduces extra challenges to process the data. Using big data stream processing platforms is necessary when mining real-world data stream. Leveraging the big data processing framework requires having scalable algorithms. In this paper, we present an enhanced high-utility data stream algorithm—called EHUI-Stream—to speed up the execution time and reduce memory usage. Utilizing our proposed algorithm, the data stream mining performance is expected to be further enhanced against both real-world datasets and synthetic datasets. Evaluation results on real-life data demonstrate the effectiveness of our platform in scalable high-utility pattern mining for dynamic profit from data streams for social and behavioral analytics. Jiaxing Jason Mai, Carson K. Leung, Connor C. J. Hryhoruk, Adam G. M. Pazdor |
DSAA | 2 |
| 2023 | Machine Learning-Based Android Malware DetectionabstractThe use of mobile phones, particularly smartphones, has been growing exponentially in recent times. From 2016 to 2021, smartphone users increased by more than 70%. With the increase in the popularity of smartphones, smartphones have become the prime target for criminal hackers. As a result, Android malware samples are coming to the market at an alarming rate. A study shows that there are more than 4 million malicious Android apps in the market, and each day around 11,000 new malwares add to this number. To combat this mass number of malware, we need a malware detection system that is efficient in detecting malicious Android apps. There are numerous existing malware detection systems, but most of them require countless features from both dynamic and static analysis. Thus, they are not scalable, lightweight, and efficient in detecting malware. Additionally, most studies that used limited features like only permission data, had done their research on much older dataset. Hence, there is a need for new research on this topic. In this paper, we build a permission-based malware detector for Android application with a new dataset and significantly less permissions. Initially, we used support vector machine (SVM) and all the extracted permission data as features to build our classification model. The model accuracy, precision, recall and F1 score were 97.41 percent, which is higher than the other state of the art similar approaches done on an older dataset. Next, we replicated this similar study with a few different machine learning algorithms: random forest, decision tree and logistic regression, and observed they all give similar results. However, tree-based algorithm performs a little better than the other algorithms. Finally, to achieve a lightweight malware detection system, we reduced the number of permissions or the features on a two-step process, and found only a slight difference in results. In the first step, even after reducing the number of permissions by about 94%, the accuracy dropped by only 2.7%. In the second step, we further reduced the number of features or permissions and observed the difference in results. We managed to prune to 9 permissions while maintaining accuracy of 93%, which is lower than technique mentioned in other literature to reduce features. David Ojo, Nusayer Masud Siddique, Carson K. Leung, Connor C. J. Hryhoruk |
DSAA | 3 |
| 2023 | Discovery of Patent Influence with Directed Acyclic Graph Network AnalysisabstractIn the domain of research, development and innovation, every new discovery usually relies on previous evidence as the basis of any novel argument. Research papers and patents are often used to support such results. However, their discovery—especially, discovery of patents—is not an easy task; it requires a lot of time and effort to find results that are actually helpful. The influence patents have on research, development and innovation is substantial because they aim to discover related and influential patents that can tremendously help drive new discoveries and inventions. In this paper, we present a database engineered solution, which showcases techniques that enable the patent discovery. The solution leads to meaningful and relevant discovery when looking for relevant and influential patents in a given domain to be used by researchers, inventors and businesses. We also explore techniques to visually inspect large amounts of data and find other interesting results that would be difficult to view otherwise. Evaluation results show the practicality of our solution in discovering and visualizing patent influence when conducting directed acyclic graph (DAG) analysis. Carson K. Leung |
IDEAS | 1 |
| 2022 | Social Network Analysis on Interpretable Compressed Sparse NetworksabstractBig data are everywhere. World Wide Web is an example of these big data. It has become a vast data production and consumption platform, at which threads of data evolve from multiple devices, by different human interactions, over worldwide locations, under divergent distributed settings. Embedded in these big web data is implicit, previously unknown and potentially useful information and knowledge that awaited to be discovered. This calls for web intelligence solutions, which make good use of data science and data mining (especially, web mining or social network mining) to discover useful knowledge and important information from the web. As a web mining task, web structure mining aims to examine incoming and outgoing links on web pages and make recommendations of frequently referenced web pages to web surfers. As another web mining task, web usage mining aims to examine web surfer patterns and make recommendations of frequently visited pages to web surfers. While the size of the web is huge, the connection among all web pages may be sparse. In other words, the number of vertex nodes (i.e., web pages) on the web is huge, the number of directed edges (i.e., incoming and outgoing hyperlinks between web pages) may be small. This leads to a sparse web. In this paper, we present a solution for interpretable mining of frequent patterns from sparse web. In particular, we represent web structure and usage information by bitmaps to capture connections to web pages. Due to the sparsity of the web, we compress the bitmaps, and use them in mining influential patterns (e.g., popular web pages). For explainability of the mining process, we ensure the compressed bitmaps are interpretable. Evaluation on real-life web data demonstrates the effectiveness, interpretability and practicality of our solution for interpretable mining of influential patterns from sparse web. Connor C. J. Hryhoruk, Carson K. Leung |
ASONAM | 2 |
| 2022 | Social Network Analysis of Popular YouTube Videos via Vertical Quantitative MiningabstractFrequent itemset (or frequent pattern) mining is a technique used in big data mining to discover frequently occurring sets of items (such as popular co-purchased merchandise) and has numerous applications in the field of databases. Traditional frequent pattern mining algorithms only look at Boolean mining; that is, considering only the presence or absence of an item in an itemset. In this paper, we present an algorithm for mining interesting quantitative frequent patterns. Our qEclat (or Q-Eclat) algorithm extends the common Eclat algorithm to be able to vertically mine quantitative patterns. When compared with the existing MQA-M algorithm (which was built for quantitative horizontal frequent pattern mining), our evaluation results show that qEclat mines quantitative frequent patterns faster. Adam G. M. Pazdor, Carson K. Leung, Thomas J. Czubryt, Denys Popov, Sanskar Raval |
ASONAM | 2 |
| 2022 | Handwritten Word Recognition using Deep Learning Approach: A Novel Way of Generating Handwritten WordsabstractA handwritten word recognition system comes with issues such as-lack of large and diverse datasets. It is necessary to resolve such issues since millions of official documents can be digitized by training deep learning models using a large and diverse dataset. Due to the lack of data availability, the trained model does not give the expected result. Thus, it has a high chance of showing poor results. This paper proposes a novel way of generating diverse handwritten word images using handwritten characters. The idea of our project is to train the BiLSTM-CTC architecture with generated synthetic handwritten words. The whole approach shows the process of generating two types of large and diverse handwritten word datasets: overlapped and non-overlapped. Since handwritten words also have issues like overlapping between two characters, we have tried to put it into our experimental part. We have also demonstrated the process of recognizing handwritten documents using the deep learning model. For the experiments, we have targeted the Bangla language, which lacks the handwritten word dataset, and can be followed for any language. Our approach is less complex and less costly than traditional GAN models. Finally, we have evaluated our model using Word Error Rate (WER), accuracy, f1-score, precision, and recall metrics. The model gives 39% WER score, 92% percent accuracy, and 92% percent f1 scores using non-overlapped data and 63% percent WER score, 83% percent accuracy, and 85% percent f1 scores using overlapped data. Mst. Shapna Akter, Hossain Shahriar, Alfredo Cuzzocrea, Nova Ahmed, Carson K. Leung |
IEEE Big Data | 5 |
| 2022 | Preserving Privacy Integration and Mining for Big Temporal Co-occurrence PatternsabstractPrivacy policy, terms of use, public consent, reusable data, and transparency are trending words associated with the world wide web data such that privacy is now the responsibility of all stakeholders. Although privacy is a concern, integrating publicly available data may be for social good. For instance, integrating emergency calls, substance use, and overdose antagonist drug may help inform policies relating to emergency resources allocation, substance overdose antagonist drug distribution, and spiral effect of reducing overdose death. Hence, in this paper, we examine privacy preserving integration of public open data in the hierarchy of time and space. Our experimental result on four open datasets demonstrate the effectiveness of temporal and location hierarchy model in preserving privacy integration of big temporal co-occurrence data. Anifat M. Olawoyin, Carson K. Leung |
IEEE Big Data | 2 |
| 2022 | Mahalanobis Distance Based K-Means Clustering
Paul O. Brown, Meng Ching Chiang, Shiqing Guo, Yingzi Jin, Carson K. Leung, Evan L. Murray, Adam G. M. Pazdor, Alfredo Cuzzocrea |
DaWaK | 5 |
| 2022 | Q-VIPER: Quantitative Vertical Bitwise Algorithm to Mine Frequent Patterns
Thomas J. Czubryt, Carson K. Leung, Adam G. M. Pazdor |
DaWaK | 2 |
| 2022 | Enhanced Sliding Window-Based Periodic Pattern Mining from Dynamic Streams
Evan Madill, Carson K. Leung, Justin M. Gouge |
DaWaK | 2 |
| 2022 | Generating Privacy Preserving Synthetic Medical DataabstractDue to the recent development in the deep learning community and the availability of state-of-the-art models, medical practitioners are getting more interested in computer vision and deep learning for diagnosis tasks. Moreover, those medical diagnostic models can also increase the reliability of conventional findings. As radiology images can convey a lot of information for a patient’s diagnosis task, the problem is that such medical data may contain sensitive private information in their content header. De-anonymization (i.e., removal of sensitive header information) does not work well due to the re-identification risk, which may link those images to essential details (e.g., birth date, SSN, institution name, etc.), and such an approach can also reduce utility. In the medical domain, utility is significant because a less accurate diagnosis may lead to the wrong course of treatment and/or loss of life. In this paper, we developed a differentially private approach that can generate high-quality and high dimensional synthetic medical image data with guaranteed differential privacy. It can be used to create sufficient quality data to train a deep model. Moreover, we used W-GAN for bounded gradient guarantee, which eliminates the need for an extensive clipping hyperparameter search. We also added noise selectively to the generator to maintain the privacy-utility trade-off. Due to a noise-free discriminator and such selective noise addition to the generator, high-quality and reliable generated radiology images can be utilized for diagnosis tasks. Moreover, our approach can work in a distributed system where different hospitals can contain their private images in the local server and use a central server to generate synthetic radiology images without storing patient data. Fahim Faisal, Noman Mohammed, Carson K. Leung, Yang Wang 0003 |
DSAA | 3 |
| 2022 | Q-Eclat: Vertical Mining of Interesting Quantitative PatternsabstractFrequent pattern mining is a popular technique in big data mining and analytics. It discovers frequently occurring sets of items (e.g., popular merchandise items, frequently co-occurring events) from big data found in numerous database engineered applications. These frequent patterns can be discovered horizontally by transaction-centric mining algorithms or vertically by item-centric mining algorithms. Regardless of their mining direction (horizontal or vertical), traditional frequent pattern mining algorithms aim to discover Boolean frequent patterns in the sense that patterns capture the presence (or absence) of items within the discovered patterns. However, there are many real-life situations, in which quantities of items within the patterns are important. For example, the quantity of items may also affect profits of selling the items within the discovered patterns. Hence, in this paper, we present an algorithm for vertical mining of interesting quantitative frequent patterns. This Q-Eclat algorithm first represents the big data as a collection of equivalence classes according to their prefix item labels. Each domain item is represented by one of these classes. Their corresponding item-centric sets capture (a) IDs of transactions containing the item, as well as (b) the quantity of that item in each transaction. With this representation, our algorithm then vertically mines quantitative frequent patterns. When compared the existing MQA-M algorithm (which was built for quantitative horizontal frequent pattern mining), evaluation results show that our quantitative vertical Q-Eclat algorithm takes shorter runtime to mine quantitative frequent patterns. Thomas J. Czubryt, Carson K. Leung, Adam G. M. Pazdor |
IDEAS | 2 |
| 2022 | Mining weighted sequential patterns in incremental uncertain databases
Kashob Kumar Roy, Md Hasibul Haque Moon, Md Mahmudur Rahman 0002, Chowdhury Farhan Ahmed, Carson K. Leung |
Inf. Sci. | 5 |
| 2021 | Compressing and mining social network dataabstractNowadays, social networking is popular. As such, numerous social networking sites (e.g., Facebook, YouTube, Instagram) are generating very large volumes of social data rapidly. Valuable knowledge and information is embedded into these big social data. As the social network can be very sparse, it is awaiting to be (a) compressed via social network data compression and (b) analyzed and mined via social network analysis and mining. We present in this paper a solution for compressing and mining social networks. It gives an interpretable compressed representation of sparse social network, and discovers interesting patterns from the social network. Results of our evaluation show the effectiveness of our solution in explaining the compression and mining of the sparse social network data. Connor C. J. Hryhoruk, Carson K. Leung |
ASONAM | 2 |
| 2021 | A mathematical model for friend discovery from dynamic social graphsabstractNowadays, social networking is popular. As such, numerous social networking sites (e.g., Facebook, YouTube, Instagram) are generating very large volumes of social data rapidly. Valuable knowledge and information is embedded into these big social data, and is awaiting to be analyzed and mined via social network analysis and mining. In general, social networks can be represented as graphs. Because of the dynamic nature of social networking, edges and/or vertices keep adding to (or deleting from) the graphs. We present in this paper a mathematical model for friend discovery from dynamic social graphs. In particular, we focus on both linear algebra and graph theory approaches to discover interesting social entities---such as active followers---from dynamic social networks represented as dynamic directional social graphs. Carson K. Leung, Sehaj Pal Singh |
ASONAM | 1 |
| 2021 | Open Data Lake to Support Machine Learning on Arctic Big DataabstractThe era of big data is evolving with the introduction of the data lake concept. While a data warehouse provides a well-structured model to manage big data, a data lake accepts data of any types and formats with or without schema and provides access to the data for diverse communities of users. A data lake provides flexible, agile, and scalable solution to manage the ever-increasing volume of big data we are witnessing in the world today, including many siloed data collected over the years by researchers through Arctic expeditions. In this paper, we present our conceptual model of a data lake for integrating the diverse huge amount of data collected by researchers during Arctic expedition. We also design a baseline metadata using a data-driven approach to manage the disparately huge structured, semi-structured, and unstructured data collected from the Arctic region. The resulting open data lake not only effectively manages big Arctic data but also supports machine learning on these big data. Anifat M. Olawoyin, Carson K. Leung, Alfredo Cuzzocrea |
IEEE BigData | 2 |
| 2021 | Privacy-Preserving Publishing and Visualization of Spatial-Temporal InformationabstractPartially due to technological advancements as well as the availability of affordable global positioning system (GPS) and cellular devices, more spatio-temporal data can be generated and collected. The presence of spatial and temporal dimensions uniquely differentiate spatio-temporal data from classical data as spatio-temporal data points are structurally related in the context of space and time. In this paper, we present a solution for privacy-preserving publishing and visualization of spatiotemporal big data information. Specifically, it consists of a spatiotemporal hierarchy model (STHM) for some common big data management tasks such as visualization. Our data visualizer provides actionable insight to enhance data-driven decision making. It also enables the discovery of hidden patterns, clusters of events, and outliers. We design two different metrics to preprocess the spatio-temporal for data visualization. Although we demonstrate the usefulness of our solution in privacy-preserving publishing and visualization of spatio-temporal information by using big real-life parking data from two cities, our solution can be applicable for publishing and visualizing spatio-temporal information from many other big data. Anifat M. Olawoyin, Carson K. Leung, Alfredo Cuzzocrea |
IEEE BigData | 2 |
| 2021 | Health Analytics on COVID-19 Data with Few-Shot Learning
Carson K. Leung, Daryl L. X. Fung, Calvin S. H. Hoi |
DaWaK | 1 |
| 2021 | Explainable Artificial Intelligence for Data Science on Customer ChurnabstractMachine learning, as a tool, has become critical for decision-making mechanisms in the modern world. It has applications in a wide range of areas, including finance, healthcare, justice, and transportation. Unfortunately, machine learning is often considered as a “black box”. As such, recommendations made by machine learning techniques, as well as the reasoning behind those recommendations, are not easily understood by humans. In this paper, we present an explainable artificial intelligence (XAI) solution that integrates and enhances state-of-the-art techniques to produce understandable and practical explanations to end-users. To evaluate the effectiveness of our XAI solution for data science, we conduct a case study on applying our solution to explaining a random forest-based predictive model on customer churn. Results show the practicality and usefulness of our XAI solution in practical applications such as data science on customer churn. Carson K. Leung, Adam G. M. Pazdor, Joglas Souza |
DSAA | 1 |
| 2021 | Explainable Data Analytics for Disease and Healthcare InformaticsabstractWith advancements in technology, huge volumes of valuable data have been generated and collected at a rapid velocity from a wide variety of rich data sources. Examples of these valuable data include healthcare and disease data such as privacy-preserving statistics on patients who suffered from diseases like the coronavirus disease 2019 (COVID-19). Analyzing these data can be for social good. For instance, data analytics on the healthcare and disease data often leads to the discovery of useful information and knowledge about the disease. Explainable artificial intelligence (XAI) further enhances the interpretability of the discovered knowledge. Consequently, the explainable data analytics helps people to get a better understanding of the disease, which may inspire them to take part in preventing, detecting, controlling and combating the disease. In this paper, we present an explainable data analytics system for disease and healthcare informatics. Our system consists of two key components. The predictor component analyzes and mines historical disease and healthcare data for making predictions on future data. Although huge volumes of disease and healthcare data have been generated, volumes of available data may vary partially due to privacy concerns. So, the predictor makes predictions with different methods. It uses random forest With sufficient data and neural network-based few-shot learning (FSL) with limited data. The explainer component provides the general model reasoning and a meaningful explanation for specific predictions. As a database engineering application, we evaluate our system by applying it to real-life COVID-19 data. Evaluation results show the practicality of our system in explainable data analytics for disease and healthcare informatics. Carson K. Leung, Daryl L. X. Fung, Daniel Mai, Qi Wen 0001, Jason Tran, Joglas Souza |
IDEAS | 1 |
| 2021 | Mining Frequent Patterns from Hypergraph Databases
Md. Tanvir Alam, Chowdhury Farhan Ahmed, Mohammad Samiullah 0001, Carson K. Leung |
PAKDD (2) | 4 |
| 2021 | Discriminating Frequent Pattern Based Supervised Graph Embedding for Classification
Md. Tanvir Alam, Chowdhury Farhan Ahmed, Mohammad Samiullah 0001, Carson K. Leung |
PAKDD (2) | 4 |
| 2021 | Mining Sequential Patterns in Uncertain Databases Using Hierarchical Index Structure
Kashob Kumar Roy, Md Hasibul Haque Moon, Md Mahmudur Rahman 0002, Chowdhury Farhan Ahmed, Carson K. Leung |
PAKDD (2) | 5 |
| 2020 | Compression for Very Sparse Big Social DataabstractTechnological advancements in the current era of big data have led to rapid generation and collection of very large amounts of valuable data from a wide variety of rich data sources. As rich data sources, social networks consist of social entities that are linked by some social relationships (e.g., kinship, colleagueship, co-authorship, friendship, followship). Usually, these networks are very big but also very sparse. Embedded in the very sparse but very big networks are implicit, previously unknown and potentially useful information and knowledge that can be discovered by social network analysis and mining. In this paper, we aim to discover interesting social relationships from very sparse but very big social network data. Due to the sparsity of the data, we effectively compress bitmaps representing social entities in the data, from which useful information can be mined and interesting knowledge can be discovered. Evaluation results show the effectiveness of our compression scheme for very sparse but very big social network data. Carson K. Leung, Yibin Zhang 0002, Fan Jiang 0001 |
ASONAM | 1 |
| 2020 | A Theoretical Approach for Discovery of Friends from Directed Social GraphsabstractSince social networking has been popular in the current era of big data, numerous social networking sites (e.g., Instagram, Twitter) have generated huge volumes of social data at a rapid rate. Embedded into these data are valuable information and knowledge. This calls for social network analysis and mining. In this paper, we specifically aim to discover interesting relationships in directed social graphs via a theoretical approach. More specifically, we examine both graph theory and linear algebra approaches to discover interesting entities (e.g., popular followees, second-degree followees) from social networks represented in the form of big directional graphs. Sehaj Pal Singh, Carson K. Leung |
ASONAM | 2 |
| 2020 | Machine Learning and OLAP on Big COVID-19 DataabstractIn the current technological era, huge amounts of big data are generated and collected from a wide variety of rich data sources. These big data can be of different levels of veracity in the sense that some of them are precise while some others are imprecise and uncertain. Embedded in these big data are useful information and valuable knowledge to be discovered. An example of these big data is healthcare and epidemiological data such as data related to patients who suffered from epidemic diseases like the coronavirus disease 2019 (COVID-19). Knowledge discovered from these epidemiological data-via data science techniques such as machine learning, data mining, and online analytical processing (OLAP)-helps researchers, epidemiologists and policy makers to get a better understanding of the disease, which may inspire them to come up ways to detect, control and combat the disease. In this paper, we present a machine learning and big data analytic tool for processing and analyzing COVID-19 epidemiological data. Specifically, the tool makes good use of taxonomy and OLAP to generalize some specific attributes into some generalized attributes for effective big data analytics. Instead of ignoring unknown or unstated values of some attributes, the tool provides users with flexibility of including or excluding these values, depending on their preference and applications. Moreover, the tool discovers frequent patterns and their related patterns, which help reveal some useful knowledge such as absolute and relative frequency of the patterns. Furthermore, the tool learns from the patterns discovered from historical data and predicts useful information such as clinical outcomes for future data. As such, the tool helps users to get a better understanding of information about the confirmed cases of COVID-19. Although this tool is designed for machine learning and analytics of big epidemiological data, it would be applicable to machine learning and analytics of big data in many other real-life applications and services. Carson K. Leung, Yubo Chen 0003, Calvin S. H. Hoi, Siyuan Shang, Alfredo Cuzzocrea |
IEEE BigData | 1 |
| 2020 | Preserving Privacy of Temporal Big DataabstractIn the current technological era, huge amounts of big data are generated and collected from a wide variety of rich data sources. Embedded in these big data are useful information and valuable knowledge to be utilized. With the popularity of initiatives of open data, more big data have been published on open data platforms and made accessible to the public. To preserve privacy while maintaining the utility of data, research on privacy-preserving publishing has focused on preserving privacy of sensitive personal data such as patient data for health related applications. However, there are many other real-life situations, in which personal data of individual citizens and their daily routines need to be preserved when publishing. In this paper, we examine the problem of preserving privacy of temporal big data. Specifically, we present a temporal hierarchy privacy preserving model (THPPM) for some common daily routines-for example, parking. The model adapts and extends temporal hierarchy to generalize temporal data related to timestamp and spatial data related to check-in location. It also makes good use of generalized temporal representative points to preserve privacy of specific temporal data points. Evaluations on two real-life datasets on parking tickets for the US city of Buffalo and the Canadian city of Toronto shows that effectiveness and practicality of our THPPM in preserving privacy of temporal big data. Although this model is demonstrated and evaluated on parking ticket data, it would be applicable to preserving privacy of temporal big data for many other real-life applications and services. Anifat M. Olawoyin, Carson K. Leung, Alfredo Cuzzocrea |
IEEE BigData | 2 |
| 2020 | Privacy-Preserving Spatio-Temporal Patient Data Publishing
Anifat M. Olawoyin, Carson K. Leung, Ratna Choudhury |
DEXA (2) | 2 |
| 2020 | Data science for healthcare predictive analyticsabstractBig data are everywhere nowadays. Many businesses possess big data for their success because big data are very useful and are considered as new oil. For instance, big data are very important in predicting the trends on what will happen in the future. Many researchers have generated or gathered data to further enhance their research and to apply them to numerous real-life applications. Examples of big data include healthcare patient data. To improve the detection of illnesses and diseases, researchers have gathered healthcare patient data, examined the diagnosis on healthcare patient data (e.g., cells, blood count, antibodies count), and compared with previous data to determine if a specific illness or disease exist. Having an automatic predictive method for healthcare and disease analytics would be desirable. In this paper, we focus on healthcare mining, which aims to computationally discover knowledge from healthcare data. In particular, we present a data science framework with two predictive analytic algorithms for accurate prediction on the trends of cancer cases. The algorithms predict cancerous cells based on the information of the cell data from some data samples. Evaluation results on several real-life datasets related to the breast cancer demosntrate the effectiveness of our data science framework and predictive algorithms in healthcare data analytics. Carson K. Leung, Daryl L. X. Fung, Saad B. Mushtaq, Owen T. Leduchowski, Robert Luc Bouchard, Alfredo Cuzzocrea, Christine Y. Zhang |
IDEAS | 1 |
| 2020 | Effective privacy preserving data publishing by vectorization
Chris Soo-Hyun Eom, Charles Cheolgi Lee, Wookey Lee, Carson K. Leung |
Inf. Sci. | 4 |
| 2019 | DeepGx: deep learning using gene expression for cancer classificationabstractThis paper aims to explore the problems associated in solving the classification of cancer in gene expression data using deep learning model. Our proposed solution for the cancer classification of ribonucleic acid sequencing (RNA-seq) extracted from the Pan-Cancer Atlas is to transform the 1-dimensional (1D) gene expression values into 2-dimensional (2D) images. This solution of embedding the gene expression values into a 2D image considers the overall features of the genes and computes features that are needed in the classification task of the deep learning model by using the convolutional neural network (CNN). When training and testing the 33 cohorts of cancer types in the convolutional neural network, our classification model led to an accuracy of 95.65%. This result is reasonably good when compared with existing works that use multiclass label classification. We also examine the genes based on their significance related to cancer types through the heat map and associate them with biomarkers. Our CNN for the classification task fosters the deep learning framework in the cancer genome analysis and leads to better understanding of complex features in cancer disease. Joseph M. de Guia, Madhavi Devaraj, Carson K. Leung |
ASONAM | 3 |
| 2019 | Flexible compression of big dataabstractHigh volumes of valuable data and information can be easily collected in the current era of big data. As rich and constant sources of big data, an incredible amount of people from different social stratum take part in social networks. Hence, social networks are desired for many research topics. In social networks, users (or social entities) are often linked by some 'following' relationships. As the social networks growing, some famous users account (or social entities) might be followed by a large number of same other users. In this situation, we call those famous users as frequently followed groups, which some researchers (or businesses) may be interested in them for investigating. However, the discovery of those frequently followed groups might be difficult and challenging because the following data in social networks are usually very big but sparse (huge number of users lead to big 'following' data, but each user is likely only following a small number of other users). As a result, in this paper, we present a new compression model, which can be used during mining these very big but sparse social networks for discovering the frequently followed groups of users/social entities. Carson K. Leung, Fan Jiang 0001, Yibin Zhang 0002 |
ASONAM | 1 |
| 2019 | Personalized DeepInf: Enhanced Social Influence Prediction with Deep Learning and Transfer LearningabstractSocial influence is referred to as the phenomenon that one's opinions or behaviors be affected by others. Nowadays, the potential impact of social influence analysis (SIA) is significant. For example, SIA applications can include viral marketing, online content recommendation. Convention social influence analysis uses hand-crafted features and requires domain expert knowledge. Such an approach is not scalable and introduces a high cost. To overcome these disadvantages, deep learning based approaches was introduced. One of the most recent approaches is DeepInf, which is an end-to-end framework for predicting social influence by learning user's latent features. We extended DeefInf in the current paper by integrating teleport probability $\alpha$ from the domain of page rank into the graph convolution network (GCN) model to enhance the performance. Furthermore, we also propose an algorithm called hybrid personalized propagation of neural predictions (HPPNP), which shows an impressive performance in terms of prediction accuracy compared to existing methods. We reused the datasets from DeepInf and performed extensive experiments on Open Academic Graph, Twitter, DIGG datasets. By optimally sampling the teleport probability $\alpha$, the experimental results show that our model performs the best when compared with existing methods on different datasets. These results demonstrates the effectiveness of our enhanced personalized DeepInf-namely, HPPNP-in social influence prediction via both deep and transfer learning. Carson K. Leung, Alfredo Cuzzocrea, Jiaxing Jason Mai, Deyu Deng, Fan Jiang 0001 |
IEEE BigData | 1 |
| 2019 | Exploiting Anti-Monotonic Constraints in Mining Palindromic Motifs from Big Genomic DataabstractThe advent of high-throughput technologies such as Illumina HiSeq X, mass spectrometry, and microarray heralds a new era of big biological datasets in computational biology. This digital revolution in bioinformatics has generated unprecedented volumes of omics data (e.g., transcriptomes, genomes, proteomes, metabolomes) with various degrees of veracities and values. These deluge of omics data are awash with a wealth of information in the form of frequently repeated contiguous patterns-namely, sequence motifs. Sequence motifs are short repeated contiguous subsequences located in the promoter region of a genome sequence. On some occasions, users are interested in mining only a particular type of sequence motifs (e.g., palindromic motifs). In genomics, palindromes are sequences from the nucleotide bases from deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) strands that are symmetrical in the sense that they read exactly the same as their complementary sequences in the reverse direction. The use of classical constraints (e.g., anti-monotonic, succinct, and/or convertible constraints) allows users to specify their interest in the universal search space, and thus enhancing distinct and effective pruning of the search space- leading to a reduction in the computational time required for the mining process. Despite several attempts made by existing algorithms for mining palindromic motifs from DNA sequences, a major drawback stems from the high volumes of the DNA sequences leading to high complexities and turnaround time of the algorithms. To this end, we propose a parallel scalable sequential mining algorithm that exploits some features of anti-monotonic constraints-using the in-memory computing model of the Apache Spark framework deployed on a cluster of a homogeneous distributed-memory system-for mining palindromic motifs from high volumes of DNA sequences. To evaluate our algorithm, we obtained the human genome (Homo sapiens) assemblies GRCh37 patch 13 (hg19), which is of size 3.2 GB from the Ensembl data repository. It contains 104,763 protein-coding sequences and 24,513 non-coding sequences. Evaluation results show that our algorithm extracts accurate palindromic motifs using a short turnaround time. Oluwafemi A. Sarumi, Carson K. Leung |
IEEE BigData | 2 |
| 2019 | Fast Privacy-Preserving Keyword Search on Encrypted Outsourced DataabstractCloud providers offer storage as a service to the data owners to store emails and files on the cloud server. However, sensitive data should be encrypted before storing on the cloud server to avoid privacy concerns. With the encryption of documents, it is not feasible for data owners to retrieve documents based on keyword search as they can do with plain text documents. Hence, it is desirable to perform a multi-keyword search on encrypted data. To achieve this goal, we present a fast privacy-preserving model for keyword search on encrypted outsourced data in this paper. Specifically, the model first performs a keyword search on encrypted data and checks its support for dynamic operations. Based on keyword search results, it then sorts all the relevant data documents using the number of keywords matched for a given query. To evaluate its performance of our model, we applied the standard metrics like precision and recall. The results show the effectiveness of our privacy-preserving keyword search on encrypted outsourced data. Bryan H. Wodi, Carson K. Leung, Alfredo Cuzzocrea, S. Sourav |
IEEE BigData | 2 |
| 2019 | Urban Analytics of Big Transportation Data for Supporting Smart Cities
Carson K. Leung, Peter Braun 0004, Calvin S. H. Hoi, Joglas Souza, Alfredo Cuzzocrea |
DaWaK | 1 |
| 2019 | A Flexible Query Answering System for Movie Analytics
Carson K. Leung, Lucas B. Eckhardt, Amanjyot Singh Sainbhi, Cong Thanh Kevin Tran, Qi Wen 0001, Wookey Lee |
FQAS | 1 |
| 2019 | Pattern mining for knowledge discoveryabstractPattern mining aims to discover from data the implicit, previously unknown and potentially useful information and knowledge in the form of patterns. Over the past 20 years, numerous pattern mining algorithms have been proposed. They focused on algorithmic efficiency, functionalities, and other aspects. These algorithms have been applied to various real-life applications running in serial, parallel, and/or high-performing computing environments. In this paper, we review many existing pattern mining algorithms and suggest some pattern mining algorithms---especially, hybrid vertical frequent pattern mining running in serial, parallel, high-performing computing, and/or edge/fog environments---to discover knowledge from dataset. Carson K. Leung |
IDEAS | 1 |
| 2019 | A Theoretical Approach to Discover Mutual Friendships from Social Graph NetworksabstractDue to popularity of social networking in the current era of big data, many social networking sites (e.g., Facebook) has been generating huge volumes of social data. In this paper, we aim to discover interesting relationships in a undirected (social) graph via a theoretical approach. Specifically, we examine the graph theory and linear algebra approaches to find mutual friendships from social networks represented in the form of big graphs or big graph databases. Sehaj Pal Singh, Carson K. Leung, Fan Jiang 0001, Alfredo Cuzzocrea |
iiWAS | 2 |
| 2019 | Mining weighted frequent sequences in uncertain databases
Md Mahmudur Rahman 0002, Chowdhury Farhan Ahmed, Carson K. Leung |
Inf. Sci. | 3 |
| 2018 | Social Network Mining for Recommendation of Friends Based on Music InterestsabstractWith the rapid development of technology and software, social media have become a necessity in our daily lives as it is a way for people to keep in touch with friends and share about current events. Some of the most popular social media and social networking sites that people use include Facebook, Instagram, Snapchat, and Twitter. Finding compatible persons to be friends on social media can be a challenge as many of the people recommended to the user by social media are people who are already friends with them or have been followed. However, when users are looking for friends, the real concern is whether they have common interests or hobbies with each other and whether they often interact with one another. In this paper, we propose friend recommendation algorithms revolving around music interests and interactions in social media. Chenxi Fan, Huizi Hao, Carson K. Leung, Leslie Yu Sun, Jennifer Tran |
ASONAM | 3 |
| 2018 | Mining 'Following' Patterns from Big but Sparsely Distributed Social Network DataabstractIn the current era of big data, advanced technology has led to easy collection or generation of high volumes of a wide variety of valuable data of different veracity. As rich sources of big data, social networks consist of users (or social entities) who are often linked by some interdependency such as `following' relationships. Since these big social networks keep growing at a high velocity, there are situations in which an individual user (or business) wants to find those frequently followed groups of social entities so that he can also follow the same groups. Discovery of these frequently followed groups can be challenging because the social networks are usually big (with lots of users/social entities) but can be sparsely distributed (with most users only know some but not all users/social entities in some portions of a social network). In this paper, we present a social network mining algorithm that uses different compressed models to space-efficiently represent social entities so as to facilitate the discovery of groups of frequently followed social entities from these big but sparsely distributed social networks. Evaluation results show the practicality of our algorithm in efficient mining of `following' patterns from big but sparsely distributed social networks. Carson K. Leung, Ryan Middleton, Adam G. M. Pazdor, Yeyoung Won |
ASONAM | 1 |
| 2018 | Privacy-Preserving Frequent Pattern Mining from Big Uncertain DataabstractAs we are living in the era of big data, high volumes of wide varieties of data which may be of different veracity (e.g., precise data, imprecise and uncertain data) are easily generated or collected at a high velocity in many real-life applications. Embedded in these big data is valuable knowledge and useful information, which can be discovered by big data science solutions. As a popular data science task, frequent pattern mining aims to discover implicit, previously unknown and potentially useful information and valuable knowledge in terms of sets of frequently co-occurring merchandise items and/or events. Many of the existing frequent pattern mining algorithms use a transaction-centric mining approach to find frequent patterns from precise data. However, there are situations in which an item-centric mining approach is more appropriate, and there are also situations in which data are imprecise and uncertain. Hence, in this paper, we present an item-centric algorithm for mining frequent patterns from big uncertain data. In recent years, big data have been gaining the attention from the research community as driven by relevant technological innovations (e.g., clouds) and novel paradigms (e.g., social networks). As big data are typically published online to support knowledge management and fruition processes, these big data are usually handled by multiple owners with possible secure multi-part computation issues. Thus, privacy and security of big data has become a fundamental problem in this research context. In this paper, we present, not only an item-centric algorithm for mining frequent patterns from big uncertain data, but also a privacy-preserving algorithm. In other words, we present- in this paper-a privacy-preserving item-centric algorithm for mining frequent patterns from big uncertain data. Results of our analytical and empirical evaluation show the effectiveness of our algorithm in mining frequent patterns from big uncertain data in a privacy-preserving manner. Carson K. Leung, Calvin S. H. Hoi, Adam G. M. Pazdor, Bryan H. Wodi, Alfredo Cuzzocrea |
IEEE BigData | 1 |
| 2018 | Effective Classification of Ground Transportation Modes for Urban Data Mining in Smart Cities
Carson K. Leung, Peter Braun 0004, Adam G. M. Pazdor |
DaWaK | 1 |
| 2018 | Scalable Vertical Mining for Big Data Analytics of Frequent Itemsets
Carson K. Leung, Hao Zhang 0027, Joglas Souza, Wookey Lee |
DEXA (1) | 1 |
| 2018 | WFSM-MaxPWS: An Efficient Approach for Mining Weighted Frequent Subgraphs from Edge-Weighted Graph Databases
Md. Ashraful Islam 0001, Chowdhury Farhan Ahmed, Carson K. Leung, Calvin S. H. Hoi |
PAKDD (3) | 3 |
| 2018 | Web Page Recommendation from Sparse Big Web DataabstractIn many real-life web applications, web surfers would like to get recommendation on which collections of web pages that would be interested to them or that they should follow. In order to discover this information and make recommendation, data analytics-and specially, association rule mining or web data mining-is in demand. Since its introduction, association rule mining has drawn attention of many researchers. Consequently, many association rule mining algorithms have been proposed for finding interesting relationships-in the form of association rules-among frequently occurring patterns. For instance, in IEEE/WIC/ACM WI 2016 and 2017, serial and parallel algorithms were proposed to find interesting web pages. However, like most of the existing association rule mining algorithms, these two algorithms also were not designed for mining big data. Moreover, the search space of web pages can sparse in the sense that web pages are connected to a small subset of all web pages in the search space. In this paper, we present a compact bitwise representation for web pages in the search space. Such a representation can then be used with a bitwise serial or parallel association rule mining system for web mining and recommendation. Evaluation results show the effectiveness of our compression and the practicality of our algorithm-which discovers popular pages on the web, which in turn gives the web surfers recommendation of web pages that might be interested to them-in real-life web applications. Carson K. Leung, Fan Jiang 0001, Joglas Souza |
WI | 1 |
| 2017 | Efficient Mining of 'Following' Patterns from Very Big but Sparse Social NetworksabstractAdvances in technology in the current era of big data has led to the high-velocity generation of high volumes of a wide variety of valuable data of different veracity. As rich sources of big data, social networks consist of users (or social entities) who are often linked by some interdependency such as 'following' relationships. Given these big social networks keep growing, there are situations in which an individual user (or business) wants to find those frequently followed groups of social entities so that he can follow the same groups. Discovery of these frequently followed groups can be challenging because the social networks are usually very big (with lots of users/social entities) but can be sparse (with most users only know some but not all users/social entities in a social network). In this paper, we present a few social network mining algorithms that use compressed models in mining these very big but sparse social networks for discovering groups of frequently followed social entities. Evaluation results show the practicality of our algorithms in efficient mining of 'following' patterns from very big but sparse social networks. Carson K. Leung, Fan Jiang 0001 |
ASONAM | 1 |
| 2017 | MapReduce-Based Complex Big Data Analytics over Uncertain and Imprecise Social Networks
Peter Braun 0004, Alfredo Cuzzocrea, Fan Jiang 0001, Carson K. Leung, Adam G. M. Pazdor |
DaWaK | 4 |
| 2017 | Social Media Mining: Prediction of Box Office RevenueabstractIn recent years, social media has played a huge role in how we share and communicate our thoughts and opinions. This information can very valuable for companies and governments as it can be used to analyze public mood and opinion which is a very powerful tool. In this paper, we present a system that mines social media content from a platform such as Twitter for predicting future outcomes. Specifically, it uses chatter from Twitter to predict box office revenue of movies by extracting features such as tweets and their sentiments. Then, by using these features, our system constructs a polynomial regression model for predicting box office revenue. Experimental results show the effectiveness of our system in mining social media and predicting box office revenue. Deepankar Choudhery, Carson K. Leung |
IDEAS | 2 |
| 2017 | Bitwise parallel association rule mining for web page recommendationabstractFor many real-life web applications, web surfers would like to get recommendation on which collections of web pages that would be interested to them or that they should follow. In order to discover this information and make recommendation, data mining---and specially, association rule mining or web mining---is in demand. Since its introduction, association rule mining has drawn attention of many researchers. Consequently, many association rule mining algorithms have been proposed for finding interesting relationships---in the form of association rules---among frequently occurring patterns. These algorithms include level-wise Apriori-based algorithms, tree-based algorithms, hyperlinked array structure based algorithms, and vertical mining algorithms. While these algorithms are popular, they suffer from some drawbacks. Moreover, as we are living in the era of big data, high volumes of a wide variety of valuable data of different veracity collected at a high velocity post another challenges to data science and big data analytics. To deal with these big data while avoiding the drawbacks of existing algorithms, we present a bitwise parallel association rule mining system for web mining and recommendation in this paper. Evaluation results show the effectiveness and practicality of our parallel algorithm---which discovers popular pages on the web, which in turn gives the web surfers recommendation of web pages that might be interested to them---in real-life web applications. Carson K. Leung, Fan Jiang 0001, Adam G. M. Pazdor |
WI | 1 |
| 2016 | B-mine: Frequent Pattern Mining and Its Application to Knowledge Discovery from Social Networks
Fan Jiang 0001, Carson K. Leung, Hao Zhang 0027 |
APWeb (1) | 2 |
| 2016 | Big data mining of social networks for friend recommendationabstractIn the current era of big data, high volumes of valuable data can be easily collected and generated. Social networks are examples of generating sources of these big data. Users in these social networks are often linked by some interdependency such as friendship. As these big social networks keep growing, there are situations in which an individual user wants to find popular groups of friends so that he can recommend the same groups to other users. In this paper, we present a big data analytic solution that uses the MapReduce model in mining these big social networks for discovering groups of frequently connected users for friend recommendation. Evaluation results show the efficiency and practicality of our data analytic solution in mining big social networks, discovering popular users, and recommending friends. Fan Jiang 0001, Carson K. Leung, Adam G. M. Pazdor |
ASONAM | 2 |
| 2016 | Mining 'following' patterns from big sparse social networksabstractIn the current era of big data, high volumes of valuable data can be easily collected and generated. Social networks are examples of generating sources of these big data. Users (or social entities) in these social networks are often linked by some interdependency such as friendship or `following' relationships. As these big social networks keep growing, there are situations in which an individual user (or business) wants to find those frequently followed groups of social entities so that he can follow the same groups. Discovery of these frequently followed groups can be challenging because the social networks are usually big (with lots of users/social entities) but sparse (with most users only know some but not all users/social entities in a social network). In this paper, we present a data analytic solution that uses a compression model in mining these big but sparse social networks for discovering groups of frequently followed social entities. Evaluation results show the efficiency and practicality of our data analytic solution in discovering `following' patterns from social networks. Carson K. Leung, Edson M. Dela Cruz, Trevor L. Cook, Fan Jiang 0001 |
ASONAM | 1 |
| 2016 | Clickstream Prediction Using Sequential Stream Mining Techniques with Markov ChainsabstractAs one of data mining tasks, sequential pattern mining provides valuable information about frequent patterns of users over time. For instance, frequent sequential patterns can be applicable to analyze user clickstreams for determination of web navigation patterns, genome sequences, and customer purchasing patterns. In many real-life situations, data to be mined are continuously changing. Moreover, these data are streaming at a high velocity, which leads to impracticality of storing all these data in memory. Hence, to handle these situations, we propose three stream mining algorithms to first find frequent sequential patterns. The algorithms then form statistical models, which are stored as Markov chains or transition matrices capturing frequent sequential patterns mined so far, to predict future user clickstream (e.g., the web page the user will visit next). Experimental results show the efficiency and prediction accuracy of our proposed Markov chain-based sequential stream mining algorithms in clickstream prediction. Shelby D. Bernhard, Carson K. Leung, Vanessa J. Reimer, Joshua Westlake |
IDEAS | 2 |
| 2016 | Computing Theoretically-Sound Upper Bounds to Expected Support for Frequent Pattern Mining Problems over Uncertain Big Data
Alfredo Cuzzocrea, Carson K. Leung |
IPMU (2) | 2 |
| 2016 | Data Mining Meets HCI: Data and Visual Analytics of Frequent Patterns
Carson K. Leung, Christopher L. Carmichael, Yaroslav Hayduk, Fan Jiang 0001, Vadim V. Kononov, Adam G. M. Pazdor |
ECML/PKDD (3) | 1 |
| 2016 | An Interactive Circular Visual Analytic Tool for Visualization of Web DataabstractVisual analytics on frequent web usage patterns aims to help users to (i) analyze the data so as to discover implicit, previously unknown and potentially useful information in the form of collections of frequently visited web pages in a single session and to (ii) visually represent the discovered knowledge so as to gain insight about the data. In this paper, we propose an interactive visual analytics tool (icVAT) for frequent pattern mining. It uses an orientation free, circular layout to show frequent patterns. Moreover, we provide users with interactive feature to explicitly show connections between superset and subsets of sets of visited web pages. Experimental results show the effectiveness of our icVAT for visual analytics of frequent patterns about web data. Patrick M. J. Dubois, Zhao Han, Fan Jiang 0001, Carson K. Leung |
WI | 4 |
| 2016 | Web Page Recommendation Based on Bitwise Frequent Pattern MiningabstractIn many applications, web surfers would like to get recommendation on which collections of web pages that would be interested to them or that they should follow. In order to discover this information and make recommendation, data mining in general-or frequent pattern mining in specific-can be applicable. Since its introduction, frequent pattern mining has drawn attention from many researchers. Consequently, many frequent pattern mining algorithms have been proposed, which include levelwise Apriori-based algorithms, tree-based algorithms, hyperlinked array structure based algorithms, as well as vertical mining algorithms. While these algorithms are popular, they also suffer from some drawbacks. To avoid these drawbacks, we propose an alternative frequent pattern mining algorithm called BW-mine in this paper. Evaluation results show that our proposed algorithm is both space-and time-efficient. Furthermore, to show the practicality of BW-mine in real-life applications, we apply BW-mine to discover popular pages on the web, which in turn gives the web surfers recommendation of web pages that might be interested to them. Fan Jiang 0001, Carson K. Leung, Adam G. M. Pazdor |
WI | 2 |
| 2016 | Mining interesting patterns from uncertain databases
Akiz Uddin Ahmed, Chowdhury Farhan Ahmed, Mohammad Samiullah 0001, Nahim Adnan, Carson K. Leung |
Inf. Sci. | 5 |
| 2015 | Probabilistic Frequent Pattern Mining by PUH-Mine
Wenzhu Tong, Carson K. Leung, Dacheng Liu, Jialiang Yu |
APWeb | 2 |
| 2015 | Big Data Analytics of Social Networks for the Discovery of "Following" Patterns
Carson K. Leung, Fan Jiang 0001 |
DaWaK | 1 |
| 2015 | Balancing Tree Size and Accuracy in Fast Mining of Uncertain Frequent Patterns
Carson K. Leung, Richard Kyle MacKinnon |
DaWaK | 1 |
| 2014 | Efficient Frequent Itemset Mining from Dense Data Streams
Alfredo Cuzzocrea, Fan Jiang 0001, Wookey Lee, Carson K. Leung |
APWeb | 4 |
| 2014 | Mining Interesting "Following" Patterns from Social Networks
Fan Jiang 0001, Carson K. Leung |
DaWaK | 2 |
| 2014 | BLIMP: A Compact Tree Structure for Uncertain Frequent Pattern Mining
Carson K. Leung, Richard Kyle MacKinnon |
DaWaK | 1 |
| 2014 | Fast Algorithms for Frequent Itemset Mining from Uncertain DataabstractThe majority of existing data mining algorithms mine frequent item sets from precise databases. A well-known algorithm is FP-growth, which builds a compact FP-tree structure to capture important contents of the database and mines frequent item sets from the FP-tree. However, there are situations in which data are uncertain. In recent years, researchers have paid attention to frequent item set mining from uncertain databases. UFP-growth is one of the frequently cited algorithms for mining uncertain data. However, the corresponding UFP-tree structure can be large. Other tree structures for handling uncertain data may achieve compactness at the expense of looser upper bounds on expected supports. To solve this problem, we propose two compact tree structures which capture uncertain data with tighter upper bounds than existing tree structures. We also designed two algorithms that mine frequent item sets from our proposed trees. Our experimental results show the tightness of bounds to expected supports provided by these algorithms. Carson K. Leung, Richard Kyle MacKinnon |
ICDM | 1 |
| 2014 | A machine learning approach for stock price predictionabstractData mining and machine learning approaches can be incorporated into business intelligence (BI) systems to help users for decision support in many real-life applications. Here, in this paper, we propose a machine learning approach for BI applications. Specifically, we apply structural support vector machines (SSVMs) to perform classification on complex inputs such as the nodes of a graph structure. We connect collaborating companies in the information technology sector in a graph structure and use an SSVM to predict positive or negative movement in their stock prices. The complexity of the SSVM cutting plane optimization problem is determined by the complexity of the separation oracle. It is shown that (i) the separation oracle performs a task equivalent to maximum a posteriori (MAP) inference and (ii) a minimum graph cutting algorithm can solve this problem in the stock price case in polynomial time. Experimental results show the practicability of our proposed machine learning approach in predicting stock prices. Carson K. Leung, Richard Kyle MacKinnon, Yang Wang 0003 |
IDEAS | 1 |
| 2013 | Finding Diverse Friends in Social Networks
Syed Khairuzzaman Tanbeer, Carson K. Leung |
APWeb | 2 |
| 2013 | Mining Frequent Patterns from Uncertain Data with MapReduce for Big Data Analytics
Carson K. Leung, Yaroslav Hayduk |
DASFAA (1) | 1 |
| 2013 | Stream Mining of Frequent Patterns from Delayed Batches of Uncertain Data
Fan Jiang 0001, Carson K. Leung |
DaWaK | 2 |
| 2013 | Mining Frequent Patterns from Human Interactions in Meetings Using Directed Acyclic Graphs
Anna Fariha, Chowdhury Farhan Ahmed, Carson K. Leung, S. M. Abdullah, Longbing Cao |
PAKDD (1) | 3 |
| 2013 | PUF-Tree: A Compact Tree Structure for Frequent Pattern Mining of Uncertain Data
Carson K. Leung, Syed Khairuzzaman Tanbeer |
PAKDD (1) | 1 |
| 2013 | Mining Frequent Itemsets from Sparse Data Streams in Limited Memory Environments
Juan J. Cameron, Alfredo Cuzzocrea, Fan Jiang 0001, Carson K. Leung |
WAIM | 4 |
| 2012 | Fast Tree-Based Mining of Frequent Itemsets from Uncertain Data
Carson K. Leung, Syed Khairuzzaman Tanbeer |
DASFAA (1) | 1 |
| 2012 | Mining Popular Patterns from Transactional Databases
Carson K. Leung, Syed Khairuzzaman Tanbeer |
DaWaK | 1 |
| 2012 | Efficient Fuzzy Ranking for Keyword Search on Graphs
Nidhi R. Arora, Wookey Lee, Carson K. Leung |
DEXA (1) | 3 |
| 2012 | A constrained frequent pattern mining system for handling aggregate constraintsabstractFrequent pattern mining searches data for sets of items that are frequently co-occurring together. Most of algorithms find all the frequent patterns. However, there are many real-life situations in which users is interested in only some small portions of the entire collection of frequent patterns. To mine patterns that satisfy the user aggregate constraints in the form of agg(X.attr)θconst, properties of constraints are exploited. When agg is sum, the mining can be complicated. Existing mining systems or algorithms usually make assumptions about the value or range of X.attr and/or const. In this paper, we propose a frequent pattern mining system that avoids making these assumptions and that effectively handles the sum constraints as well as other aggregate constraints. Carson K. Leung, Fan Jiang 0001, Lijing Sun |
IDEAS | 1 |
| 2012 | Mining probabilistic datasets verticallyabstractAs frequent pattern mining plays an important role in various real-life applications, it has been the subject of numerous studies. Most of the studies mine transactional datasets of precise data. However, there are situations in which data are uncertain. Over the few years, Apriori-based, tree-based, and hyperlinked array structure based mining algorithms have been proposed to mine frequent patterns from these probabilistic datasets of uncertain data. These algorithms view the datasets "horizontally" as collections of transactions, and each records a set of items contained in that transaction. In this paper, we consider an alternative representation such that probabilistic datasets of uncertain data can be viewed "vertically" as collections of vectors. The vector for each item indicates which transactions contain that item. We also propose an algorithm called U-VIPER to mine these probabilistic datasets "vertically for frequent patterns. Carson K. Leung, Syed Khairuzzaman Tanbeer, Bhavek P. Budhia, Lauren C. Zacharias |
IDEAS | 1 |
| 2012 | RadialViz: An Orientation-Free Frequent Pattern Visualizer
Carson K. Leung, Fan Jiang 0001 |
PAKDD (2) | 1 |
| 2011 | Categorical Data Skyline Using Classification Tree
Wookey Lee, Justin JongSu Song, Carson K. Leung |
APWeb | 3 |
| 2011 | Frequent Pattern Mining from Time-Fading Streams of Uncertain Data
Carson K. Leung, Fan Jiang 0001 |
DaWaK | 1 |
| 2011 | A landmark-model based system for mining frequent patterns from uncertain data streamsabstractHuge volumes of streaming data have been generated by sensors for applications such as environment surveillance. Partially due to the inherited limitation of sensors, these continuous streaming data can be uncertain. Over the past few years, algorithms have been proposed to apply the sliding window or time-fading window model to mine frequent patterns from streams of uncertain data. However, there are also other models to process data streams. In this paper, we propose a landmark-model based system for mining frequent patterns from streams of uncertain data. Carson K. Leung, Fan Jiang 0001, Yaroslav Hayduk |
IDEAS | 1 |
| 2010 | uCFS2: an enhanced system that mines uncertain data for constrained frequent setsabstractFrequent set mining searches for sets of items that are frequently co-occurring together. Existing algorithms mainly find all the frequent sets from precise data. However, there are real-life situations in which users are interested in only some tiny portions of the entire collection of frequent sets and/or the data to be mined are uncertain. Recently, a tree-based system was proposed to mine uncertain data for frequent sets that satisfy user-specified succinct constraints. However, non-succinct constraints exist. In this paper, we extend such a system to mine uncertain data for frequent sets that satisfy succinct as well as non-succinct constraints by effectively exploiting properties of these constraints. Carson K. Leung, Dale A. Brajczuk |
IDEAS | 1 |
| 2009 | AnchorWoman: top-k structured mobile web search engineabstractWith advances in technology, mobile handheld devices-such as PDAs-have become very popular. In many real-life situations, users want to find structuring information using these mobile devices, which are convenient to use but have relatively limited resources. In this paper, we present a top-k structured mobile Web search engine. It uses a top-k adaptable search-tree method that utilizes hierarchical structure of hypermedia objects to effectively look for structuring information from the mobile Web model. The engine, which is implemented in the mobile environment, provides users with top-k adaptive Web search recommendations for mobile handheld devices. Wookey Lee, James Jung-Hoon Lee, Carson K. Leung |
CIKM | 4 |
| 2009 | Mining of Frequent Itemsets from Streams of Uncertain DataabstractFrequent itemset mining plays an essential role in the mining of various patterns and is in demand in many real-life applications. Hence, mining of frequent itemsets has been the subject of numerous studies since its introduction. Generally, most of these studies find frequent itemsets from traditional transaction databases, in which the content of each transaction--namely, items--is definitely known and precise. However, there are many real-life situations in which ones are uncertain about the content of transactions. This calls for the mining of uncertain data. Moreover, due to advances in technology, a flood of precise or uncertain data can be produced in many situations. This calls for the mining of data streams. To deal with these situations, we propose two tree-based mining algorithms to efficiently find frequent itemsets from streams of uncertain data, where each item in the transactions in the streams is associated with an existential probability. Experimental results show the effectiveness of our algorithms in mining frequent itemsets from streams of uncertain data. Carson K. Leung, Boyu Hao |
ICDE | 1 |
| 2009 | Mining uncertain data for constrained frequent setsabstractData mining aims to search for implicit, previously unknown, and potentially useful pieces of information---such as sets of items that are frequently co-occurring together---that are embedded in data. The mined frequent sets can be used in the discovery of correlation or casual relations, analysis of sequences, and formation of association rules. Since its introduction, frequent set mining has been the subject of numerous studies. Most of these studies find all the frequent sets from transaction databases of precise data, in which items within each transaction are definitely known and precise. However, there are many real-life situations in which the user is interested in only some tiny portions of the entire frequent sets, and there are also many situations in which data in the transaction databases are uncertain. This calls for both (i) constrained frequent set mining (which finds frequent sets that satisfy user constraints indicating the user interest) and (ii) frequent set mining from uncertain data. In this paper, we propose a tree-based system that integrates these two kinds of frequent set mining. The resulting mining system avoids candidate generation; it pushes the user constraints inside the mining process, which avoids unnecessary computation. Consequently, the system effectively mines from transaction databases of uncertain data for only those frequent sets satisfying the user-specified constraints. Carson K. Leung, Dale A. Brajczuk |
IDEAS | 1 |
| 2008 | WiFIsViz: Effective Visualization of Frequent ItemsetsabstractFrequent itemset mining plays an essential role in the mining of many different patterns. Most existing frequent itemset mining algorithms return the mined results--namely, frequent itemsets--in the form of textual lists. However, the use of visual representation can enhance the user understanding of the inherent relations in a collection of frequent itemsets. In this paper, we propose an effective visualizer, called WiFIsViz, to display the mined frequent itemsets. WiFIsViz provides users with an overview and details about the itemsets. Moreover, this visualizer is also equipped with several interactive features for effective visualization of the frequent itemsets mined from various real-life applications. Carson K. Leung, Pourang Irani, Christopher L. Carmichael |
ICDM | 1 |
| 2008 | Efficient algorithms for stream mining of constrained frequent patterns in a limited memory environmentabstractAs technology advances, streams of data can be rapidly generated in many real-life applications. This calls for stream mining, which searches for implicit, previously unknown, and potentially useful information---such as frequent patterns---that might be embedded in continuous data streams. However, most of the existing algorithms do not allow users to express the patterns to be mined according to their intentions, via the use of constraints. As a result, these unconstrained mining algorithms can yield numerous patterns that are not interesting to the users. Moreover, many existing tree-based algorithms assume that all the trees constructed during the mining process can fit into memory. While this assumption holds for many situations, there are many other situations in which it does not hold. Hence, in this paper, we develop efficient algorithms for stream mining of constrained frequent patterns in a limited memory environment. Our algorithms allow users to impose a certain focus on the mining process, discover from data streams all those frequent patterns that satisfy the user constraints, and handle situations where the available memory space is limited. Carson K. Leung, Dale A. Brajczuk, Jialiang Yu |
IDEAS | 1 |
| 2008 | FIsViz: A Frequent Itemset Visualizer
Carson K. Leung, Pourang Irani, Christopher L. Carmichael |
PAKDD | 1 |
| 2008 | A Tree-Based Approach for Frequent Pattern Mining from Uncertain Data
Carson K. Leung, Mark Anthony F. Mateo, Dale A. Brajczuk |
PAKDD | 1 |
| 2007 | An EffectiveMulti-Layer Model for Controlling the Quality of DataabstractData mining aims to search for implicit, previously unknown, and potentially useful information that might be embedded in the data. It is well known that "garbage in, garbage out". Hence, to get meaningful mining results, a clean set of data is essential. In this paper, we propose an effective model for controlling the quality of data. Specifically, this three-layer model focuses on data validity and data consistency. To elaborate, the internal layer ensures that the observed data are valid and their values fall within reasonable ranges. The temporal layer ensures that data are consistent with their temporal behaviour. The spatial layer ensures that data are consistent with their spatial neighbours. A case study on applying our proposed model to real-life weather data for an agricultural application shows that our model is effective in controlling and improving data quality, and thus leading to better mining results. It is important to note the application of our proposed model is not confined to the weather data for agricultural applications. We also discuss, in this paper, how the proposed three-layer model can be effectively applicable to control the quality of data in some other real-life situations. Carson K. Leung, Mark Anthony F. Mateo, Andrew J. Nadler |
IDEAS | 1 |
| 2007 | CanTree: a canonical-order tree for incremental frequent-pattern mining
Carson K. Leung, Quamrul I. Khan, Tariqul Hoque |
Knowl. Inf. Syst. | 1 |
| 2006 | DSTree: A Tree Structure for the Mining of Frequent Sets from Data StreamsabstractWith advances in technology, a flood of data can be produced in many applications such as sensor networks and Web click streams. This calls for efficient techniques for extracting useful information from streams of data. In this paper, we propose a novel tree structure, called DSTree (Data Stream Tree), that captures important data from the streams. By exploiting its nice properties, the DSTree can be easily maintained and mined for frequent itemsets as well as various other patterns like constrained itemsets. Carson K. Leung, Quamrul I. Khan |
ICDM | 1 |
| 2006 | Efficient Mining of Constrained Frequent Patterns from StreamsabstractWith advances in technology, a flood of data can be produced in many applications such as sensor networks and Web click streams. This calls for stream mining, which searches for implicit, previously unknown, and potentially useful information (such as frequent patterns) that might be embedded in continuous data streams. However, most of the existing algorithms do not allow users to express the patterns to be mined according to their intentions, via the use of constraints. Consequently, these unconstrained mining algorithms can yield numerous patterns that are not interesting to users. In this paper, we develop algorithms - which use a tree-based framework to capture the important portion of the streaming data, and allow human users to impose a certain focus on the mining process - for mining frequent patterns that satisfy user constraints from the flood of data Carson K. Leung, Quamrul I. Khan |
IDEAS | 1 |
| 2005 | CanTree: A Tree Structure for Efficient Incremental Mining of Frequent PatternsabstractSince its introduction, frequent-pattern mining has been the subject of numerous studies, including incremental updating. Many existing incremental mining algorithms are Apriori-based, which are not easily adoptable to FP-tree based frequent-pattern mining. In this paper, we propose a novel tree structure, called CanTree (canonical-order tree), that captures the content of the transaction database and orders tree nodes according to some canonical order. By exploiting its nice properties, the CanTree can be easily maintained when database transactions are inserted, deleted, and/or modified. For example, the CanTree does not require adjustment, merging, and/or splitting of tree nodes during maintenance. No rescan of the entire updated database or reconstruction of a new tree is needed for incremental updating. Experimental results show the effectiveness of our CanTree. Carson K. Leung, Quamrul I. Khan, Tariqul Hoque |
ICDM | 1 |
| 2004 | Interactive Constrained Frequent-Pattern Mining System
Carson K. Leung |
IDEAS | 1 |
| 2003 | Efficient dynamic mining of constrained frequent setsabstractData mining is supposed to be an iterative and exploratory process. In this context, we are working on a project with the overall objective of developing a practical computing environment for the human-centered exploratory mining of frequent sets. One critical component of such an environment is the support for the dynamic mining of constrained frequent sets of items. Constraints enable users to impose a certain focus on the mining process; dynamic means that, in the middle of the computation, users are able to (i) change (such as tighten or relax) the constraints and/or (ii) change the minimum support threshold, thus having a decisive influence on subsequent computations. In a real-life situation, the available buffer space may be limited, thus adding another complication to the problem.In this article, we develop an algorithm, called DCF, for Dynamic Constrained Frequent-set computation . This algorithm is enhanced with a few optimizations, exploiting a lightweight structure called a segment support map . It enables DCF to (i) obtain sharper bounds on the support of sets of items, and to (ii) better exploit properties of constraints. Furthermore, when handling dynamic changes to constraints, DCF relies on the concept of a delta member generating function , which generates precisely the sets of items that satisfy the new but not the old constraints. Our experimental results show the effectiveness of these enhancements. Laks V. S. Lakshmanan, Carson K. Leung, Raymond T. Ng |
ACM Trans. Database Syst. | 2 |
| 2002 | OSSM: A Segmentation Approach to Optimize Frequency CountingabstractComputing the frequency of a pattern is one of the key operations in data mining algorithms. We describe a simple yet powerful way of speeding up any form of frequency counting satisfying the monotonicity condition. Our method, the optimized segment support map (OSSM), is a light-weight structure which partitions the collection of transactions into m segments, so as to reduce the number of candidate patterns that require frequency counting. We study the following problems: (1) what is the optimal number of segments to be used; and (2) given a user-determined m, what is the best segmentation/composition of the m segments? For Problem 1, we provide a thorough analysis and a theorem establishing the minimum value of m for which there is no accuracy lost in using the OSSM. For Problem 2, we develop various algorithms and heuristics, which efficiently generate OSSMs that are compact and effective, to help facilitate segmentation. Carson K. Leung, Raymond T. Ng, Heikki Mannila |
ICDE | 1 |