Yuefeng Li 0001

dblp:74/4581-1 · DBLP profile ↗
← Back
118ranked-venue papers
18as first author
29since 2021 · last 2026
0000-0002-3594-8980ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 76 · 13 first-author · 20 since 2021Databases, data management, data science and information retrieval · 75 · 15 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Theory of computation · 1
YearPublicationVenuePosition
2026 Exploring Selective Avoidance for Online User Behavior Analysis: A Forest of Thought Explanation
abstract
The response behaviors observed in online user-generated content (UGC) frequently demonstrate non-linear characteristics, such as conditional branching and selective avoidance. These patterns present additional challenges for ensuring the trustworthiness of Large Language Model (LLMs) reasoning, particularly as their unidirectional, left-to-right inference mechanisms may not adequately capture such complex reasoning dynamics. To address this, we propose a Forest of Thought Explanation (FoTE), a novel prompting that models the selective avoidance in UGC while ensuring explanation consensus through reasoning paths across all decision sub-trees. FoTE firstly generates various reasoning paths through an adaptive CoT prompting. Each generated thought is subsequently evaluated through cooperative game theory to quantify its fair influence. The thoughts with the top-k contribution scores are preserved and randomly sampled to emulate selective avoidance for the next reasoning iteration. Through extensive evaluations across three open-source LLMs and two established social science problems (spanning four benchmark datasets), FoTE demonstrates superior success rates compared to competing prompting strategies. Notably, its performance gains increase with the strength of selective avoidance in social problems. The trustworthiness of our FoTE is enhanced by the incorporation of (1) a solid theoretical foundation and (2) a transparent reasoning path that converges toward consensus.
Lin Li 0001, Kaize Shi, Xiaohui Tao 0001, Jianwei Zhang 0002, Yuefeng Li 0001
AAAI6
2026 Balancing privacy and performance: An empirical study of machine unlearning in deep learning models
abstract
In the field of Artificial Intelligence (AI), the data used to train models may contain private information that could potentially be exposed in the model’s output. Machine Unlearning (MU) has emerged as a promising solution for removing private or obsolete data from trained models, along with their influence, thereby enforcing the “right to be forgotten” under the General Data Protection Regulation (GDPR). However, achieving a balance between privacy guarantee and model performance remains a fundamental challenge. This paper contributes to the field of AI by presenting an empirical evaluation of key families, i.e., data deletion, data perturbation, and model update of MU for privacy preservation, focusing on their impact on both classification accuracy and privacy in deep learning (DL) models. The study assesses changes in the classification performance of the convolutional neural network (CNN) architecture and the long short-term memory (LSTM) and bidirectional LSTM (Bi-LSTM) recurrent neural network (RNN) architectures when used with data deletion, data perturbation, and model update families of MU. This study also assesses these architectures’ susceptibility to membership inference attacks (MIA) before and after unlearning on PPG-DaLiA and MHEALTH (Mobile HEALTH) datasets, providing a quantitative measure of privacy leakage. Experimental results show that model update techniques offer more scalable alternatives to data deletion and perturbation, though they introduce varying levels of privacy leakage risk. In doing so, this research highlights the strengths and limitations of current targeted unlearning methods and underscores the need for more efficient and flexible approaches to privacy protection in DL models.
Tazeem Ahmad, Xiaohui Tao 0001, Jianming Yong, Thanveer Shaik, Haoran Xie 0001, Yuefeng Li 0001, U. Rajendra Acharya
Eng. Appl. Artif. Intell.6
2026 Discover Class-Based Feature Distribution by Encoding Discrete Data for Classification
abstract
ABSTRACT The self‐organisation map is an unsupervised learning technique that discovers patterns and relationships in data without requiring labelled training data. Inspired by the self‐organisation map, Self‐Organised Granular encoding has been shown to be effective for generating reliable discrete data clustering results as it is a data encoding technique that uses fuzzy sets and granularity to handle uncertain and imprecise information within discrete data. However, it is mainly useful for unsupervised learning, and its feasibility for supervised learning has not yet been studied. Also, discrete data classification is still under‐researched. This paper proposes a new discrete data classification method called Transposed Fuzzy Class Granular classification. This method aims to transform discrete data into fuzzy partitions by considering all available classes and generating representations of the trained class's Transposed Fuzzy Class Granular distribution by measuring the total divergence from the average of each fuzzy class's membership degree distribution. The paper introduces a novel approach to discrete data classification by adapting class granules for classification and improving performance by tackling uncertainty, ambiguity, and the unique characteristics of discrete datasets. The study examined seven discrete datasets and compared their performance with eight commonly used classifiers as the baseline. These datasets were naturally discrete or created by discrete partitions of real datasets. The experimental results demonstrate that the proposed classifier outperforms the baseline classifiers in discrete data classification.
Yuefeng Li 0001, Xiaohui Tao 0001, Jianming Yong
Expert Syst. J. Knowl. Eng.2
2026 Overcoming BERT's limitations in uncertainty: A novel two-stage solution for multi-class medical text classification
abstract
• Introduced a principled definition of an uncertain decision boundary for BERT-based classifiers using a discriminative confidence margin, enabling systematic identification of ambiguous predictions. • Integrated word–class probability embeddings (WCPE) to enrich document representations with explicit class-discriminative signals, particularly effective for uncertain and noisy medical texts. • Proposed a two-stage, uncertainty-aware decision framework that routes confident samples to a baseline BERT model and uncertain samples to a WCPE-enhanced variant, improving robustness under ambiguity. • Developed a data-driven strategy for optimizing the uncertainty threshold (UncertainT), balancing coverage of ambiguous samples and overall classification accuracy. • Established the general applicability of the proposed uncertainty-aware decision framework across multiple datasets and pretrained backbones, while revealing that statistically significant improvements emerge selectively depending on dataset ambiguity and backbone pretraining. Bidirectional Encoder Representations from Transformers (BERT) has achieved state-of-the-art performance in Natural Language Processing (NLP) tasks but struggles with uncertainty management in multi-class medical text classification, where overlapping categories and domain-specific critical terms pose challenges. BERT’s attention mechanism may distribute focus across irrelevant contextual patterns, reducing its ability to prioritize medically significant terms. To address these limitations, we propose a two-stage decision-making framework incorporating an “uncertain boundary” mechanism to separate high-confidence (“certain”) and low-confidence (“uncertain”) cases based on an optimized uncertainty threshold. A baseline BERT model processes high-confidence cases, while uncertain cases are handled by an enhanced BERT model integrating Word-Class Probabilistic Embedding (WCPE) to improve domain-specific representation learning. We evaluate our framework across three datasets-Ohsumed, PubMed, and Drug Review and further examine its robustness under multiple pretrained language model settings. Using standard BERT as the primary backbone, results show consistent accuracy improvements from 0.770 to 0.773 on Ohsumed ( p < 0.05), 0.973 to 0.975 on PubMed ( p < 0.1), and 0.596 to 0.609 on Drug Review ( p < 0.05) with stronger gains on uncertain subsets. These findings show that confidence-based partitioning and model specialization enhance BERT’s classification performance while effectively handling uncertainty in medical NLP.
Prabhashrini Dhanushika Manage, Yutong Wu 0001, Jinglan Zhang, Anupama Udayangani Gunathilaka, Yuefeng Li 0001
Knowl. Based Syst.5
2025 BERT-TCPE: A Term-Class Probability-Enhanced BERT Framework for Uncertainty-Aware Medical Text Classification
Prabhashrini Dhanushika Manage, Yutong Wu 0001, Jinglan Zhang, Yuefeng Li 0001
PRICAI (4)4
2024 Aligning Bytes with Bliss: Integrating Happiness Computing with Sociological Insight
Xiaojun Wu 0001, Lin Li 0001, Xiaohui Tao 0001, Yuefeng Li 0001
ADMA (1)4
2024 Refined Sentiment Analysis Using POS Features and LDA: Mitigating Polysemy and Sparsity with BERT Contextual Embedding
Thennakoon Mudiyanselage Anupama Udayangani Gunathilaka, Yuefeng Li 0001, Jinglan Zhang, Prabhashrini Dhanushika Manage
ICONIP (6)2
2024 Granule-specific feature selection for continuous data classification using neighborhood rough sets
Mahawaga Arachchige Nayomi Dulanjala Sewwandi, Yuefeng Li 0001, Jinglan Zhang
Expert Syst. Appl.2
2024 k-outlier removal based on contextual label information and cluster purity for continuous data classification
abstract
Outlier detection is extensively adopted in machine learning applications to identify rare but significant objects in a data distribution that deviate from the majority of objects. However, none of the existing outlier definitions use the available contextual label information in a dataset to improve the identification of outliers, though it could improve the performance of data analysis. In this study, we propose a novel definition for outliers considering the available contextual label information along with a method to remove outliers based on the proposed definition to improve classification performance and granulation purity. The experimental results on eight public datasets show that the removal of outliers using the novel method highly improves the classification accuracy with Support Vector Machine, k-Nearest Neighbors, and Classification and Regression Tree algorithms while identifying 100% pure subclasses of the labelled major classes. This method is beneficial in identifying the outliers automatically when the data contains classification information about a specific application instead of outlierness information.
Mahawaga Arachchige Nayomi Dulanjala Sewwandi, Yuefeng Li 0001, Jinglan Zhang
Expert Syst. Appl.2
2024 WSNMF: Weighted Symmetric Nonnegative Matrix Factorization for attributed graph clustering
Kamal Berahmand, Mehrnoush Mohammadi, Razieh Sheikhpour, Yuefeng Li 0001, Yue Xu 0001
Neurocomputing4
2024 A Deep Semi-Supervised Community Detection Based on Point-Wise Mutual Information
abstract
Network clustering is one of the fundamental unsupervised methods of knowledge discovery. Its goal is to group similar nodes together without supervision or prior knowledge of the nature of the clusters. Among various clustering methods, semi-supervised clustering detection is one of the most promising approaches for community detection because of its ability to employ side information to better understand network topology. However, most of the previous work faces two problems: the use of linear methods to reduce dimensionality and the random selection of side information, and as a result of these two drawbacks, semi-supervised community detection methods are less efficient. To fill these gaps, we developed an end-to-end deep semi-supervisor community detection (DSSC) for complex networks. A new learning objective is designed that uses a semi-autoencoder (SeAE) with a defined pair-wise constraint matrix based on point-wise mutual information (PMI) in the representation layer to accurately learn distinctive features and, in the clustering layer, adds a pair-wise constraint as a term to minimize distance within the cluster while the distance between clusters increases. The results show that our method performs unexpectedly well in comparison to the existing state-of-the-art community detection methods in complex networks.
Kamal Berahmand, Yuefeng Li 0001, Yue Xu 0001
IEEE Trans. Comput. Soc. Syst.2
2024 A Semantics-enhanced Topic Modelling Technique: Semantic-LDA
abstract
Topic modelling is a beneficial technique used to discover latent topics in text collections. But to correctly understand the text content and generate a meaningful topic list, semantics are important. By ignoring semantics, that is, not attempting to grasp the meaning of the words, most of the existing topic modelling approaches can generate some meaningless topic words. Even existing semantic-based approaches usually interpret the meanings of words without considering the context and related words. In this article, we introduce a semantic-based topic model called semantic-LDA that captures the semantics of words in a text collection using concepts from an external ontology. A new method is introduced to identify and quantify the concept–word relationships based on matching words from the input text collection with concepts from an ontology without using pre-calculated values from the ontology that quantify the relationships between the words and concepts. These pre-calculated values may not reflect the actual relationships between words and concepts for the input collection, because they are derived from datasets used to build the ontology rather than from the input collection itself. Instead, quantifying the relationship based on the word distribution in the input collection is more realistic and beneficial in the semantic capture process. Furthermore, an ambiguity handling mechanism is introduced to interpret the unmatched words, that is, words for which there are no matching concepts in the ontology. Thus, this article makes a significant contribution by introducing a semantic-based topic model that calculates the word–concept relationships directly from the input text collection. The proposed semantic-based topic model and an enhanced version with the disambiguation mechanism were evaluated against a set of state-of-the-art systems, and our approaches outperformed the baseline systems in both topic quality and information filtering evaluations.
Dakshi T. K. Geeganage, Yue Xu 0001, Yuefeng Li 0001
ACM Trans. Knowl. Discov. Data3
2024 SDAC-DA: Semi-Supervised Deep Attributed Clustering Using Dual Autoencoder
abstract
Attributed graph clustering aims to group nodes into disjoint categories using deep learning to represent node embeddings and has shown promising performance across various applications. However, two main challenges hinder further performance improvement. Firstly, reliance on unsupervised methods impedes the learning of low-dimensional, clustering-specific features in the representation layer, thus impacting clustering performance. Secondly, the predominant use of separate approaches leads to suboptimal learned embeddings that are insufficient for subsequent clustering steps. To address these limitations, we propose a novel method called Semi-supervised Deep Attributed Clustering using Dual Autoencoder (SDAC-DA). This approach enables semi-supervised deep end-to-end clustering in attributed networks, promoting high structural cohesiveness and attribute homogeneity. SDAC-DA transforms the attribute network into a dual-view network, applies a semi-supervised autoencoder layering approach to each view, and integrates dimensionality reduction matrices by considering complementary views. The resulting representation layer contains high clustering-friendly embeddings, which are optimized through a unified end-to-end clustering process for effectively identifying clusters. Extensive experiments on both synthetic and real networks demonstrate the superiority of our proposed method over seven state-of-the-art approaches.
Kamal Berahmand, Sondos Bahadori, Maryam Nooraei Abadeh, Yuefeng Li 0001, Yue Xu 0001
IEEE Trans. Knowl. Data Eng.4
2024 Supervised fusion content-based framework for breakdown detection in task-oriented conversational systems
abstract
Conversational agents (CAs) have been widely used for many domains, such as healthcare, education, and business. One main category of CAs is task-oriented CAs, which aim to help users to complete a set of specific tasks. However, task-oriented CAs can fail to answer the user’s question, which can lead to a breakdown in the dialogue (when it is not possible to complete a conversation with a CA). Breakdown detection is an essential task for developing better CAs. Several related studies have focused on breakdown detection using different sets of features, for example, topic transition, word-based similarity and clustering; but, the existing studies develop features mainly from the system’s outputs or user’s inputs, whereas the features can be extracted from both sides, as well as from the interaction between them. Therefore, in this work, we developed a new supervised fusion machine learning (ML) model that combines the prediction from two machine learning algorithms for breakdown detection CAs services system. We developed features from different groups focusing on both the user input and the system response. Then we select the optimal combined features. The features are based on sentence similarity, sentiment features, and count-based features. The developed fusion model is mainly based on the two best performances of the single classifiers (SVM and RF). We explore several single ML algorithms using different sets of features and the combined features. To verify the effectiveness of the proposed fusion model, we compared the proposed models against baseline methods using four sets of data. We conclude that the proposed fusion model with the combined features outperforms the baselines and all other models in terms of prediction accuracy and f-score measures.
Mohammed Aldahash, Yuefeng Li 0001, Yue Xu 0001
Web Intell.2
2023 Gynecological cancer prognosis using machine learning techniques: A systematic review of the last three decades (1990-2022)
Joshua Sheehy, Hamish Rutledge, U. Rajendra Acharya, Hui Wen Loh, Raj Gururajan, Xiaohui Tao 0001, Xujuan Zhou, Yuefeng Li 0001, Tiana Gurney, Srinivas Kondalsamy-Chennakesavan
Artif. Intell. Medicine8
2023 A new method for recommendation based on embedding spectral clustering in heterogeneous networks (RESCHet)
Saman Forouzandeh, Kamal Berahmand, Razieh Sheikhpour, Yuefeng Li 0001
Expert Syst. Appl.4
2023 Robust graph regularization nonnegative matrix factorization for link prediction in attributed networks
Elahe Nasiri, Kamal Berahmand, Yuefeng Li 0001
Multim. Tools Appl.3
2023 DAC-HPP: deep attributed clustering with high-order proximity preserve
abstract
Abstract Attributed graph clustering, the task of grouping nodes into communities using both graph structure and node attributes, is a fundamental problem in graph analysis. Recent approaches have utilized deep learning for node embedding followed by conventional clustering methods. However, these methods often suffer from the limitations of relying on the original network structure, which may be inadequate for clustering due to sparsity and noise, and using separate approaches that yield suboptimal embeddings for clustering. To address these limitations, we propose a novel method called Deep Attributed Clustering with High-order Proximity Preserve (DAC-HPP) for attributed graph clustering. DAC-HPP leverages an end-to-end deep clustering framework that integrates high-order proximities and fosters structural cohesiveness and attribute homogeneity. We introduce a modified Random Walk with Restart that captures k-order structural and attribute information, enabling the modelling of interactions between network structure and high-order proximities. A consensus matrix representation is constructed by combining diverse proximity measures, and a deep joint clustering approach is employed to leverage the complementary strengths of embedding and clustering. In summary, DAC-HPP offers a unique solution for attributed graph clustering by incorporating high-order proximities and employing an end-to-end deep clustering framework. Extensive experiments demonstrate its effectiveness, showcasing its superiority over existing methods. Evaluation on synthetic and real networks demonstrates that DAC-HPP outperforms seven state-of-the-art approaches, confirming its potential for advancing attributed graph clustering research.
Kamal Berahmand, Yuefeng Li 0001, Yue Xu 0001
Neural Comput. Appl.2
2022 Enhanced Topic Representation by Ambiguity Handling
Dakshi T. K. Geeganage, Yue Xu 0001, Darshika N. Koggalahewa, Yuefeng Li 0001
WISE4
2022 Integration of fuzzy logic and a convolutional neural network in three-way decision-making
L. D. C. S. Subhashini, Yuefeng Li 0001, Jinglan Zhang, Ajantha S. Atukorale
Expert Syst. Appl.2
2022 Assessing the effectiveness of a three-way decision-making framework with multiple features in simulating human judgement of opinion classification
L. D. C. S. Subhashini, Yuefeng Li 0001, Jinglan Zhang, Ajantha S. Atukorale
Inf. Process. Manag.2
2022 Integration of semantic patterns and fuzzy concepts to reduce the boundary region in three-way decision-making
L. D. C. S. Subhashini, Yuefeng Li 0001, Jinglan Zhang, Ajantha S. Atukorale
Inf. Sci.2
2022 Graph-based multi-label disease prediction model learning from medical data and domain knowledge
Thuan Pham, Xiaohui Tao 0001, Ji Zhang 0001, Jianming Yong, Yuefeng Li 0001, Haoran Xie 0001
Knowl. Based Syst.5
2022 FedStack: Personalized activity monitoring using stacked federated learning
abstract
Recent advances in remote patient monitoring (RPM) systems can recognize various human activities to measure vital signs, including subtle motions from superficial vessels. There is a growing interest in applying artificial intelligence (AI) to this area of healthcare by addressing known limitations and challenges such as predicting and classifying vital signs and physical movements, which are considered crucial tasks. Federated learning is a relatively new AI technique designed to enhance data privacy by decentralizing traditional machine learning modeling. However, traditional federated learning requires identical architectural models to be trained across the local clients and global servers. This limits global model architecture due to the lack of local models’ heterogeneity. To overcome this, a novel federated learning architecture, FedStack, which supports ensembling heterogeneous architectural client models was proposed in this study. This work offers a protected privacy system for hospitalized in-patients in a decentralized approach and identifies optimum sensor placement. The proposed architecture was applied to a mobile health sensor benchmark dataset from 10 different subjects to classify 12 routine activities. Three AI models, artificial neural network (ANN), convolutional neural network (CNN), and bidirectional long short-term memory (Bi-LSTM) were trained on individual subject data. The federated learning architecture was applied to these models to build local and global models capable of state-of-the-art performances. The local CNN model outperformed ANN and Bi-LSTM models on each subject data. Our proposed work has demonstrated better performance for heterogeneous stacking of the local models compared to homogeneous stacking. Further analysis of the global heterogeneous CNN model determined that the optimum placement of the sensors on human limbs resulted in better activity recognition. This work sets the stage to build an enhanced RPM system that incorporates client privacy to assist with clinical observations for patients in an acute mental health facility and ultimately help to prevent unexpected death.
Thanveer Shaik, Xiaohui Tao 0001, Niall Higgins, Raj Gururajan, Yuefeng Li 0001, Xujuan Zhou, U. Rajendra Acharya
Knowl. Based Syst.5
2022 Application of CycleGAN and transfer learning techniques for automated detection of COVID-19 using X-ray images
Ghazal Bargshady, Xujuan Zhou, Prabal Datta Barua, Raj Gururajan, Yuefeng Li 0001, U. Rajendra Acharya
Pattern Recognit. Lett.5
2021 Emerging Applications in Healthcare and Their Implications to Academia and Practice
Raj Gururajan, Xiaohui Tao 0001, Yuefeng Li 0001, Xujuan Zhou, Soman Elangovan, Srinivas Kondalsamy-Chennakesavan, Revathi Venkataraman
WISE (2)3
2021 Using back-and-forth translation to create artificial augmented textual data for sentiment analysis models
abstract
Sentiment analysis classification models trained using neural networks require large amounts of data, but collecting these datasets requires significant time and resources. Although artificial data has been used successfully in computer vision, there are few effective and generalizable methods for creating artificial augmented text data. In this paper, a text based data augmentation method is proposed called back-and-forth translation that can be used to artificially increase the size of any natural language dataset. By creating augmented text data and adding it to the original dataset, it is demonstrated by empirical experiments that back-and-forth translation data augmentation can reduce the error rate in binary sentiment classification models by up to 3.4%. These results are shown to be statistically significant.
Thomas Body, Xiaohui Tao 0001, Yuefeng Li 0001, Lin Li 0001, Ning Zhong 0001
Expert Syst. Appl.3
2021 Predicting Alzheimer's Disease from Spoken and Written Language Using Fusion-Based Stacked Generalization
Ahmed H. Alkenani, Yuefeng Li 0001, Yue Xu 0001, Qing Zhang 0001
J. Biomed. Informatics2
2021 Semantic-based topic representation using frequent semantic patterns
Dakshi T. K. Geeganage, Yue Xu 0001, Yuefeng Li 0001
Knowl. Based Syst.3
2020 Review selection based on content quality
Nan Tian, Yue Xu 0001, Yuefeng Li 0001
Knowl. Inf. Syst.3
2020 A survey on text classification and its applications
abstract
Text classification (a.k.a text categorisation) is an effective and efficient technology for information organisation and management. With the explosion of information resources on the Web and corporate intranets continues to increase, it has being become more and more important and has attracted wide attention from many different research fields. In the literature, many feature selection methods and classification algorithms have been proposed. It also has important applications in the real world. However, the dramatic increase in the availability of massive text data from various sources is creating a number of issues and challenges for text classification such as scalability issues. The purpose of this report is to give an overview of existing text classification technologies for building more reliable text classification applications, to propose a research direction for addressing the challenging problems in text mining.
Xujuan Zhou, Raj Gururajan, Yuefeng Li 0001, Revathi Venkataraman, Xiaohui Tao 0001, Ghazal Bargshady, Prabal Datta Barua, Srinivas Kondalsamy-Chennakesavan
Web Intell.3
2020 Query-based unsupervised learning for improving social media search
Khaled Albishre, Yuefeng Li 0001, Yue Xu 0001
World Wide Web2
2019 Semi-supervised text classification with deep convolutional neural network using feature fusion approach
abstract
Supervised learning algorithms employ labeled training data for classification purposes while obtaining labeled data for large datasets is costly and time consuming. Semi-supervised learning algorithms, on the contrary, use a small set of labeled data and a large set of unlabeled data to improve predication performance and thus may be a good alternative to supervised learning algorithms for large text datasets. Although many semi-supervised learning algorithms have been proposed in the data science literature, most of these algorithms are not feasible for discrete and unstructured text data.
Parvaneh Shayegh, Yuefeng Li 0001, Jinglan Zhang, Qing Zhang 0001
WI2
2019 Dual pattern-enhanced representations model for query-focused multi-document summarisation
Yutong Wu 0001, Yuefeng Li 0001, Yue Xu 0001
Knowl. Based Syst.2
2018 Enhancing Binary Classification by Modeling Uncertain Boundary in Three-Way Decisions (Extended Abstract)
abstract
Text classification techniques are playing a crucial role in identifying relevant texts from a large data set, e.g., various online crimes such as Cyberbullying, terrorist recruiting, propaganda or attack planning. Until now, supervised deep learning has brought about breakthroughs in processing multimedia data; however, there was no good practical way to harvest this opportunity for text classification because acquiring and maintaining a massive amount of training examples are too expensive for a large number of categories (e.g., Yahoo! taxonomy contains nearly 300,000 categories and the Library of Congress Subject Headings (LCSH) contains 394,070 subjects). Therefore, the question of how to effectively learn from sparse or small set of training examples is crucial for the true success of text classification. Semi-supervised approaches have been proposed for this challenge, which usually use a pair or several existing classifiers to extend a small training set. However, extracted pseudo training samples are uncertain because they are determined by a machine rather than people. Also, the massive volume and high variability of text data are creating a number of challenging issues such as the scalability and complicated relations between words. There are two fundamental issues with regards to the performance of existing classifiers: overlook and overload. Overlook means that some objects relevant to a class have been omitted, whereas overload means that some objects assigned to a class are actually not relevant to that class. The two issues are even more serious in the following two cases: (1) large uncertain boundary - the decision boundary between two classes includes many mixed examples (e.g., relevant and nonrelevant documents together), and (2) unbalanced classes - one class (e.g., information about terrorist attacks) is much smaller than another class (e.g., normal descriptions). We propose a three-way decision model [1] for dealing with the uncertain boundary for improving text classification performance based on rough set techniques and centroid solution. It aims to understand the uncertain boundary through partitioning the training samples into three regions (the positive, boundary and negative regions) by two main boundary vectors created from the labeled positive and negative training subsets, respectively, and further resolve the objects in the boundary region by two derived boundary vectors produced according to the structure of the boundary region. Four decision rules are proposed from the training process and applied to the incoming documents for more precise classification. The experimental results on the standard data sets RCV1 and Reuters-21578 show that the usage of boundary vectors is very effective and efficient for dealing with uncertainties of the decision boundary, and the proposed model has significantly improved the performance of binary text classification in terms of F1 measure and AUC area compared with six other popular baseline models.
Yuefeng Li 0001, Libiao Zhang, Yue Xu 0001, Yiyu Yao, Raymond Y. K. Lau, Yutong Wu 0001
ICDE1
2018 Query-Based Automatic Training Set Selection for Microblog Retrieval
Khaled Albishre, Yuefeng Li 0001, Yue Xu 0001
PAKDD (2)2
2018 An Extended Random-Sets Model for Fusion-Based Text Feature Selection
Abdullah Semran Alharbi, Yuefeng Li 0001, Yue Xu 0001
PAKDD (3)2
2018 A Semantic Similarity Based Topic Evaluation for Enhancing Information Filtering
abstract
Topic Modelling has been applied in many successful applications in data mining, text mining, machine learning and information filtering. The limitation is that the quality of topics generated from modelled corpus are not always good because many topics contain intrusive and ambiguous words. This negative drawback would affect the performance of text based application systems based on topic models. Hence, topic evaluation to assess and to rank the topics is really important for the good quality topics before applying those topics to text based applications. In this study, we proposed an ontology-based topic evaluation method for enhancing information filtering, named as STRbTCM. This new model assesses the quality of topics by matching topic models with headings in Library Congress Subject Heading (LCSH) ontology. To evaluate the effectiveness of our proposed model, we compare the model with two existing topic evaluation methods applied to information filtering system. In addition, we also compare our proposed model to term-based model BM25 and two other models based on topics: TNG and LDA_words. Through extensive experiments, we find that our proposed model performed better than other baseline models according to four main evaluating measures.
Hanh Nguyen, Yue Xu 0001, Yuefeng Li 0001
WI3
2018 Investigation of the Quality of Topic Models for Noisy Data Sources
abstract
Latent Dirichlet Allocation (LDA) has become the most stable and widely used topic model to derive topics from collections of documents where it depicts different levels of success based on diversified domains of inputs. Nevertheless, it is a vital requirement to evaluate the LDA against the quality of the input. The noise and uncertainty of the content create a negative influence on the topic model. The major contribution of this investigation is to critically evaluate the LDA based on the quality of input sources and human perception. The empirical study shows the relationship between the quality of the input and the accuracy of the output generated by LDA. Perplexity and coherence have been evaluated with three data-sets (RCV1, conference data set, tweets) which contain different level of complexities and uncertainty in their contents. Human perception in generating topics has been compared with the LDA in terms of human defined topics. Results of the analysis demonstrate a strong relationship between the quality of the input and generated topics. Thus, highly relevant topics were generated from formally written contents while noisy and messy contents lead to generate meaningless topics. A considerable gap is noticed between human defined topics and LDA generated topics. Finally, a concept-based topic modeling technique is proposed to improve the quality of topics by capturing the meaning of the content and eliminating the irrelevant and meaningless topics.
Yue Xu 0001, Yuefeng Li 0001, Dakshi T. K. Geeganage
WI2
2018 Interpretation of text patterns
Md. Abul Bashar, Yuefeng Li 0001
Data Min. Knowl. Discov.2
2017 Topical term weighting based on extended random sets for relevance feature selection
abstract
It is challenging to discover relevant features from long documents that describe user information needs due to the nature of text where synonymy, polysemy noise, and high dimensionality are inherited problems. Traditional feature selection methods could not effectively deal with these problems, because they assume that documents describe one topic only. Topic-based techniques, such as Latent Dirichlet Allocation (LDA), relax this assumption. They have been developed on the basis that a document can exhibit multiple hidden topics. However, LDA does not show encouraging results in selecting relevant features, because LDA calculates the weight of terms based on their local documents and does not generalise it globally at the collection level. So as to address this problem, we propose an innovative and effective extended random set model to generalise LDA weight for local document terms. The model is used as a weighting scheme for topical terms. It can assign a more discriminately accurate weight to these terms based on their appearance in LDA topics and relevant documents. The experimental results, based on the standard RCV1 dataset, TREC topics, and five standard performance measures, show that the proposed model significantly outperforms eight state-of-the-art baseline models in information filtering.
Abdullah Semran Alharbi, Yuefeng Li 0001, Yue Xu 0001
WI2
2017 Conceptual annotation of text patterns
abstract
Abstract Patterns are used as a fundamental means for analyzing data in many data mining applications. Many efficient techniques have been developed to discover patterns. However, the excessive number of discovered patterns and the lack of semantic information have made it difficult for a user to interpret and explore the patterns. A rough idea of the meanings of patterns can benefit the user in the process of exploring them. To address this issue, this paper presents a model for automatically annotating patterns with concepts. In addition, in a given context, the relative importance of each term that defines a concept is not the same. To define a context, there are a number of related information sources, such as documents, patterns, concepts, and an ontology. The question is which information sources are useful for estimating the relative importance of the terms? Should the most accurate one to be focused on or all of them be used to define the context? This research investigated these questions and defined an effective annotation context to estimate the relative importance of the terms, where the aim is to improve the performance of a machine that relies on the subject matter of a pattern set. The model is evaluated by comparing it with different baseline models on 2 standard datasets. The results show that the performance of the proposed model is significantly better.
Md. Abul Bashar, Yuefeng Li 0001, Yang Gao 0016
Comput. Intell.2
2017 Finding Semantically Valid and Relevant Topics by Association-Based Topic Selection Model
abstract
Topic modelling methods such as Latent Dirichlet Allocation (LDA) have been successfully applied to various fields, since these methods can effectively characterize document collections by using a mixture of semantically rich topics. So far, many models have been proposed. However, the existing models typically outperform on full analysis on the whole collection to find all topics but difficult to capture coherent and specifically meaningful topic representations. Furthermore, it is very challenging to incorporate user preferences into existing topic modelling methods to extract relevant topics. To address these problems, we develop a novel personalized Association-based Topic Selection (ATS) model, which can identify semantically valid and relevant topics from a set of raw topics based on the semantical relatedness between users’ preferences and the structured patterns captured in topics. The advantage of the proposed ATS model is that it enables an interactive topic modelling process driven by users’ specific interests. Based on three benchmark datasets, namely, RCV1, R8, and WT10G under the context of information filtering (IF) and information retrieval (IR), our rigorous experiments show that the proposed ATS model can effectively identify relevant topics with respect to users’ specific interests, and hence to improve the performance of IF and IR.
Yang Gao 0016, Yuefeng Li 0001, Raymond Y. K. Lau, Yue Xu 0001, Md. Abul Bashar
ACM Trans. Intell. Syst. Technol.2
2017 Enhancing Binary Classification by Modeling Uncertain Boundary in Three-Way Decisions
abstract
Text classification is a process of classifying documents into predefined categories through different classifiers learned from labelled or unlabelled training samples. Many researchers who work on binary text classification attempt to find a more effective way to separate relevant texts from a large data set. However, current text classifiers cannot unambiguously describe the decision boundary between positive and negative objects because of uncertainties caused by text feature selection and the knowledge learning process. This paper proposes a three-way decision model for dealing with the uncertain boundary to improve the binary text classification performance based on therough settechniques and centroid solution. It aims to understand the uncertain boundary through partitioning the training samples into three regions (the positive, boundary, and negative regions) by two main boundary vectors$\vec{C_{P}}$and$\vec{C_{N}}$, created from the labeled positive and negative training subsets, respectively, and further resolve the objects in the boundary region by two derived boundary vectors$\vec{B_{P}}$and$\vec{B_{N}}$, produced according to the structure of the boundary region. It involves an indirect strategy which is composed of two successive steps in the whole classification process: ‘two-way to three-way’ and ‘three-way to two-way’. Four decision rules are proposed from the training process and applied to the incoming documents for more precise classification. A large number of experiments have been conducted based on the standard data sets RCV1 and Reuters-21578. The experimental results show that the usage of boundary vectors is very effective and efficient for dealing with uncertainties of the decision boundary, and the proposed model has significantly improved the performance of binary text classification in terms of$F_{1}$measure and$AUC$area compared with six other popular baseline models.
Yuefeng Li 0001, Libiao Zhang, Yue Xu 0001, Yiyu Yao, Raymond Y. K. Lau, Yutong Wu 0001
IEEE Trans. Knowl. Data Eng.1
2016 Specialized Review Selection Using Topic Models
Nan Tian, Yue Xu 0001, Yuefeng Li 0001
PKAW4
2016 A Framework for Automatic Personalised Ontology Learning
abstract
Understanding or acquiring a user's information needs from their local information repository (e.g. a set of example-documents that are relevant to user information needs) is important in many applications. However, acquiring the user's information needs from the local information repository is very challenging. Personalised ontology is emerging as a powerful tool to acquire the information needs of users. However, its manual or semi-automatic construction is expensive and time-consuming. To address this problem, this paper proposes a model to automatically learn personalised ontology by labelling topic models with concepts, where the topic models are discovered from a user's local information repository. The proposed model is evaluated by comparing against ten baseline models on the standard dataset RCV1 and a large ontology LCSH. The results show that the model is effective and its performance is significantly improved.
Md. Abul Bashar, Yuefeng Li 0001, Yang Gao 0016
WI2
2016 Mining Topically Coherent Patterns for Unsupervised Extractive Multi-document Summarization
abstract
Addressing the problem of information overload, automatic multi-document summarization (MDS) has been widely utilized in the various real-world applications. Most of existing approaches adopt term-based representation for documents which limit the performance of MDS systems. In this paper, we proposed a novel unsupervised pattern-enhanced topic model (PETMSum) for the MDS task. PETMSum combining pattern mining techniques with LDA topic modelling could generate discriminative and semantic rich representations for topics and documents so that the most representative, non-redundant, and topically coherent sentences can be selected automatically to form a succinct and informative summary. Extensive experiments are conducted on the data of document understanding conference (DUC) 2006 and 2007. The results prove the effectiveness and efficiency of our proposed approach.
Yutong Wu 0001, Yuefeng Li 0001, Yue Xu 0001
WI2
2015 Pattern-based Topics for Document Modelling in Information Filtering
abstract
Many mature term-based or pattern-based approaches have been used in the field of information filtering to generate users' information needs from a collection of documents. A fundamental assumption for these approaches is that the documents in the collection are all about one topic. However, in reality users' interests can be diverse and the documents in the collection often involve multiple topics. Topic modelling, such as Latent Dirichlet Allocation (LDA), was proposed to generate statistical models to represent multiple topics in a collection of documents, and this has been widely utilized in the fields of machine learning and information retrieval, etc. But its effectiveness in information filtering has not been so well explored. Patterns are always thought to be more discriminative than single terms for describing documents. However, the enormous amount of discovered patterns hinder them from being effectively and efficiently used in real applications, therefore, selection of the most discriminative and representative patterns from the huge amount of discovered patterns becomes crucial. To deal with the above mentioned limitations and problems, in this paper, a novel information filtering model, Maximum matched Pattern-based Topic Model (MPBTM), is proposed. The main distinctive features of the proposed model include: (1) user information needs are generated in terms of multiple topics; (2) each topic is represented by patterns; (3) patterns are generated from topic models and are organized in terms of their statistical and taxonomic features; and (4) the most discriminative and representative patterns, called Maximum Matched Patterns, are proposed to estimate the document relevance to the user's information needs in order to filter out irrelevant documents. Extensive experiments are conducted to evaluate the effectiveness of the proposed model by using the TREC data collection Reuters Corpus Volume 1. The results show that the proposed model significantly outperforms both state-of-the-art term-based models and pattern-based models.
Yang Gao 0016, Yue Xu 0001, Yuefeng Li 0001
IEEE Trans. Knowl. Data Eng.3
2015 Relevance Feature Discovery for Text Mining
abstract
It is a big challenge to guarantee the quality of discovered relevance features in text documents for describing user preferences because of large scale terms and data patterns. Most existing popular text mining and classification methods have adopted term-based approaches. However, they have all suffered from the problems of polysemy and synonymy. Over the years, there has been often held the hypothesis that pattern-based methods should perform better than term-based ones in describing user preferences; yet, how to effectively use large scale patterns remains a hard problem in text mining. To make a breakthrough in this challenging issue, this paper presents an innovative model for relevance feature discovery. It discovers both positive and negative patterns in text documents as higher level features and deploys them over low-level features (terms). It also classifies terms into categories and updates term weights based on their specificity and their distributions in patterns. Substantial experiments using this model on RCV1, TREC topics and Reuters-21578 show that the proposed model significantly outperforms both the state-of-the-art term-based methods and the pattern based methods.
Yuefeng Li 0001, Abdulmohsen Algarni, Mubarak Albathan, Moch Arif Bijaksana
IEEE Trans. Knowl. Data Eng.1
2014 Centroid Training to achieve effective text classification
abstract
Traditional text classification technology based on machine learning and data mining techniques has made a big progress. However, it is still a big problem on how to draw an exact decision boundary between relevant and irrelevant objects in binary classification due to much uncertainty produced in the process of the traditional algorithms. The proposed model CTTC (Centroid Training for Text Classification) aims to build an uncertainty boundary to absorb as many indeterminate objects as possible so as to elevate the certainty of the relevant and irrelevant groups through the centroid clustering and training process. The clustering starts from the two training subsets labelled as relevant or irrelevant respectively to create two principal centroid vectors by which all the training samples are further separated into three groups: POS, NEG and BND, with all the indeterminate objects absorbed into the uncertain decision boundary BND. Two pairs of centroid vectors are proposed to be trained and optimized through the subsequent iterative multi-learning process, all of which are proposed to collaboratively help predict the polarities of the incoming objects thereafter. For the assessment of the proposed model, F1and Accuracy have been chosen as the key evaluation measures. We stress the F1measure because it can display the overall performance improvement of the final classifier better than Accuracy. A large number of experiments have been completed using the proposed model on the Reuters Corpus Volume 1 (RCV1) which is important standard dataset in the field. The experiment results show that the proposed model has significantly improved the binary text classification performance in both F1and Accuracy compared with three other influential baseline models.
Libiao Zhang, Yuefeng Li 0001, Yue Xu 0001, Dian Tjondronegoro
DSAA2
2014 Product Feature Taxonomy Learning based on User Reviews
abstract
In recent years, the Web 2.0 has provided considerable facilities for people to create, share and exchange information and ideas. Upon this, the user generated content, such as reviews, has exploded. Such data provide a rich source to exploit in order to identify the information associated with specific reviewed items. Opinion mining has been widely used to identify the significant features of items (e.g., cameras) based upon user reviews. Feature extraction is the most critical step to identify useful information from texts. Most existing approaches only find individual features about a product without revealing the structural relationships between the features which usually exist. In this paper, we propose an approach to extract features and feature relationships, represented as a tree structure called feature taxonomy, based on frequent patterns and associations between patterns derived from user reviews. The generated feature taxonomy profiles the product at multiple levels and provides more detailed information about the product. Our experiment results based on some popularly used review datasets show that our proposed approach is able to capture the product features and relations effectively.
Nan Tian, Yue Xu 0001, Yuefeng Li 0001, Ahmad Abdel-Hafez, Audun Jøsang
WEBIST (2)3
2014 Topical Pattern Based Document Modelling and Relevance Ranking
Yang Gao 0016, Yue Xu 0001, Yuefeng Li 0001
WISE (1)3
2014 A Review Selection Method Using Product Feature Taxonomy
Nan Tian, Yue Xu 0001, Yuefeng Li 0001
WISE (1)3
2014 Extracting news blog hot topics based on the W2T Methodology
Erzhong Zhou, Ning Zhong 0001, Yuefeng Li 0001
World Wide Web3
2013 Scoring-Thresholding Pattern Based Text Classifier
Moch Arif Bijaksana, Yuefeng Li 0001, Abdulmohsen Algarni
ACIIDS (1)2
2013 Mining Specific Features for Acquiring User Information Needs
Abdulmohsen Algarni, Yuefeng Li 0001
PAKDD (1)2
2013 A Two-Stage Approach for Generating Topic Models
Yang Gao 0016, Yue Xu 0001, Yuefeng Li 0001
PAKDD (2)3
2012 Personalization in tag ontology learning for recommendation making
abstract
Due to the explosive growth of the Web, the domain of Web personalization has gained great momentum both in the research and commercial areas. One of the most popular web personalization systems is recommender systems. In recommender systems choosing user information that can be used to profile users is very crucial for user profiling. In Web 2.0, one facility that can help users organize Web resources of their interest is user tagging systems. Exploring user tagging behavior provides a promising way for understanding users' information needs since tags are given directly by users. However, free and relatively uncontrolled vocabulary makes the user self-defined tags lack of standardization and semantic ambiguity. Also, the relationships among tags need to be explored since there are rich relationships among tags which could provide valuable information for us to better understand users. In this paper, we propose a novel approach for learning tag ontology based on the widely used lexical database WordNet for capturing the semantics and the structural relationships of tags. We present personalization strategies to disambiguate the semantics of tags by combining the opinion of WordNet lexicographers and users' tagging behavior together. To personalize further, clustering of users is performed to generate a more accurate ontology for a particular group of users. In order to evaluate the usefulness of the tag ontology, we use the tag ontology in a pilot tag recommendation experiment for improving the recommendation performance by exploiting the semantic information in the tag ontology. The initial result shows that the personalized information has improved the accuracy of the tag recommendation.
Endang Djuana, Yue Xu 0001, Yuefeng Li 0001, Clive Cox
iiWAS3
2012 Unsupervised Multi-label Text Classification Using a World Knowledge Ontology
Xiaohui Tao 0001, Yuefeng Li 0001, Raymond Y. K. Lau, Hua Wang 0002
PAKDD (1)2
2012 Using Patterns Co-occurrence Matrix for Cleaning Closed Sequential Patterns for Text Mining
abstract
With the overwhelming increase in the amount of texts on the web, it is almost impossible for people to keep abreast of up-to-date information. Text mining is a process by which interesting information is derived from text through the discovery of patterns and trends. Text mining algorithms are used to guarantee the quality of extracted knowledge. However, the extracted patterns using text or data mining algorithms or methods leads to noisy patterns and inconsistency. Thus, different challenges arise, such as the question of how to understand these patterns, whether the model that has been used is suitable, and if all the patterns that have been extracted are relevant. Furthermore, the research raises the question of how to give a correct weight to the extracted knowledge. To address these issues, this paper presents a text post-processing method, which uses a pattern co-occurrence matrix to find the relation between extracted patterns in order to reduce noisy patterns. The main objective of this paper is not only reducing the number of closed sequential patterns, but also improving the performance of pattern mining as well. The experimental results on Reuters Corpus Volume 1 data collection and TREC filtering topics show that the proposed method is promising.
Mubarak Albathan, Yuefeng Li 0001, Abdulmohsen Algarni
Web Intelligence2
2012 Semantic Labelling for Document Feature Patterns Using Ontological Subjects
abstract
Finding and labelling semantic features patterns of documents in a large, spatial corpus is a challenging problem. Text documents have characteristics that make semantic labelling difficult, the rapidly increasing volume of online documents makes a bottleneck in finding meaningful textual patterns. Aiming to deal with these issues, we propose an unsupervised documnent labelling approach based on semantic content and feature patterns. A world ontology with extensive topic coverage is exploited to supply controlled, structured subjects for labelling. An algorithm is also introduced to reduce dimensionality based on the study of ontological structure. The proposed approach was promisingly evaluated by compared with typical machine learning methods including SVMs, Rocchio, and kNN.
Xiaohui Tao 0001, Yuefeng Li 0001
Web Intelligence2
2012 A two-stage decision model for information filtering
Yuefeng Li 0001, Xujuan Zhou, Peter Bruza, Yue Xu 0001, Raymond Y. K. Lau
Decis. Support Syst.1
2012 Effective Pattern Discovery for Text Mining
abstract
Many data mining techniques have been proposed for mining useful patterns in text documents. However, how to effectively use and update discovered patterns is still an open research issue, especially in the domain of text mining. Since most existing text mining methods adopted term-based approaches, they all suffer from the problems of polysemy and synonymy. Over the years, people have often held the hypothesis that pattern (or phrase)-based approaches should perform better than the term-based ones, but many experiments do not support this hypothesis. This paper presents an innovative and effective pattern discovery technique which includes the processes of pattern deploying and pattern evolving, to improve the effectiveness of using and updating discovered patterns for finding relevant and interesting information. Substantial experiments on RCV1 data collection and TREC topics demonstrate that the proposed solution achieves encouraging performance.
Ning Zhong 0001, Yuefeng Li 0001, Sheng-Tang Wu
IEEE Trans. Knowl. Data Eng.2
2012 Text mining in negative relevance feedback
abstract
It is a big challenge to clearly identify the boundary between positive and negative streams. Several attempts have used negative feedback to solve this challenge; however, there are two issues for using negative relevance feedback to improve the eff
Abdulmohsen Algarni, Yuefeng Li 0001, Sheng-Tang Wu, Yue Xu 0001
Web Intell. Agent Syst.2
2012 Personalized recommender systems integrating tags and item taxonomy
abstract
Tags in Web 2.0 are becoming another important information source to profile users' interests and preferences to make personalized recommendations. To solve the problem of low information sharing caused by the free-style vocabulary of tags and the lo
Huizhi Liang 0001, Yue Xu 0001, Yuefeng Li 0001
Web Intell. Agent Syst.3
2011 Effective Hybrid Recommendation Combining Users-Searches Correlations Using Tensors
Rakesh Rawat, Richi Nayak, Yuefeng Li 0001
APWeb3
2011 Aggregate Distance Based Clustering Using Fibonacci Series-FIBCLUS
Rakesh Rawat, Richi Nayak, Yuefeng Li 0001, Slah Alsaleh
APWeb3
2011 XML Documents Clustering Using a Tensor Space Model
Sangeetha Kutty, Richi Nayak, Yuefeng Li 0001
PAKDD (1)3
2011 Pattern Mining for a Two-Stage Information Filtering System
Xujuan Zhou, Yuefeng Li 0001, Peter Bruza, Yue Xu 0001, Raymond Y. K. Lau
PAKDD (1)2
2011 Reliable representations for association rules
Yue Xu 0001, Yuefeng Li 0001, Gavin Shaw
Data Knowl. Eng.2
2011 A pattern mining approach for information filtering systems
Yuefeng Li 0001, Abdulmohsen Algarni, Yue Xu 0001
Inf. Retr.1
2011 A Personalized Ontology Model for Web Information Gathering
abstract
As a model for knowledge description and formalization, ontologies are widely used to represent user profiles in personalized web information gathering. However, when representing user profiles, many models have utilized only knowledge from either a global knowledge base or a user local information. In this paper, a personalized ontology model is proposed for knowledge representation and reasoning over user profiles. This model learns ontological user profiles from both a world knowledge base and user local instance repositories. The ontology model is evaluated by comparing it against benchmark models in web information gathering. The results show that this ontology model is successful.
Xiaohui Tao 0001, Yuefeng Li 0001, Ning Zhong 0001
IEEE Trans. Knowl. Data Eng.2
2010 Selected new training documents to update user profile
abstract
Relevance Feedback (RF) has been proven very effective for improving retrieval accuracy. Adaptive information filtering (AIF) technology has benefited from the improvements achieved in all the tasks involved over the last decades. A difficult problem in AIF has been how to update the system with new feedback efficiently and effectively. In current feedback methods, the updating processes focus on updating system parameters. In this paper, we developed a new approach, the Adaptive Relevance Features Discovery (ARFD). It automatically updates the system's knowledge based on a sliding window over positive and negative feedback to solve a nonmonotonic problem efficiently. Some of the new training documents will be selected using the knowledge that the system currently obtained. Then, specific features will be extracted from selected training documents. Different methods have been used to merge and revise the weights of features in a vector space. The new model is designed for Relevance Features Discovery (RFD), a pattern mining based approach, which uses negative relevance feedback to improve the quality of extracted features from positive feedback. Learning algorithms are also proposed to implement this approach on Reuters Corpus Volume 1 and TREC topics. Experiments show that the proposed approach can work efficiently and achieves the encouragement performance.
Abdulmohsen Algarni, Yuefeng Li 0001, Yue Xu 0001
CIKM2
2010 Personalized recommender system based on item taxonomy and folksonomy
abstract
Item folksonomy or tag information is popularly available on the web now. However, since tags are arbitrary words given by users, they contain a lot of noise such as tag synonyms, semantic ambiguities and personal tags. Such noise brings difficulties to improve the accuracy of item recommendations. In this paper, we propose to combine item taxonomy and folksonomy to reduce the noise of tags and make personalized item recommendations. The experiments conducted on the dataset collected from Amazon.com demonstrated the effectiveness of the proposed approaches. The results suggested that the recommendation accuracy can be further improved if we consider the viewpoints and the vocabularies of both experts and users.
Huizhi Liang 0001, Yue Xu 0001, Yuefeng Li 0001, Richi Nayak
CIKM3
2010 Rough sets based reasoning and pattern mining for a two-stage information filtering system
abstract
This paper presents a novel two-stage information filtering model which combines the merits of term-based and pattern- based approaches to effectively filter sheer volume of infor- mation. In particular, the first filtering stage is supported by a novel rough analysis model which efficiently removes a large number of irrelevant documents, thereby addressing the overload problem. The second filtering stage is empow- ered by a semantically rich pattern taxonomy mining model which effectively fetches incoming documents according to the specific information needs of a user, thereby addressing the mismatch problem. The experiments have been conducted to compare the proposed two-stage filtering (T-SM) model with other possible "term-based + pattern-based" or "term-based + term-based" IF models. The results based on the RCV1 corpus show that the T-SM model significantly outperforms other types of "two-stage" IF models.
Xujuan Zhou, Yuefeng Li 0001, Peter Bruza, Yue Xu 0001, Raymond Y. K. Lau
CIKM2
2010 Mining positive and negative patterns for relevance feature discovery
abstract
It is a big challenge to guarantee the quality of discovered relevance features in text documents for describing user preferences because of the large number of terms, patterns, and noise. Most existing popular text mining and classification methods have adopted term-based approaches. However, they have all suffered from the problems of polysemy and synonymy. Over the years, people have often held the hypothesis that pattern-based methods should perform better than term-based ones in describing user preferences, but many experiments do not support this hypothesis. The innovative technique presented in paper makes a breakthrough for this difficulty. This technique discovers both positive and negative patterns in text documents as higher level features in order to accurately weight low-level features (terms) based on their specificity and their distributions in the higher level features. Substantial experiments using this technique on Reuters Corpus Volume 1 and TREC topics show that the proposed approach significantly outperforms both the state-of-the-art term-based methods underpinned by Okapi BM25, Rocchio or Support Vector Machine and pattern based methods on precision, recall and F measures.
Yuefeng Li 0001, Abdulmohsen Algarni, Ning Zhong 0001
KDD1
2010 Ontology-Based Specific and Exhaustive User Profiles for Constraint Information Fusion for Multi-agents
abstract
Intelligent agents are an advanced technology utilized in Web Intelligence. When searching information from a distributed Web environment, information is retrieved by multi-agents on the client site and fused on the broker site. The current information fusion techniques rely on cooperation of agents to provide statistics. Such techniques are computationally expensive and unrealistic in the real world. In this paper, we introduce a model that uses a world ontology constructed from the Dewey Decimal Classification to acquire user profiles. By search using specific and exhaustive user profiles, information fusion techniques no longer rely on the statistics provided by agents. The model has been successfully evaluated using the large INEX data set simulating the distributed Web environment.
Xiaohui Tao 0001, Yuefeng Li 0001, Raymond Y. K. Lau, Shlomo Geva
Web Intelligence2
2010 A knowledge-based model using ontologies for personalized web information gathering
abstract
Nowadays, how to gather useful and meaningful information from the Web has become challenging to all users because of the explosion in the amount of Web information. However, the mainstream of Web information gathering techniques has many drawbacks,
Xiaohui Tao 0001, Yuefeng Li 0001, Ning Zhong 0001
Web Intell. Agent Syst.2
2009 An effective model of using negative relevance feedback for information filtering
abstract
Over the years, people have often held the hypothesis that negative feedback should be very useful for largely improving the performance of information filtering systems; however, we have not obtained very effective models to support this hypothesis. This paper, proposes an effective model that use negative relevance feedback based on a pattern mining approach to improve extracted features. This study focuses on two main issues of using negative relevance feedback: the selection of constructive negative examples to reduce the space of negative examples; and the revision of existing features based on the selected negative examples. The former selects some offender documents, where offender documents are negative documents that are most likely to be classified in the positive group. The later groups the extracted features into three groups: the positive specific category, general category and negative specific category to easily update the weight. An iterative algorithm is also proposed to implement this approach on RCV1 data collections, and substantial experiments show that the proposed approach achieves encouraging performance.
Abdulmohsen Algarni, Yuefeng Li 0001, Yue Xu 0001, Raymond Y. K. Lau
CIKM2
2009 XCFS: an XML documents clustering approach using both the structure and the content
abstract
This paper introduces a clustering approach, XML Clustering using Frequent Substructures (XCFS) that considers both the structural and the content information of XML documents in clustering. XCFS uses frequent substructures in the form of a novel representation, Closed Frequent Embedded (CFE) subtrees to constrain the content in the clustering process. The empirical analysis ascertains that XCFS can effectively cluster even very large XML datasets and outperforms other existing methods.
Sangeetha Kutty, Richi Nayak, Yuefeng Li 0001
CIKM3
2009 HCX: an efficient hybrid clustering approach for XML documents
abstract
This paper proposes a novel Hybrid Clustering approach for XML documents (HCX) that first determines the structural similarity in the form of frequent subtrees and then uses these frequent subtrees to represent the constrained content of the XML documents in order to determine the content similarity. The empirical analysis reveals that the proposed method is scalable and accurate.
Sangeetha Kutty, Richi Nayak, Yuefeng Li 0001
ACM Symposium on Document Engineering3
2009 Concept-Based, Personalized Web Information Gathering: A Survey
Xiaohui Tao 0001, Yuefeng Li 0001
KSEM2
2009 Mining Negative Relevance Feedback for Information Filtering
abstract
It is a big challenge to clearly identify the boundary between positive and negative streams. Several attempts have used negative feedback to solve this challenge; however, there are two issues for using negative relevance feedback to improve the effectiveness of information filtering. The first one is how to select constructive negative samples in order to reduce the space of negative documents. The second issue is how to decide noisy extracted features that should be updated based on the selected negative samples. This paper proposes a pattern mining based approach to select some offenders from the negative documents, where an offender can be used to reduce the side effects of noisy features. It also classifies extracted features (i.e., terms) into three categories: positive specific terms, general terms, and negative specific terms. In this way, multiple revising strategies can be used to update extracted features. An iterative learning algorithm is also proposed to implement this approach on RCV1, and substantial experiments show that the proposed approach achieves encouraging performance.
Yuefeng Li 0001, Abdulmohsen Algarni, Sheng-Tang Wu, Yue Xue
Web Intelligence1
2009 Personalized Recommender Systems Integrating Social Tags and Item Taxonomy
abstract
The social tags in web 2.0 are becoming another important information source to profile users' interests and preferences to make personalized recommendations. To solve the problem of low information sharing caused by the free-style vocabulary of tags and the long tails of the distribution of tags and items, this paper proposes an approach to integrate the social tags given by users and the item taxonomy with standard vocabulary and hierarchical structure provided by experts to make personalized recommendations. The experimental results show that the proposed approach can effectively improve the information sharing and recommendation accuracy.
Huizhi Liang 0001, Yue Xu 0001, Yuefeng Li 0001, Richi Nayak, Li-Tung Weng
Web Intelligence3
2009 Toward a Fuzzy Domain Ontology Extraction Method for Adaptive e-Learning
abstract
With the widespread applications of electronic learning (e-Learning) technologies to education at all levels, increasing number of online educational resources and messages are generated from the corresponding e-Learning environments. Nevertheless, it is quite difficult, if not totally impossible, for instructors to read through and analyze the online messages to predict the progress of their students on the fly. The main contribution of this paper is the illustration of a novel concept map generation mechanism which is underpinned by a fuzzy domain ontology extraction algorithm. The proposed mechanism can automatically construct concept maps based on the messages posted to online discussion forums. By browsing the concept maps, instructors can quickly identify the progress of their students and adjust the pedagogical sequence on the fly. Our initial experimental results reveal that the accuracy and the quality of the automatically generated concept maps are promising. Our research work opens the door to the development and application of intelligent software tools to enhance e-Learning.
Raymond Y. K. Lau, Dawei Song 0001, Yuefeng Li 0001, Chun-Ho Cheung, Jin-Xing Hao
IEEE Trans. Knowl. Data Eng.3
2008 Effective pattern taxonomy mining in text documents
abstract
Many data mining techniques have been proposed for mining useful patterns in databases. However, how to effectively utilize discovered patterns is still an open research issue, especially in the domain of text mining. Most existing methods adopt term-based approaches. However, they all suffer from the problems of polysemy and synonymy. This paper presents an innovative technique, pattern taxonomy mining, to improve the effectiveness of using discovered patterns for finding useful information. Substantial experiments on RCV1 demonstrate that the proposed solution achieves encouraging performance.
Yuefeng Li 0001, Sheng-Tang Wu, Xiaohui Tao 0001
CIKM1
2008 A two-stage text mining model for information filtering
abstract
Mismatch and overload are the two fundamental issues regarding the effectiveness of information filtering. Both term-based and pattern (phrase) based approaches have been employed to address these issues. However, they all suffer from some limitations with regard to effectiveness. This paper proposes a novel solution that includes two stages: an initial topic filtering stage followed by a stage involving pattern taxonomy mining. The objective of the first stage is to address mismatch by quickly filtering out probable irrelevant documents. The threshold used in the first stage is motivated theoretically. The objective of the second stage is to address overload by apply pattern mining techniques to rationalize the data relevance of the reduced document set after the first stage. Substantial experiments on RCV1 show that the proposed solution achieves encouraging performance.
Yuefeng Li 0001, Xujuan Zhou, Peter Bruza, Yue Xu 0001, Raymond Y. K. Lau
CIKM1
2008 A User Driven Data Mining Process Model and Learning System
Esther Ge, Richi Nayak, Yue Xu 0001, Yuefeng Li 0001
DASFAA4
2008 Exploiting Item Taxonomy for Solving Cold-Start Problem in Recommendation Making
abstract
Recommender systems' performance can be easily affected when there are no sufficient item preferences data provided by previous users, and it is commonly referred to as cold-start problem. This paper suggests another information source, item taxonomies, in addition to item preference data for assisting recommendation making. Item taxonomy information has been popularly applied in diverse ecommerce domains for product or content classification, and therefore can be easily obtained and adapted by recommender systems. In this paper, we investigate the implicit relations between users' item preferences and taxonomic preferences, suggest and verify that users who share similar item preferences may also share similar taxonomic preferences. Under this assumption, a novel recommendation technique is proposed that combines the users' item preferences and the additional taxonomic preferences together to make better quality recommendations as well as alleviate the cold-start problem. Empirical evaluations to this approach are conducted and the results show that the proposed technique outperforms other existing techniques in both recommendation quality and computation efficiency.
Li-Tung Weng, Yue Xu 0001, Yuefeng Li 0001, Richi Nayak
ICTAI (2)3
2008 Concise representations for approximate association rules
abstract
The quality of association rule mining has drawn more and more attention recently. One problem with the quality of the discovered association rules is the huge size of the extracted rule set. Often for a dataset, a huge number of rules can be extracted, but many of them can be redundant to other rules and thus useless in practice. Mining non-redundant rules is a promising approach to solve this problem. In this paper, we firstly propose a definition for redundancy; then we propose a concise representation called reliable basis for representing non-redundant association rules for both exact rules and approximate rules. We prove that the redundancy elimination based on the reliable basis does not reduce the belief to the extracted rules. We also prove that all association rules can be deduced from the reliable basis. Therefore the reliable basis is a lossless representation of association rules. Experimental results show that the reliable basis significantly reduces the number of extracted rules.
Yue Xu 0001, Yuefeng Li 0001, Gavin Shaw
SMC2
2008 An Ontology-Based Framework for Knowledge Retrieval
abstract
Retrieving accurate information from the Web is a great challenge to users. The existing information retrieval systems are mostly term-based and thus need to be enhanced toward knowledge-based. User information needs need to be better captured in order to deliver personalized search results. In this paper, an ontology-based framework is proposed for capturing user information needs using a world knowledge base and the user's local instance repository. The framework aims to discover a user's background knowledge for knowledge retrieval. The evaluation result is encouraging, in which the proposed model achieved the same performance as a manual user model.
Xiaohui Tao 0001, Yuefeng Li 0001, Ning Zhong 0001, Richi Nayak
Web Intelligence2
2008 Knowledge discovery for adaptive negotiation agents in e-marketplaces
Raymond Y. K. Lau, Yuefeng Li 0001, Dawei Song 0001, Ron Chi-Wai Kwok
Decis. Support Syst.2
2007 Generating concise association rules
abstract
Association rule mining has made many achievements in the area of knowledge discovery. However, the quality of the extracted association rules is a big concern. One problem with the quality of the extracted association rules is the huge size of the extracted rule set. As a matter of fact, very often tens of thousands of association rules are extracted among which many are redundant thus useless. Mining non-redundant rules is a promising approach to solve this problem. The Min-max exact basis proposed by Pasquier et al [Pasquier05] has showed exciting results by generating only non-redundant rules. In this paper, we first propose a relaxing definition for redundancy under which the Min-max exact basis still contains redundant rules; then we propose a condensed representation called Reliable exact basis for exact association rules. The rules in the Reliable exact basis are not only non-redundant but also more succinct than the rules in Min-max exact basis. We prove that the redundancy eliminated by the Reliable exact basis does not reduce the belief to the Reliable exact basis. The size of the Reliable exact basis is much smaller than that of the Min-max exact basis. Moreover, we prove that all exact association rules can be deduced from the Reliable exact basis. Therefore the Reliable exact basis is a lossless representation of exact association rules. Experimental results show that the Reliable exact basis significantly reduces the number of non-redundant rules.
Yue Xu 0001, Yuefeng Li 0001
CIKM2
2007 Granule Based Intertransaction Association Rule Mining
abstract
Intertransaction association rule mining is used to discover patterns between different transactions. It breaks the scope of association rule mining on the same transaction. Currently the FITI algorithm is the state of the art in intertransaction association rule mining. However, the FTTI introduces many unneeded combinations of items because the set of extended items is much larger than the set of items. Thus, we propose an alternative approach of granule based intertransaction association rule mining, where a granule is a group of transactions that meet a certain constraint. The experimental results show that this approach is promising in real-world industry.
Wanzhong Yang, Yuefeng Li 0001, Yue Xu 0001
ICTAI (1)2
2007 Ontology Mining for Semantic Interpretation of Information Needs
Xiaohui Tao 0001, Yuefeng Li 0001, Richi Nayak
KSEM2
2007 Mining Fuzzy Domain Ontology from Textual Databases
abstract
Ontology plays an essential role in the formalization of common information (e.g., products, services, relationships of businesses) for effective human-computer interactions. However, engineering of these ontologies turns out to be very labor intensive and time consuming. Although some text mining methods have been proposed for automatic or semi-automatic discovery of crisp ontologies, the robustness, accuracy, and computational efficiency of these methods need to be improved to support large scale ontology construction for real-world applications. This paper illustrates a novel fuzzy domain ontology mining algorithm for supporting real-world ontology engineering. In particular, contextual information of the knowledge sources is exploited for the extraction of high quality domain ontologies and the uncertainty embedded in the knowledge sources is modeled based on the notion of fuzzy sets. Empirical studies have confirmed that the proposed method can discover high quality fuzzy domain ontology which leads to significant improvement in information retrieval performance.
Raymond Y. K. Lau, Yuefeng Li 0001, Yue Xu 0001
Web Intelligence2
2007 Ontology Mining for PersonalizedWeb Information Gathering
abstract
It is well accepted that ontology is useful for personalized Web information gathering. However, it is challenging to use semantic relations of "kind-of", "part-of", and "related-to" and synthesize commonsense and expert knowledge in a single computational model. In this paper, a personalized ontology model is proposed attempting to answer this challenge. A two-dimensional (Exhaustivity and Specificity) method is also presented to quantitatively analyze these semantic relations in a single framework. The proposals are successfully evaluated by applying the model to a Web information gathering system. The model is a significant contribution to personalized ontology engineering and concept-based Web information gathering in Web Intelligence.
Xiaohui Tao 0001, Yuefeng Li 0001, Ning Zhong 0001, Richi Nayak
Web Intelligence2
2007 Using Information Filtering in Web Data Mining Process
abstract
The amount of Web information is growing rapidly, improving the efficiency and accuracy of Web information retrieval is uphill battle. There are two fundamental issues regarding the effectiveness of Web information gathering: information mismatch and overload. To tackle these difficult issues, an integrated information filtering and sophisticated data processing model has been presented in this paper. In the first phase of the proposed scheme, an information filter that based on user search intents was incorporated in Web search process to quickly filter out irrelevant data. In the second data processing phase, a pattern taxonomy model (PTM) was carried out using the reduced data. PTM rationalizes the data relevance by applying data mining techniques that involves more rigorous computations. Several experiments have been conducted and the results show that more effective and efficient access Web information has been achieved using the new scheme.
Xujuan Zhou, Yuefeng Li 0001, Peter Bruza, Sheng-Tang Wu, Yue Xu 0001, Raymond Y. K. Lau
Web Intelligence2
2007 Sequential Pattern Mining and Nonmonotonic Reasoning for Intelligent Information Agents
abstract
With the explosive growth of information available on the Internet, more effective data mining and data reasoning mechanism is required to process the sheer volume of information. Belief revision logic offers the expressive power to represent information retrieval contexts, and it also provides a sound inference mechanism to model the nonmonotonicity arising in changing retrieval contexts. Contextual knowledge for information retrieval can be extracted via efficient sequential pattern mining. We present a pattern taxonomy extraction model which efficiently performs the task of discovering descriptive frequent sequential patterns by pruning the noisy associations. This paper illustrates a novel approach of integrating the sequential data mining method into the belief revision based adaptive information agents to improve the agents' learning autonomy and prediction power. Initial experiments show that our belief revision logic and sequential pattern mining based intelligent information agents outperform the vector space model based information agents. Our work opens the door to the development of next generation of intelligent information agents to alleviate the information overload problem.
Raymond Y. K. Lau, Yuefeng Li 0001, Sheng-Tang Wu, Xujuan Zhou
Int. J. Pattern Recognit. Artif. Intell.2
2007 Editorial
Yuefeng Li 0001, Ning Zhong 0001
Int. J. Pattern Recognit. Artif. Intell.1
2007 Mining Non-Redundant Association Rules Based on Concise Bases
abstract
Association rule mining has many achievements in the area of knowledge discovery. However, the quality of the extracted association rules has not drawn adequate attention from researchers in data mining community. One big concern with the quality of association rule mining is the size of the extracted rule set. As a matter of fact, very often tens of thousands of association rules are extracted among which many are redundant, thus useless. In this paper, we first analyze the redundancy problem in association rules and then propose a reliable exact association rule basis from which more concise nonredundant rules can be extracted. We prove that the redundancy eliminated using the proposed reliable association rule basis does not reduce the belief to the extracted rules. Moreover, this paper proposes a level wise approach for efficiently extracting closed itemsets and minimal generators — a key issue in closure based association rule mining.
Yue Xu 0001, Yuefeng Li 0001
Int. J. Pattern Recognit. Artif. Intell.2
2007 Mining world knowledge for analysis of search engine content
John D. King, Yuefeng Li 0001, Xiaohui Tao 0001, Richi Nayak
Web Intell. Agent Syst.2
2006 Multi-Tier Granule Mining for Representations of Multidimensional Association Rules
abstract
It is a big challenge to promise the quality of multidimensional association mining. The essential issue is how to represent meaningful multidimensional association rules efficiently. Currently we have not found satisfactory approaches for solving this challenge because of the complicated correlation between attributes. Multi-tier granule mining is an initiative for solving this challenging issue. It divides attributes into some tiers and then compresses the large multidimensional database into granules at each tier. It also builds association mappings to illustrate the correlation between tiers. In this way, the meaningful association rules can be justified according to these association mappings.
Yuefeng Li 0001, Wanzhong Yang, Yue Xu 0001
ICDM1
2006 Deploying Approaches for Pattern Refinement in Text Mining
abstract
Text mining is the technique that helps users find useful information from a large amount of digital text documents on the Web or databases. Instead of the keyword-based approach which is typically used in this field, the pattern-based model containing frequent sequential patterns is employed to perform the same concept of tasks. However, how to effectively use these discovered patterns is still a big challenge. In this study, we propose two approaches based on the use of pattern deploying strategies. The performance of the pattern deploying algorithms for text mining is investigated on the Reuters dataset RCVI and the results show that the effectiveness is improved by using our proposed pattern refinement approaches.
Sheng-Tang Wu, Yuefeng Li 0001, Yue Xu 0001
ICDM2
2006 Rough Association Rule Mining in Text Documents for Acquiring Web User Information Needs
abstract
It is a big challenge to apply data mining techniques for effective Web information gathering because of duplications and ambiguities of data values (e.g., terms). To provide an effective solution to this challenge, this paper first explains the relationship between association rules and rough set based decision rules. It proves that a decision pattern is a kind of closed pattern. It also presents a novel concept of rough association rules in order to improve the effectiveness of association rule mining. The premise of a rough association rule consists of a set of terms and a frequency distribution of terms. The distinct advantage of rough association rules is that they contain more specific information than normal association rules. It is also feasible to update rough association rules dynamically to produce effective results
Yuefeng Li 0001, Ning Zhong 0001
Web Intelligence1
2006 Automatically Acquiring Training Sets for Web Information Gathering
abstract
The traditional techniques rely on human effort to acquire training sets, which is expensive and inefficient. In this paper we present an alternative method to automatically acquire training sets without heavy investment of user efforts. The proposed method tends to fill a gap for effectiveness of using Web data in Web mining, and contributes to Web information gathering. The evaluation shows that the method is adequate to yield an promising achievement.
Xiaohui Tao 0001, Yuefeng Li 0001, Ning Zhong 0001, Richi Nayak
Web Intelligence2
2006 Distributed Recommender Profiling and Selection with Gittins Indices
abstract
Most existing recommender systems nowadays operate in a single organizational base, and very often they do not have sufficient resources to be used in order to generate quality recommendations. Therefore, it would be beneficial if recommender systems of different organizations can cooperate together to share their resources and recommendations. In this paper, we present a distributed recommender system model that consists of multiple recommender systems from different organizations. With the hope to provide better recommendation service to users, the recommender systems can improve their performances by sharing their recommendations cooperatively. A recommender selection technique based on the Gittins indices is presented in this paper, and it makes selections based on the stability, average performance and selection frequency of the recommenders
Li-Tung Weng, Yue Xu 0001, Yuefeng Li 0001, Richi Nayak
Web Intelligence3
2006 Utilizing Search Intent in Topic Ontology-Based User Profile for Web Mining
abstract
It is well known that taking the Web user profiles into account can enhance the effectiveness of Web mining systems. However, due to the dynamic and complex nature of Web users, automatically acquiring worthwhile user profiles was found to be very challenging. Ontology-based user profile can possess more accurate user information. This research emphasizes on acquiring search intentions information. This paper presents a new approach of developing user profile for Web searching. The model considers the user's search intentions by the process of PTM (Pattern-Taxonomy Model). Initial experiments show that the user profile based on search intention is more useful than the generic PTM user profile. Developing user profile that contains user search intentions is essential for effective Web search and retrieval.
Xujuan Zhou, Sheng-Tang Wu, Yuefeng Li 0001, Yue Xu 0001, Raymond Y. K. Lau, Peter Bruza
Web Intelligence3
2006 Mining Ontology for Automatically Acquiring Web User Information Needs
Yuefeng Li 0001, Ning Zhong 0001
IEEE Trans. Knowl. Data Eng.1
2005 An Effective Deploying Algorithm for Using Pattern-Taxonomy
Sheng-Tang Wu, Yuefeng Li 0001, Yue Xu 0001
iiWAS2
2005 Information Fusion With Subject-Based Information Gathering Method for Intelligent Multi-Agent Models
Xiaohui Tao 0001, John D. King, Yuefeng Li 0001
iiWAS3
2005 Mining Interesting Topics for Web Information Gathering and Web Personalization
abstract
The quality of discovery patterns is crucial for building satisfactory systems of Web text mining. It is no doubt that we can find numerous frequent patterns from Web documents. However, there are many meaningless frequent patterns. This paper presents a novel method to improve the quality of discovered patterns. It generalizes discovered patterns into interesting topics in order to acquire the necessary useful information. The experimental results also verify the proposed method is promising.
Yuefeng Li 0001, Ben Murphy, Ning Zhong 0001
Web Intelligence1
2004 Modelling Traditional Chinese Paintings for Content-Based Image Classification and Retrieval
abstract
Content-based image retrieval (CBIR) has been investigated extensively in the past decade in order to classify and search images according to similarities derived from automatically extracted visual features, such as colours, textures and object shapes. It has now been realised that two fundamental problems in CBIR, namely, feature extraction and similarity measure, are likely to be domain specific. In this paper, we present some early results of applying CBIR to traditional Chinese paintings. Our research is motivated by three main goals: (1) to develop tools for art historians to study evolution and cross-influences of oriental paintings by automatically identifying visual artistic clues from digitised paintings; (2) to verify and further advance the existing CBIR techniques by limiting the images studied to a specific and simpler domain of traditional Chinese paintings; and (3) to verify and further advance the problem of high-dimensional data clustering (especially in relation to the "dimensional curse" problem). We present a framework for modelling traditional Chinese paintings, and examine various existing CBIR proposals and algorithms for their suitability for traditional Chinese paintings. In this paper, we also present a research agenda to study the problems of Chinese paintings classification and retrieval based on the framework.
Danqing Zhang, Binh Pham 0001, Yuefeng Li 0001
MMM3
2004 Capturing Evolving Patterns for Ontology-based Web Mining
abstract
An ontology-based Web mining model tends to extract an ontology from user feedback and use it to search the right data from the Web to answer what users want. It is indubitable that we can obtain numerous discovered patterns using a Web mining model. However, some discovered patterns might include uncertainties when we extract them. Also user profiles are changeable. Therefore, the difficult issue is how to use and maintain the discovered patterns. This paper presents a theoretical framework for this issue, which consists of automatic ontology extraction, reasoning on the ontology and capturing evolving patterns. The experimental results show that all objectives we expect for the theoretical framework are achievable.
Yuefeng Li 0001, Ning Zhong 0001
Web Intelligence1
2004 Automatic Pattern-Taxonomy Extraction for Web Mining
abstract
In this paper, we propose a model for discovering frequent sequential patterns, phrases, which can be used as profile descriptors of documents. It is indubitable that we can obtain numerous phrases using data mining algorithms. However, it is difficult to use these phrases effectively for answering what users want. Therefore, we present a pattern taxonomy extraction model which performs the task of extracting descriptive frequent sequential patterns by pruning the meaningless ones. The model then is extended and tested by applying it to the information filtering system. The results of the experiment show that pattern-based methods outperform the keyword-based methods. The results also indicate that removal of meaningless patterns not only reduces the cost of computation but also improves the effectiveness of the system.
Sheng-Tang Wu, Yuefeng Li 0001, Yue Xu 0001, Binh Pham 0001, Yi-Ping Phoebe Chen
Web Intelligence2
2004 Web mining model and its applications for information gathering
Yuefeng Li 0001, Ning Zhong 0001
Knowl. Based Syst.1
2003 Interpretations of Association Rules by Granular Computing
abstract
We present interpretations for association rules. We first introduce Pawlak's method, and the corresponding algorithm of finding decision rules (a kind of association rules). We then use extended random sets to present a new algorithm of finding interesting rules. We prove that the new algorithm is faster than Pawlak's algorithm. The extended random sets are easily to include more than one criterion for determining interesting rules. We also provide two measures for dealing with uncertainties in association rules.
Yuefeng Li 0001, Ning Zhong 0001
ICDM1
2003 Ontology-Based Web Mining Model: Representations of User Profiles
abstract
Web mining is used to search the right information from the Web to meet user information needs. Acquiring correct user profiles is difficult, since users may be unsure of their interests and may not wish to invest a great deal of effort in creating such a profile. Our aim is to present a foundation for representations of user profiles on ontology for designing efficient Web mining models. We assume the user concept can be constructed from some primary ones; hence, we use "part-of" relation to describe the relationships between classes. We also present set-valued relevance functions on such ontology to unravel the relationships between facts and the existing classes. A numerical interpretation is also presented for the set-valued relation functions.
Yuefeng Li 0001, Ning Zhong 0001
Web Intelligence1