EDBT 2026 Demo / reviewers in the wild / expert
Yue Xu 0001
dblp:29/3925-1
· DBLP profile ↗
95ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0002-1137-0272ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 59 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 52 · 5 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 3 since 2021Security and privacy · 6 · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-authorTheory of computation · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-agent reinforcement curriculum learning for real unmanned ground vehiclesabstractThis paper investigates the use of deep reinforcement learning (DRL) for the control of mobile robot teams within the context of navigation and task-based collaborative scenarios. We apply a DRL policy with a tailored neural network architecture as a solution to control, path planning, and higher-level guidance tasks. Our network architecture was trained using a unique multi-stage curriculum that progresses from single-agent navigation, to multi-agent pathfinding with obstacles, and finally to a complex collaborative firefighting scenario. This structured approach accelerates training convergence by systematically building sophisticated collaborative behaviours upon foundational skills, which enhances training stability and guides the agents towards learning effective and coordinated strategies The policy evaluation was conducted in both simulation and hybrid simulation-physical demonstrations utilising a real unmanned ground vehicle (UGV). The policy presented is capable of achieving multi-agent navigation tasks with a 95.83% accuracy in our testing environments, and has demonstrated emergent multi-agent behaviours. In more complex collaborative firefighting scenarios, the policy also demonstrated superior performance than baselines in reaching goals, e.g., navigating and extinguishing two fires with a 99.67% success rate, suggesting its strong potential for real-world deployment. Timothy Mead, Zhe Wang 0001, Ernest Foo, Jin Song Dong 0001, Naipeng Dong, Ryan Kok Leong Ko, Abigail M. Y. Koay, Kien Nguyen Thanh, Yue Xu 0001, Junae Kim, Stephen Bornstein |
Eng. Appl. Artif. Intell. | 9 |
| 2025 | Developing guidelines for functionally-grounded evaluation of explainable artificial intelligence using tabular dataabstractExplainable Artificial Intelligence (XAI) techniques are used to provide transparency to complex, opaque predictive models. However, these techniques are often designed for image and text data, and it is unclear how fit-for-purpose they are when applied to tabular data. As XAI techniques are rarely evaluated in the context of tabular data, the applicability of existing evaluation criteria and methods are also unclear and needs re-examination. For example, some works suggest that evaluation methods may unduly influence the evaluation results when using tabular data. This lack of clarity on evaluation procedures can lead to reduced transparency and ineffective use of XAI techniques in real world settings. In this study, we examine literature on XAI evaluation to derive guidelines on functionally-grounded assessment of local, post hoc XAI techniques. We identify 20 evaluation criteria and associated evaluation methods, and derive guidelines on when and how each criterion should be evaluated. We also identify key research gaps to be addressed by future work. Our study contributes to the body of knowledge on XAI evaluation through in-depth examination of functionally-grounded XAI evaluation protocols, and has laid the groundwork for future research on XAI evaluation. Mythreyi Velmurugan, Chun Ouyang 0001, Yue Xu 0001, Renuka Sindhgatta, Bemali Wickramanayake, Catarina Moreira |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Contextual Transformer-based Node Embedding for Vulnerability Detection using Graph LearningabstractAutomated source code vulnerability detection using code graphs has seen major improvements in recent years, however one critical, but oft-overlooked, element of this problem is producing embeddings for graph nodes. Before graph-based classifiers can be used for vulnerability detection, the nodes in the graph must first be given vector representations. Graphlearning models propagate information from these embeddings through the graph before classification, and so the initial states of these embeddings are vital for all subsequent learning. While a variety of solutions to this problem have been proposed in existing literature, this is typically not the focus of these works. We propose a novel node embedding strategy for graph-based vulnerability discovery, which takes advantage of richly-learned information about the code contained in each node. We also implement and test several existing node embedding strategies, comparing them to each other and our new strategy under a standard graph-learning architecture. We find that our strategy outperforms existing methods by 10.47-50.70%. Joseph Gear, Yue Xu 0001, Ernest Foo, Praveen Gauravaram, Zahra Jadidi, Leonie Ruth Simpson |
TrustCom | 2 |
| 2024 | Improved Packet-Level Synthetic Network Traffic GenerationabstractWhile using generative models to create synthetic network traffic is faster and cheaper than traditional testbeds, synthetic traffic suffers from problems with realism and structural completeness. State of the art traffic generation frameworks usually omit payloads because of the difficulties in representing their high-dimensional data, which makes the synthetic traffic unrealistic and limits its usefulness. This work proposes a two-stage process that takes advantage of the high repetition of some protocols, particularly those used by Industrial Control Systems, to selectively simplify payloads, greatly reducing the number of classes and reducing model loss and consequently the ability of the model to handle sequences of payloads. Model training loss was reduced by 47.796%, and payload class selection was improved up to 69% over state of the art approaches, allowing for more realistic synthetic network traffic with reduced memory and computation overheads. Jacob Soper, Yue Xu 0001, Ernest Foo, Zahra Jadidi, Kien Nguyen Thanh |
TrustCom | 2 |
| 2024 | Learning medical concept representation based on semantic information in medical textural dataabstractElectric Health Records (EHR) have been widely adopted by many hospitals to improve clinical decision making and re-admission prediction. Each patient admission usually contains both multiple medical codes as well as clinical notes. Accurate learning of the representations (also called embeddings) of medical concepts from EHR is a key strategy to improve prediction performance in healthcare. Existing works employ medical ontologies to improve the quality of representations but focus solely on the relationships amongst the medical codes, ignoring textual data such medical code descriptions, clinical notes, and patient demographics. In this paper, we propose a new model called Semantic-based Attention model using Textual data for Medical Concept Embedding (SATexMCE). SATexMCE consists of three parts: medical codes embedding, admission embedding, and prediction model. First, we generate representations of medical codes by using both the textual description of medical codes and also the relationships among medical codes. Then, we generate the admission representation based on the representation of medical codes in the admission, the clinical notes and the demographic information associated with the admission via several attention mechanisms. The admission embeddings are used to construct a Recurrent neural network model which is used to predict patients’ readmission and a disease in the next admission based on patient admission data. Experimental results show that the proposed SATexMCE model improves not only the performance of readmission prediction but also the quality of medical concept representations. Attention mechanism helps us measure the importance of different medical codes and understand the meaning of the words in clinical notes for predictions. Sea Jung Im, Yue Xu 0001, Jason Watson |
Expert Syst. Appl. | 2 |
| 2024 | WSNMF: Weighted Symmetric Nonnegative Matrix Factorization for attributed graph clustering
Kamal Berahmand, Mehrnoush Mohammadi, Razieh Sheikhpour, Yuefeng Li 0001, Yue Xu 0001 |
Neurocomputing | 5 |
| 2024 | A Deep Semi-Supervised Community Detection Based on Point-Wise Mutual InformationabstractNetwork clustering is one of the fundamental unsupervised methods of knowledge discovery. Its goal is to group similar nodes together without supervision or prior knowledge of the nature of the clusters. Among various clustering methods, semi-supervised clustering detection is one of the most promising approaches for community detection because of its ability to employ side information to better understand network topology. However, most of the previous work faces two problems: the use of linear methods to reduce dimensionality and the random selection of side information, and as a result of these two drawbacks, semi-supervised community detection methods are less efficient. To fill these gaps, we developed an end-to-end deep semi-supervisor community detection (DSSC) for complex networks. A new learning objective is designed that uses a semi-autoencoder (SeAE) with a defined pair-wise constraint matrix based on point-wise mutual information (PMI) in the representation layer to accurately learn distinctive features and, in the clustering layer, adds a pair-wise constraint as a term to minimize distance within the cluster while the distance between clusters increases. The results show that our method performs unexpectedly well in comparison to the existing state-of-the-art community detection methods in complex networks. Kamal Berahmand, Yuefeng Li 0001, Yue Xu 0001 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | A Semantics-enhanced Topic Modelling Technique: Semantic-LDAabstractTopic modelling is a beneficial technique used to discover latent topics in text collections. But to correctly understand the text content and generate a meaningful topic list, semantics are important. By ignoring semantics, that is, not attempting to grasp the meaning of the words, most of the existing topic modelling approaches can generate some meaningless topic words. Even existing semantic-based approaches usually interpret the meanings of words without considering the context and related words. In this article, we introduce a semantic-based topic model called semantic-LDA that captures the semantics of words in a text collection using concepts from an external ontology. A new method is introduced to identify and quantify the concept–word relationships based on matching words from the input text collection with concepts from an ontology without using pre-calculated values from the ontology that quantify the relationships between the words and concepts. These pre-calculated values may not reflect the actual relationships between words and concepts for the input collection, because they are derived from datasets used to build the ontology rather than from the input collection itself. Instead, quantifying the relationship based on the word distribution in the input collection is more realistic and beneficial in the semantic capture process. Furthermore, an ambiguity handling mechanism is introduced to interpret the unmatched words, that is, words for which there are no matching concepts in the ontology. Thus, this article makes a significant contribution by introducing a semantic-based topic model that calculates the word–concept relationships directly from the input text collection. The proposed semantic-based topic model and an enhanced version with the disambiguation mechanism were evaluated against a set of state-of-the-art systems, and our approaches outperformed the baseline systems in both topic quality and information filtering evaluations. Dakshi T. K. Geeganage, Yue Xu 0001, Yuefeng Li 0001 |
ACM Trans. Knowl. Discov. Data | 2 |
| 2024 | SDAC-DA: Semi-Supervised Deep Attributed Clustering Using Dual AutoencoderabstractAttributed graph clustering aims to group nodes into disjoint categories using deep learning to represent node embeddings and has shown promising performance across various applications. However, two main challenges hinder further performance improvement. Firstly, reliance on unsupervised methods impedes the learning of low-dimensional, clustering-specific features in the representation layer, thus impacting clustering performance. Secondly, the predominant use of separate approaches leads to suboptimal learned embeddings that are insufficient for subsequent clustering steps. To address these limitations, we propose a novel method called Semi-supervised Deep Attributed Clustering using Dual Autoencoder (SDAC-DA). This approach enables semi-supervised deep end-to-end clustering in attributed networks, promoting high structural cohesiveness and attribute homogeneity. SDAC-DA transforms the attribute network into a dual-view network, applies a semi-supervised autoencoder layering approach to each view, and integrates dimensionality reduction matrices by considering complementary views. The resulting representation layer contains high clustering-friendly embeddings, which are optimized through a unified end-to-end clustering process for effectively identifying clusters. Extensive experiments on both synthetic and real networks demonstrate the superiority of our proposed method over seven state-of-the-art approaches. Kamal Berahmand, Sondos Bahadori, Maryam Nooraei Abadeh, Yuefeng Li 0001, Yue Xu 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | Supervised fusion content-based framework for breakdown detection in task-oriented conversational systemsabstractConversational agents (CAs) have been widely used for many domains, such as healthcare, education, and business. One main category of CAs is task-oriented CAs, which aim to help users to complete a set of specific tasks. However, task-oriented CAs can fail to answer the user’s question, which can lead to a breakdown in the dialogue (when it is not possible to complete a conversation with a CA). Breakdown detection is an essential task for developing better CAs. Several related studies have focused on breakdown detection using different sets of features, for example, topic transition, word-based similarity and clustering; but, the existing studies develop features mainly from the system’s outputs or user’s inputs, whereas the features can be extracted from both sides, as well as from the interaction between them. Therefore, in this work, we developed a new supervised fusion machine learning (ML) model that combines the prediction from two machine learning algorithms for breakdown detection CAs services system. We developed features from different groups focusing on both the user input and the system response. Then we select the optimal combined features. The features are based on sentence similarity, sentiment features, and count-based features. The developed fusion model is mainly based on the two best performances of the single classifiers (SVM and RF). We explore several single ML algorithms using different sets of features and the combined features. To verify the effectiveness of the proposed fusion model, we compared the proposed models against baseline methods using four sets of data. We conclude that the proposed fusion model with the combined features outperforms the baselines and all other models in terms of prediction accuracy and f-score measures. Mohammed Aldahash, Yuefeng Li 0001, Yue Xu 0001 |
Web Intell. | 3 |
| 2023 | Generating multi-level explanations for process outcome predictionsabstractProcess mining focuses on the analysis of event log data to build various process analytical capabilities. Predictive process analytics has emerged as one of such key capabilities and it uses machine learning techniques to construct process prediction models. In recent years, deep neural networks have gained increasing interest in process prediction since they can handle multi-dimensional sequential inputs with minimal information loss. However, they are considered black-box models and existing studies in explaining deep neural network-based process predictions rely on only event-level features for explanation. In this paper, we propose a new approach for generating explanations for process outcome predictions at multiple levels. The approach is underpinned by three different prediction models: a transparent model for generating global explanations based on case-level features, an attention-based deep neural network for generating local explanations based on event-level features, and a novel eXplainable Dual-learning Deep network (XD2-net) for generating local explanations based on case-level features. Using three publicly available datasets, we have tested the applicability of the approach and further examined the multi-level explanations generated by the approach through an elaborate case study. Unlike others, the design of our approach promotes the idea of leveraging the complementary capabilities of different models and utilizing their strengths, rather than focusing on model performance competition. This will contribute towards generating more comprehensive explanations that meet the needs of different end users and purposes in the future. Bemali Wickramanayake, Chun Ouyang 0001, Yue Xu 0001, Catarina Moreira |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | DAC-HPP: deep attributed clustering with high-order proximity preserveabstractAbstract Attributed graph clustering, the task of grouping nodes into communities using both graph structure and node attributes, is a fundamental problem in graph analysis. Recent approaches have utilized deep learning for node embedding followed by conventional clustering methods. However, these methods often suffer from the limitations of relying on the original network structure, which may be inadequate for clustering due to sparsity and noise, and using separate approaches that yield suboptimal embeddings for clustering. To address these limitations, we propose a novel method called Deep Attributed Clustering with High-order Proximity Preserve (DAC-HPP) for attributed graph clustering. DAC-HPP leverages an end-to-end deep clustering framework that integrates high-order proximities and fosters structural cohesiveness and attribute homogeneity. We introduce a modified Random Walk with Restart that captures k-order structural and attribute information, enabling the modelling of interactions between network structure and high-order proximities. A consensus matrix representation is constructed by combining diverse proximity measures, and a deep joint clustering approach is employed to leverage the complementary strengths of embedding and clustering. In summary, DAC-HPP offers a unique solution for attributed graph clustering by incorporating high-order proximities and employing an end-to-end deep clustering framework. Extensive experiments demonstrate its effectiveness, showcasing its superiority over existing methods. Evaluation on synthetic and real networks demonstrates that DAC-HPP outperforms seven state-of-the-art approaches, confirming its potential for advancing attributed graph clustering research. Kamal Berahmand, Yuefeng Li 0001, Yue Xu 0001 |
Neural Comput. Appl. | 3 |
| 2023 | H-DAC: discriminative associative classification in data streamsabstractAbstract In this paper, we propose an efficient and highly accurate method for data stream classification, called discriminative associative classification. We define class discriminative association rules (CDARs) as the class association rules (CARs) in one data stream that have higher support compared with the same rules in the rest of the data streams. Compared to associative classification mining in a single data stream, there are additional challenges in the discriminative associative classification mining in multiple data streams, as the Apriori property of the subset is not applicable. The proposed single-pass H-DAC algorithm is designed based on distinguishing features of the rules to improve classification accuracy and efficiency. Continuously arriving transactions are inserted at fast speed and large volume, and CDARs are discovered in the tilted-time window model. The data structures are dynamically adjusted in offline time intervals to reflect each rule supported in different periods. Empirical analysis shows the effectiveness of the proposed method in the large fast speed data streams. Good efficiency is achieved for batch processing of small and large datasets, plus 0–2% improvements in classification accuracy using the tilted-time window model (i.e., almost with zero overhead). These improvements are seen only for the first 32 incoming batches in the scale of our experiments and we expect better results as the data streams grow. Majid Seyfi, Yue Xu 0001 |
Soft Comput. | 2 |
| 2022 | Semantic-based Attention model for Hospital Readmission PredictionabstractUnexpected hospital readmissions are problematic to both hospitals and patients, and also costly in terms of money and resources as well. Thus, identifying whether or not a patient will be readmitted becomes an important task. Due to the limitation of traditional machine learning techniques, Recurrent neural networks (RNN) and attention mechanisms have been proposed to learn temporal relationships between patient admissions for readmission prediction. Especially, learning accurate representations to represent medical concepts plays a crucial role for generating accurate predictions. Existing works demonstrate that incorporating medical ontologies can be beneficial to learn accurate medical code representations. However, current works ignore the importance of the textual descriptions associated with the medical codes in medical ontologies. In fact, these textual descriptions provide valuable information about the meaning of the medical codes and also the semantic relations among them. In this paper, we propose a new model called Semantic-based Attention model for Medical Concept Embedding (SAMCE). It takes into consideration the semantic information provided in code descriptions to learn medical code representations based on which to generate readmission prediction via RNN and attention mechanisms. Experimental results show that the proposed SAMCE model improves not only the performance of readmission prediction but also the quality of medical code representations. Sea Jung Im, Yue Xu 0001, Jason Watson |
IEEE Big Data | 2 |
| 2022 | SCEVD: Semantic-enhanced Code Embedding for Vulnerability DiscoveryabstractSource code vulnerability detection is a major goal in security research. In recent years, deep learning methods have been applied to this end, however the task of embedding code into vector representations as input for deep learning models has yet to be definitively solved. The use of graphs, specifically Abstract Syntax Trees and Code Property Graphs, is a promising research direction for this task, however learning from graphs grows prohibitively computationally expensive for large graphs. No close examination of intelligent ways to prune this input to only vulnerability-relevant information has yet been performed. Additionally, most existing works focus largely on structural information from graphs, often neglecting information contained within the nodes themselves. We address these gaps in the prior research by proposing SCEVD: a deep learning model for vulnerability discovery which utilises semantic information to intelligently select features in source code graphs for learning. It uses information contained within code graph nodes, as well as information about their relationships with one another to select the code graph features which are most relevant to code vulnerability. We implement SCEVD and conduct experiments using the SARD Juliet test suite, finding that we are able to improve vulnerability discovery results using this process of semantic-enhanced code graph feature selection. Joseph Gear, Yue Xu 0001, Ernest Foo, Praveen Gauravaram, Zahra Jadidi, Leonie Ruth Simpson |
TrustCom | 2 |
| 2022 | Enhanced Topic Representation by Ambiguity Handling
Dakshi T. K. Geeganage, Yue Xu 0001, Darshika N. Koggalahewa, Yuefeng Li 0001 |
WISE | 2 |
| 2022 | Building interpretable models for business process prediction using shared and specialised attention mechanisms
Bemali Wickramanayake, Zhipeng He 0002, Chun Ouyang 0001, Catarina Moreira, Yue Xu 0001, Renuka Sindhgatta |
Knowl. Based Syst. | 5 |
| 2021 | Predicting Alzheimer's Disease from Spoken and Written Language Using Fusion-Based Stacked Generalization
Ahmed H. Alkenani, Yuefeng Li 0001, Yue Xu 0001, Qing Zhang 0001 |
J. Biomed. Informatics | 3 |
| 2021 | Mining discriminative itemsets in data streams using the tilted-time window model
Majid Seyfi, Richi Nayak, Yue Xu 0001, Shlomo Geva |
Knowl. Inf. Syst. | 3 |
| 2021 | Semantic-based topic representation using frequent semantic patterns
Dakshi T. K. Geeganage, Yue Xu 0001, Yuefeng Li 0001 |
Knowl. Based Syst. | 2 |
| 2020 | Review selection based on content quality
Nan Tian, Yue Xu 0001, Yuefeng Li 0001 |
Knowl. Inf. Syst. | 2 |
| 2020 | Jointly Learning Topics in Sentence Embedding for Document SummarizationabstractSummarization systems for various applications, such as opinion mining, online news services, and answering questions, have attracted increasing attention in recent years. These tasks are complicated, and a classic representation using bag-of-words does not adequately meet the comprehensive needs of applications that rely on sentence extraction. In this paper, we focus on representing sentences as continuous vectors as a basis for measuring relevance between user needs and candidate sentences in source documents. Embedding models based on distributed vector representations are often used in the summarization community because, through cosine similarity, they simplify sentence relevance when comparing two sentences or a sentence/query and a document. However, the vector-based embedding models do not typically account for the salience of a sentence, and this is a very necessary part of document summarization. To incorporate sentence salience, we developed a model, called CCTSenEmb, that learns latent discriminative Gaussian topics in the embedding space and extended the new framework by seamlessly incorporating both topic and sentence embedding into one summarization system. To facilitate the semantic coherence between sentences in the framework of prediction-based tasks for sentence embedding, the CCTSenEmb further considers the associations between neighboring sentences. As a result, this novel sentence embedding framework combines sentence representations, word-based content, and topic assignments to predict the representation of the next sentence. A series of experiments with the DUC datasets validate CCTSenEmb's efficacy in document summarization in a query-focused extraction-based setting and an unsupervised ILP-based setting. Yang Gao 0016, Yue Xu 0001, Heyan Huang, Qian Liu 0012, Linjing Wei |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2020 | Query-based unsupervised learning for improving social media search
Khaled Albishre, Yuefeng Li 0001, Yue Xu 0001 |
World Wide Web | 3 |
| 2019 | Dual pattern-enhanced representations model for query-focused multi-document summarisation
Yutong Wu 0001, Yuefeng Li 0001, Yue Xu 0001 |
Knowl. Based Syst. | 3 |
| 2018 | Enhancing Binary Classification by Modeling Uncertain Boundary in Three-Way Decisions (Extended Abstract)abstractText classification techniques are playing a crucial role in identifying relevant texts from a large data set, e.g., various online crimes such as Cyberbullying, terrorist recruiting, propaganda or attack planning. Until now, supervised deep learning has brought about breakthroughs in processing multimedia data; however, there was no good practical way to harvest this opportunity for text classification because acquiring and maintaining a massive amount of training examples are too expensive for a large number of categories (e.g., Yahoo! taxonomy contains nearly 300,000 categories and the Library of Congress Subject Headings (LCSH) contains 394,070 subjects). Therefore, the question of how to effectively learn from sparse or small set of training examples is crucial for the true success of text classification. Semi-supervised approaches have been proposed for this challenge, which usually use a pair or several existing classifiers to extend a small training set. However, extracted pseudo training samples are uncertain because they are determined by a machine rather than people. Also, the massive volume and high variability of text data are creating a number of challenging issues such as the scalability and complicated relations between words. There are two fundamental issues with regards to the performance of existing classifiers: overlook and overload. Overlook means that some objects relevant to a class have been omitted, whereas overload means that some objects assigned to a class are actually not relevant to that class. The two issues are even more serious in the following two cases: (1) large uncertain boundary - the decision boundary between two classes includes many mixed examples (e.g., relevant and nonrelevant documents together), and (2) unbalanced classes - one class (e.g., information about terrorist attacks) is much smaller than another class (e.g., normal descriptions). We propose a three-way decision model [1] for dealing with the uncertain boundary for improving text classification performance based on rough set techniques and centroid solution. It aims to understand the uncertain boundary through partitioning the training samples into three regions (the positive, boundary and negative regions) by two main boundary vectors created from the labeled positive and negative training subsets, respectively, and further resolve the objects in the boundary region by two derived boundary vectors produced according to the structure of the boundary region. Four decision rules are proposed from the training process and applied to the incoming documents for more precise classification. The experimental results on the standard data sets RCV1 and Reuters-21578 show that the usage of boundary vectors is very effective and efficient for dealing with uncertainties of the decision boundary, and the proposed model has significantly improved the performance of binary text classification in terms of F1 measure and AUC area compared with six other popular baseline models. Yuefeng Li 0001, Libiao Zhang, Yue Xu 0001, Yiyu Yao, Raymond Y. K. Lau, Yutong Wu 0001 |
ICDE | 3 |
| 2018 | Query-Based Automatic Training Set Selection for Microblog Retrieval
Khaled Albishre, Yuefeng Li 0001, Yue Xu 0001 |
PAKDD (2) | 3 |
| 2018 | An Extended Random-Sets Model for Fusion-Based Text Feature Selection
Abdullah Semran Alharbi, Yuefeng Li 0001, Yue Xu 0001 |
PAKDD (3) | 3 |
| 2018 | A Semantic Similarity Based Topic Evaluation for Enhancing Information FilteringabstractTopic Modelling has been applied in many successful applications in data mining, text mining, machine learning and information filtering. The limitation is that the quality of topics generated from modelled corpus are not always good because many topics contain intrusive and ambiguous words. This negative drawback would affect the performance of text based application systems based on topic models. Hence, topic evaluation to assess and to rank the topics is really important for the good quality topics before applying those topics to text based applications. In this study, we proposed an ontology-based topic evaluation method for enhancing information filtering, named as STRbTCM. This new model assesses the quality of topics by matching topic models with headings in Library Congress Subject Heading (LCSH) ontology. To evaluate the effectiveness of our proposed model, we compare the model with two existing topic evaluation methods applied to information filtering system. In addition, we also compare our proposed model to term-based model BM25 and two other models based on topics: TNG and LDA_words. Through extensive experiments, we find that our proposed model performed better than other baseline models according to four main evaluating measures. Hanh Nguyen, Yue Xu 0001, Yuefeng Li 0001 |
WI | 2 |
| 2018 | Investigation of the Quality of Topic Models for Noisy Data SourcesabstractLatent Dirichlet Allocation (LDA) has become the most stable and widely used topic model to derive topics from collections of documents where it depicts different levels of success based on diversified domains of inputs. Nevertheless, it is a vital requirement to evaluate the LDA against the quality of the input. The noise and uncertainty of the content create a negative influence on the topic model. The major contribution of this investigation is to critically evaluate the LDA based on the quality of input sources and human perception. The empirical study shows the relationship between the quality of the input and the accuracy of the output generated by LDA. Perplexity and coherence have been evaluated with three data-sets (RCV1, conference data set, tweets) which contain different level of complexities and uncertainty in their contents. Human perception in generating topics has been compared with the LDA in terms of human defined topics. Results of the analysis demonstrate a strong relationship between the quality of the input and generated topics. Thus, highly relevant topics were generated from formally written contents while noisy and messy contents lead to generate meaningless topics. A considerable gap is noticed between human defined topics and LDA generated topics. Finally, a concept-based topic modeling technique is proposed to improve the quality of topics by capturing the meaning of the content and eliminating the irrelevant and meaningless topics. Yue Xu 0001, Yuefeng Li 0001, Dakshi T. K. Geeganage |
WI | 1 |
| 2018 | A Hybrid Approach for Detecting Spammers in Online Social Networks
Bandar Alghamdi, Yue Xu 0001, Jason Watson |
WISE (1) | 2 |
| 2018 | Cost-sensitive and hybrid-attribute measure multi-decision tree over imbalanced data sets
Fenglian Li, Xiqian Zhang, Chunlei Du, Yue Xu 0001, Yu-Chu Tian |
Inf. Sci. | 5 |
| 2017 | Topical term weighting based on extended random sets for relevance feature selectionabstractIt is challenging to discover relevant features from long documents that describe user information needs due to the nature of text where synonymy, polysemy noise, and high dimensionality are inherited problems. Traditional feature selection methods could not effectively deal with these problems, because they assume that documents describe one topic only. Topic-based techniques, such as Latent Dirichlet Allocation (LDA), relax this assumption. They have been developed on the basis that a document can exhibit multiple hidden topics. However, LDA does not show encouraging results in selecting relevant features, because LDA calculates the weight of terms based on their local documents and does not generalise it globally at the collection level. So as to address this problem, we propose an innovative and effective extended random set model to generalise LDA weight for local document terms. The model is used as a weighting scheme for topical terms. It can assign a more discriminately accurate weight to these terms based on their appearance in LDA topics and relevant documents. The experimental results, based on the standard RCV1 dataset, TREC topics, and five standard performance measures, show that the proposed model significantly outperforms eight state-of-the-art baseline models in information filtering. Abdullah Semran Alharbi, Yuefeng Li 0001, Yue Xu 0001 |
WI | 3 |
| 2017 | Efficient mining of discriminative itemsetsabstractDiscriminative itemsets can be more useful than frequent itemsets as the former identifies the frequent itemsets in one dataset with much higher frequencies than the same itemsets in other datasets. The discriminative itemsets can distinguish the target dataset from all others. The discriminative itemsets are a small subset of frequent itemsets. The efficient mining of discriminative itemsets is a challenging problem, since the Apriori property of frequent itemsets is not applicable, and the designed algorithms must deal with the exponential number of itemset combinations in more than one dataset. In this paper, a novel algorithm, called DISSparse, is proposed for efficient mining of discriminative itemsets. Two determinative heuristics are proposed for limiting the mining of discriminative itemsets to the potential discriminative itemsets. Our experiments show the efficient time and space usage of the proposed algorithm in the large and complex datasets. Majid Seyfi, Richi Nayak, Yue Xu 0001, Shlomo Geva |
WI | 3 |
| 2017 | An empirical study on the susceptibility to social engineering in social networking sites: the case of FacebookabstractResearch suggests that social engineering attacks pose a significant security risk, with social networking sites (SNSs) being the most common source of these attacks. Recent studies showed that social engineers could succeed even among those organizations that identify themselves as being aware of social engineering techniques. Although organizations recognize the serious risks of social engineering, there is little understanding and control of such threats. This may be partly due to the complexity of human behaviors in failing to recognize attackers in SNSs. Due to the vital role that impersonation plays in influencing users to fall victim to social engineering deception, this paper aims to investigate the impact of source characteristics on users’ susceptibility to social engineering victimization on Facebook. In doing so, we identify source credibility dimensions in terms of social engineering on Facebook, Facebook-based source characteristics that influence users to judge an attacker as per these dimensions, and mediation effects that these dimensions play between Facebook-based source characteristics and susceptibility to social engineering victimization. Abdullah Algarni, Yue Xu 0001, Taizan Chan |
Eur. J. Inf. Syst. | 2 |
| 2017 | Finding Semantically Valid and Relevant Topics by Association-Based Topic Selection ModelabstractTopic modelling methods such as Latent Dirichlet Allocation (LDA) have been successfully applied to various fields, since these methods can effectively characterize document collections by using a mixture of semantically rich topics. So far, many models have been proposed. However, the existing models typically outperform on full analysis on the whole collection to find all topics but difficult to capture coherent and specifically meaningful topic representations. Furthermore, it is very challenging to incorporate user preferences into existing topic modelling methods to extract relevant topics. To address these problems, we develop a novel personalized Association-based Topic Selection (ATS) model, which can identify semantically valid and relevant topics from a set of raw topics based on the semantical relatedness between users’ preferences and the structured patterns captured in topics. The advantage of the proposed ATS model is that it enables an interactive topic modelling process driven by users’ specific interests. Based on three benchmark datasets, namely, RCV1, R8, and WT10G under the context of information filtering (IF) and information retrieval (IR), our rigorous experiments show that the proposed ATS model can effectively identify relevant topics with respect to users’ specific interests, and hence to improve the performance of IF and IR. Yang Gao 0016, Yuefeng Li 0001, Raymond Y. K. Lau, Yue Xu 0001, Md. Abul Bashar |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2017 | Enhancing Binary Classification by Modeling Uncertain Boundary in Three-Way DecisionsabstractText classification is a process of classifying documents into predefined categories through different classifiers learned from labelled or unlabelled training samples. Many researchers who work on binary text classification attempt to find a more effective way to separate relevant texts from a large data set. However, current text classifiers cannot unambiguously describe the decision boundary between positive and negative objects because of uncertainties caused by text feature selection and the knowledge learning process. This paper proposes a three-way decision model for dealing with the uncertain boundary to improve the binary text classification performance based on therough settechniques and centroid solution. It aims to understand the uncertain boundary through partitioning the training samples into three regions (the positive, boundary, and negative regions) by two main boundary vectors$\vec{C_{P}}$and$\vec{C_{N}}$, created from the labeled positive and negative training subsets, respectively, and further resolve the objects in the boundary region by two derived boundary vectors$\vec{B_{P}}$and$\vec{B_{N}}$, produced according to the structure of the boundary region. It involves an indirect strategy which is composed of two successive steps in the whole classification process: ‘two-way to three-way’ and ‘three-way to two-way’. Four decision rules are proposed from the training process and applied to the incoming documents for more precise classification. A large number of experiments have been conducted based on the standard data sets RCV1 and Reuters-21578. The experimental results show that the usage of boundary vectors is very effective and efficient for dealing with uncertainties of the decision boundary, and the proposed model has significantly improved the performance of binary text classification in terms of$F_{1}$measure and$AUC$area compared with six other popular baseline models. Yuefeng Li 0001, Libiao Zhang, Yue Xu 0001, Yiyu Yao, Raymond Y. K. Lau, Yutong Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2016 | Finding Anomalies in SCADA Logs Using Rare Sequential Pattern Mining
Anisur Rahman, Yue Xu 0001, Kenneth Radke, Ernest Foo |
NSS | 2 |
| 2016 | Specialized Review Selection Using Topic Models
Nan Tian, Yue Xu 0001, Yuefeng Li 0001 |
PKAW | 3 |
| 2016 | Mining Topically Coherent Patterns for Unsupervised Extractive Multi-document SummarizationabstractAddressing the problem of information overload, automatic multi-document summarization (MDS) has been widely utilized in the various real-world applications. Most of existing approaches adopt term-based representation for documents which limit the performance of MDS systems. In this paper, we proposed a novel unsupervised pattern-enhanced topic model (PETMSum) for the MDS task. PETMSum combining pattern mining techniques with LDA topic modelling could generate discriminative and semantic rich representations for topics and documents so that the most representative, non-redundant, and topically coherent sentences can be selected automatically to form a succinct and informative summary. Extensive experiments are conducted on the data of document understanding conference (DUC) 2006 and 2007. The results prove the effectiveness and efficiency of our proposed approach. Yutong Wu 0001, Yuefeng Li 0001, Yue Xu 0001 |
WI | 3 |
| 2015 | An accurate rating aggregation method for generating item reputationabstractMany websites presently provide the facility for users to rate items quality based on user opinion. These ratings are used later to produce item reputation scores. The majority of websites apply the mean method to aggregate user ratings. This method is very simple and is not considered as an accurate aggregator. Many methods have been proposed to make aggregators produce more accurate reputation scores. In the majority of proposed methods the authors use extra information about the rating providers or about the context (e.g. time) in which the rating was given. However, this information is not available all the time. In such cases these methods produce reputation scores using the mean method or other alternative simple methods. In this paper, we propose a novel reputation model that generates more accurate item reputation scores based on collected ratings only. Our proposed model embeds statistical data, previously disregarded, of a given rating dataset in order to enhance the accuracy of the generated reputation scores. In more detail, we use the Beta distribution to produce weights for ratings and aggregate ratings using the weighted mean method. Experiments show that the proposed model exhibits performance superior to that of current state-of-the-art models. Ahmad Abdel-Hafez, Yue Xu 0001, Audun Jøsang |
DSAA | 2 |
| 2015 | Learning Higher-Order Interactions for User and Item Profiling Based on Tensor FactorizationabstractUser profiling techniques play a central role in many Recommender Systems (RS). In recent years, multidimensional data are getting increasing attention for making recommendations. Additional metadata help algorithms better understanding users' behaviors and decisions. Existing user/item profiling techniques for Collaborative Filtering (CF) RS in multidimensional environment mostly analyze data through splitting the multidimensional relations. However, this leads to the loss of multidimensionality in user-item interactions; whereas the interactions are naturally multidimensional since users' choices are often affected by contextual information. In this paper, we propose a unified profiling approach which models users/items with latent higher-order interaction factors. We demonstrate that the proposed profiling approach is intimately related to two-dimensional profiling based on Matrix Factorization techniques. We further propose to integrate the profiling approach into three neighborhood-based CF recommenders for item recommendation. Finally, we empirically show on real-world social tagging datasets that the proposed recommenders outperform state-of-the-art CF recommendation approaches in accuracy. Yue Xu 0001, Shlomo Geva |
IUI | 2 |
| 2015 | Pattern-based Topics for Document Modelling in Information FilteringabstractMany mature term-based or pattern-based approaches have been used in the field of information filtering to generate users' information needs from a collection of documents. A fundamental assumption for these approaches is that the documents in the collection are all about one topic. However, in reality users' interests can be diverse and the documents in the collection often involve multiple topics. Topic modelling, such as Latent Dirichlet Allocation (LDA), was proposed to generate statistical models to represent multiple topics in a collection of documents, and this has been widely utilized in the fields of machine learning and information retrieval, etc. But its effectiveness in information filtering has not been so well explored. Patterns are always thought to be more discriminative than single terms for describing documents. However, the enormous amount of discovered patterns hinder them from being effectively and efficiently used in real applications, therefore, selection of the most discriminative and representative patterns from the huge amount of discovered patterns becomes crucial. To deal with the above mentioned limitations and problems, in this paper, a novel information filtering model, Maximum matched Pattern-based Topic Model (MPBTM), is proposed. The main distinctive features of the proposed model include: (1) user information needs are generated in terms of multiple topics; (2) each topic is represented by patterns; (3) patterns are generated from topic models and are organized in terms of their statistical and taxonomic features; and (4) the most discriminative and representative patterns, called Maximum Matched Patterns, are proposed to estimate the document relevance to the user's information needs in order to filter out irrelevant documents. Extensive experiments are conducted to evaluate the effectiveness of the proposed model by using the TREC data collection Reuters Corpus Volume 1. The results show that the proposed model significantly outperforms both state-of-the-art term-based models and pattern-based models. Yang Gao 0016, Yue Xu 0001, Yuefeng Li 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2015 | A normal-distribution based rating aggregation method for generating product reputationsabstractWith the extensive use of rating systems in the web, and their significance in decision making process by users, the need for more accurate aggregation methods has emerged. The Naïve aggregation method, using the simple mean, is not adequate anymore in providing accurate reputation scores for items [6], hence, several researches where conducted in order to provide more accurate alternative aggregation methods. Most of the current reputation models do not consider the distribution of ratings across the different possible ratings values. In this paper, we propose a novel reputation model, which generates more accurate reputation scores for items by deploying the normal distribution over ratings. Experiments show promising results for our proposed model over state-of-the-art ones on sparse and dense datasets. Ahmad Abdel-Hafez, Yue Xu 0001, Audun Jøsang |
Web Intell. | 2 |
| 2015 | PrefaceabstractThe explosive growth of resources available through the Internet, especially the emergence of social media, has created highly interactive platforms for users to create, share, exchange information and build social networks. This special issue includ Yue Xu 0001, Gabriella Pasi |
Web Intell. | 1 |
| 2014 | A Reputation-Enhanced Recommender System
Ahmad Abdel-Hafez, Nan Tian, Yue Xu 0001 |
ADMA | 4 |
| 2014 | Centroid Training to achieve effective text classificationabstractTraditional text classification technology based on machine learning and data mining techniques has made a big progress. However, it is still a big problem on how to draw an exact decision boundary between relevant and irrelevant objects in binary classification due to much uncertainty produced in the process of the traditional algorithms. The proposed model CTTC (Centroid Training for Text Classification) aims to build an uncertainty boundary to absorb as many indeterminate objects as possible so as to elevate the certainty of the relevant and irrelevant groups through the centroid clustering and training process. The clustering starts from the two training subsets labelled as relevant or irrelevant respectively to create two principal centroid vectors by which all the training samples are further separated into three groups: POS, NEG and BND, with all the indeterminate objects absorbed into the uncertain decision boundary BND. Two pairs of centroid vectors are proposed to be trained and optimized through the subsequent iterative multi-learning process, all of which are proposed to collaboratively help predict the polarities of the incoming objects thereafter. For the assessment of the proposed model, F1and Accuracy have been chosen as the key evaluation measures. We stress the F1measure because it can display the overall performance improvement of the final classifier better than Accuracy. A large number of experiments have been completed using the proposed model on the Reuters Corpus Volume 1 (RCV1) which is important standard dataset in the field. The experiment results show that the proposed model has significantly improved the binary text classification performance in both F1and Accuracy compared with three other influential baseline models. Libiao Zhang, Yuefeng Li 0001, Yue Xu 0001, Dian Tjondronegoro |
DSAA | 3 |
| 2014 | Item Reputation-Aware Recommender SystemsabstractRecommender systems provide personalized advice for online customers based on their own preferences, while reputation systems generate a community advice on the quality of items on the Web. Both systems employ users' ratings to generate their output. In this paper, we aim to combine reputation models with recommender systems to enhance the accuracy of recommendations. Our proposed methods make two contributions. First of all, we propose two methods for merging two ranked item lists which are generated based on recommendation scores and reputation scores, respectively. In addition, a novel personalized reputation method is designed in order to generate item reputations based upon users' interests. The proposed merging methods can be applicable to any recommendation methods and reputation methods, i.e., they are independent from generating recommendation scores and reputation scores. The experiments we conducted showed that the proposed methods could enhance the accuracy of existing recommender systems. Ahmad Abdel-Hafez, Yue Xu 0001, Nan Tian |
iiWAS | 2 |
| 2014 | Refining User and Item Profiles based on Multidimensional Data for Top-N Item RecommendationabstractIn recommender systems based on multidimensional data, additional metadata provides algorithms with more information for better understanding the interaction between users and items. However, most of the profiling approaches in neighbourhood-based recommendation approaches for multidimensional data merely split or project the dimensional data and lack the consideration of latent interaction between the dimensions of the data. In this paper, we propose a novel user/item profiling approach for Collaborative Filtering (CF) item recommendation on multidimensional data. We further present incremental profiling method for updating the profiles. For item recommendation, we seek to delve into different types of relations in data to understand the interaction between users and items more fully, and propose three multidimensional CF recommendation approaches for top-N item recommendations based on the proposed user/item profiles. The proposed multidimensional CF approaches are capable of incorporating not only localized relations of user-user and/or item-item neighbourhoods but also latent interaction between all dimensions of the data. Experimental results show significant improvements in terms of recommendation accuracy. Yue Xu 0001, Shlomo Geva |
iiWAS | 2 |
| 2014 | A Normal-Distribution Based Reputation Model
Ahmad Abdel-Hafez, Yue Xu 0001, Audun Jøsang |
TrustBus | 2 |
| 2014 | Product Feature Taxonomy Learning based on User ReviewsabstractIn recent years, the Web 2.0 has provided considerable facilities for people to create, share and exchange information and ideas. Upon this, the user generated content, such as reviews, has exploded. Such data provide a rich source to exploit in order to identify the information associated with specific reviewed items. Opinion mining has been widely used to identify the significant features of items (e.g., cameras) based upon user reviews. Feature extraction is the most critical step to identify useful information from texts. Most existing approaches only find individual features about a product without revealing the structural relationships between the features which usually exist. In this paper, we propose an approach to extract features and feature relationships, represented as a tree structure called feature taxonomy, based on frequent patterns and associations between patterns derived from user reviews. The generated feature taxonomy profiles the product at multiple levels and provides more detailed information about the product. Our experiment results based on some popularly used review datasets show that our proposed approach is able to capture the product features and relations effectively. Nan Tian, Yue Xu 0001, Yuefeng Li 0001, Ahmad Abdel-Hafez, Audun Jøsang |
WEBIST (2) | 2 |
| 2014 | Topical Pattern Based Document Modelling and Relevance Ranking
Yang Gao 0016, Yue Xu 0001, Yuefeng Li 0001 |
WISE (1) | 2 |
| 2014 | A Review Selection Method Using Product Feature Taxonomy
Nan Tian, Yue Xu 0001, Yuefeng Li 0001 |
WISE (1) | 2 |
| 2014 | An evaluation framework for cross-lingual link discovery
Ling-Xiang Tang, Shlomo Geva, Andrew Trotman, Yue Xu 0001, Kelly Y. Itakura |
Inf. Process. Manag. | 4 |
| 2013 | Combining Recommender and Reputation Systems to Produce Better Online Advice
Audun Jøsang, Guibing Guo, Maria Silvia Pini, Francesco Santini 0001, Yue Xu 0001 |
MDAI | 5 |
| 2013 | A Two-Stage Approach for Generating Topic Models
Yang Gao 0016, Yue Xu 0001, Yuefeng Li 0001 |
PAKDD (2) | 2 |
| 2012 | Time-aware topic recommendation based on micro-blogsabstractTopic recommendation can help users deal with the information overload issue in micro-blogging communities. This paper proposes to use the implicit information network formed by the multiple relationships among users, topics and micro-blogs, and the temporal information of micro-blogs to find semantically and temporally relevant topics of each topic, and to profile users' time-drifting topic interests. The Content based, Nearest Neighborhood based and Matrix Factorization models are used to make personalized recommendations. The effectiveness of the proposed approaches is demonstrated in the experiments conducted on a real world dataset that collected from Twitter.com. Huizhi Liang 0001, Yue Xu 0001, Dian Tjondronegoro, Peter Christen |
CIKM | 2 |
| 2012 | Grouping people in social networks using a weighted multi-constraints clustering methodabstractGrouping users in social networks is an important process that improves matching and recommendation activities in social networks. The data mining methods of clustering can be used in grouping the users in social networks. However, the existing general purpose clustering algorithms perform poorly on the social network data due to the special nature of users' data in social networks. One main reason is the constraints that need to be considered in grouping users in social networks. Another reason is the need of capturing large amount of information about users which imposes computational complexity to an algorithm. In this paper, we propose a scalable and effective constraint-based clustering algorithm based on a global similarity measure that takes into consideration the users' constraints and their importance in social networks. Each constraint's importance is calculated based on the occurrence of this constraint in the dataset. Performance of the algorithm is demonstrated on a dataset obtained from an online dating website using internal and external evaluation measures. Results show that the proposed algorithm is able to increases the accuracy of matching users in social networks by 10% in comparison to other algorithms. Slah Alsaleh, Richi Nayak, Yue Xu 0001 |
FUZZ-IEEE | 3 |
| 2012 | Personalization in tag ontology learning for recommendation makingabstractDue to the explosive growth of the Web, the domain of Web personalization has gained great momentum both in the research and commercial areas. One of the most popular web personalization systems is recommender systems. In recommender systems choosing user information that can be used to profile users is very crucial for user profiling. In Web 2.0, one facility that can help users organize Web resources of their interest is user tagging systems. Exploring user tagging behavior provides a promising way for understanding users' information needs since tags are given directly by users. However, free and relatively uncontrolled vocabulary makes the user self-defined tags lack of standardization and semantic ambiguity. Also, the relationships among tags need to be explored since there are rich relationships among tags which could provide valuable information for us to better understand users. In this paper, we propose a novel approach for learning tag ontology based on the widely used lexical database WordNet for capturing the semantics and the structural relationships of tags. We present personalization strategies to disambiguate the semantics of tags by combining the opinion of WordNet lexicographers and users' tagging behavior together. To personalize further, clustering of users is performed to generate a more accurate ontology for a particular group of users. In order to evaluate the usefulness of the tag ontology, we use the tag ontology in a pilot tag recommendation experiment for improving the recommendation performance by exploiting the semantic information in the tag ontology. The initial result shows that the personalized information has improved the accuracy of the tag recommendation. Endang Djuana, Yue Xu 0001, Yuefeng Li 0001, Clive Cox |
iiWAS | 2 |
| 2012 | A two-stage decision model for information filtering
Yuefeng Li 0001, Xujuan Zhou, Peter Bruza, Yue Xu 0001, Raymond Y. K. Lau |
Decis. Support Syst. | 4 |
| 2012 | Text mining in negative relevance feedbackabstractIt is a big challenge to clearly identify the boundary between positive and negative streams. Several attempts have used negative feedback to solve this challenge; however, there are two issues for using negative relevance feedback to improve the eff Abdulmohsen Algarni, Yuefeng Li 0001, Sheng-Tang Wu, Yue Xu 0001 |
Web Intell. Agent Syst. | 4 |
| 2012 | Personalized recommender systems integrating tags and item taxonomyabstractTags in Web 2.0 are becoming another important information source to profile users' interests and preferences to make personalized recommendations. To solve the problem of low information sharing caused by the free-style vocabulary of tags and the lo Huizhi Liang 0001, Yue Xu 0001, Yuefeng Li 0001 |
Web Intell. Agent Syst. | 2 |
| 2011 | Improving Matching Process in Social Network Using Implicit and Explicit User Information
Slah Alsaleh, Richi Nayak, Yue Xu 0001, Lin Chen 0012 |
APWeb | 3 |
| 2011 | Finding and Matching Communities in Social Networks Using Data MiningabstractThe rapid growth in the number of users using social networks and the information that a social network requires about their users make the traditional matching systems insufficiently adept at matching users within social networks. This paper introduces the use of clustering to form communities of users and, then, uses these communities to generate matches. Forming communities within a social network helps to reduce the number of users that the matching system needs to consider, and helps to overcome other problems from which social networks suffer, such as the absence of user activities' information about a new user. The proposed system has been evaluated on a dataset obtained from an online dating website. Empirical analysis shows that accuracy of the matching process is increased using the community information. Slah Alsaleh, Richi Nayak, Yue Xu 0001 |
ASONAM | 3 |
| 2011 | A Recommendation Method for Online Dating Networks Based on Social Relations and Demographic InformationabstractA new relationship type of social networks - online dating - are gaining popularity. With a large member base, users of a dating network are overloaded with choices about their ideal partners. Recommendation methods can be utilized to overcome this problem. However, traditional recommendation methods do not work effectively for online dating networks where the dataset is sparse and large, and a two-way matching is required. This paper applies social networking concepts to solve the problem of developing a recommendation method for online dating networks. We propose a method by using clustering, SimRank and adapted SimRank algorithms to recommend matching candidates. Empirical results show that the proposed method can achieve nearly double the performance of the traditional collaborative filtering and common neighbor methods of recommendation. Lin Chen 0012, Richi Nayak, Yue Xu 0001 |
ASONAM | 3 |
| 2011 | Pattern Mining for a Two-Stage Information Filtering System
Xujuan Zhou, Yuefeng Li 0001, Peter Bruza, Yue Xu 0001, Raymond Y. K. Lau |
PAKDD (1) | 4 |
| 2011 | Reliable representations for association rules
Yue Xu 0001, Yuefeng Li 0001, Gavin Shaw |
Data Knowl. Eng. | 1 |
| 2011 | A pattern mining approach for information filtering systems
Yuefeng Li 0001, Abdulmohsen Algarni, Yue Xu 0001 |
Inf. Retr. | 3 |
| 2010 | Selected new training documents to update user profileabstractRelevance Feedback (RF) has been proven very effective for improving retrieval accuracy. Adaptive information filtering (AIF) technology has benefited from the improvements achieved in all the tasks involved over the last decades. A difficult problem in AIF has been how to update the system with new feedback efficiently and effectively. In current feedback methods, the updating processes focus on updating system parameters. In this paper, we developed a new approach, the Adaptive Relevance Features Discovery (ARFD). It automatically updates the system's knowledge based on a sliding window over positive and negative feedback to solve a nonmonotonic problem efficiently. Some of the new training documents will be selected using the knowledge that the system currently obtained. Then, specific features will be extracted from selected training documents. Different methods have been used to merge and revise the weights of features in a vector space. The new model is designed for Relevance Features Discovery (RFD), a pattern mining based approach, which uses negative relevance feedback to improve the quality of extracted features from positive feedback. Learning algorithms are also proposed to implement this approach on Reuters Corpus Volume 1 and TREC topics. Experiments show that the proposed approach can work efficiently and achieves the encouragement performance. Abdulmohsen Algarni, Yuefeng Li 0001, Yue Xu 0001 |
CIKM | 3 |
| 2010 | Personalized recommender system based on item taxonomy and folksonomyabstractItem folksonomy or tag information is popularly available on the web now. However, since tags are arbitrary words given by users, they contain a lot of noise such as tag synonyms, semantic ambiguities and personal tags. Such noise brings difficulties to improve the accuracy of item recommendations. In this paper, we propose to combine item taxonomy and folksonomy to reduce the noise of tags and make personalized item recommendations. The experiments conducted on the dataset collected from Amazon.com demonstrated the effectiveness of the proposed approaches. The results suggested that the recommendation accuracy can be further improved if we consider the viewpoints and the vocabularies of both experts and users. Huizhi Liang 0001, Yue Xu 0001, Yuefeng Li 0001, Richi Nayak |
CIKM | 2 |
| 2010 | Rough sets based reasoning and pattern mining for a two-stage information filtering systemabstractThis paper presents a novel two-stage information filtering model which combines the merits of term-based and pattern- based approaches to effectively filter sheer volume of infor- mation. In particular, the first filtering stage is supported by a novel rough analysis model which efficiently removes a large number of irrelevant documents, thereby addressing the overload problem. The second filtering stage is empow- ered by a semantically rich pattern taxonomy mining model which effectively fetches incoming documents according to the specific information needs of a user, thereby addressing the mismatch problem. The experiments have been conducted to compare the proposed two-stage filtering (T-SM) model with other possible "term-based + pattern-based" or "term-based + term-based" IF models. The results based on the RCV1 corpus show that the T-SM model significantly outperforms other types of "two-stage" IF models. Xujuan Zhou, Yuefeng Li 0001, Peter Bruza, Yue Xu 0001, Raymond Y. K. Lau |
CIKM | 4 |
| 2010 | SimTrust: A New Method of Trust Network GenerationabstractTrust can be used for neighbor formation to generate automated recommendations. User assigned explicit rating data can be used for this purpose. However, the explicit rating data is not always available. In this paper we present a new method of generating trust network based on user’s interest similarity. To identify the interest similarity, we use user’s personalized tag information. This trust network can be used to find the neighbors to make automated recommendation. Our experiment result shows that the precision of the proposed method outperforms the traditional collaborative filtering approach. Touhid Bhuiyan, Yue Xu 0001, Audun Jøsang |
EUC | 2 |
| 2010 | Using Association Rules to Solve the Cold-Start Problem in Recommender Systems
Gavin Shaw, Yue Xu 0001, Shlomo Geva |
PAKDD (1) | 2 |
| 2010 | Developing Trust Networks Based on User Tagging Information for Recommendation Making
Touhid Bhuiyan, Yue Xu 0001, Audun Jøsang, Huizhi Liang 0001, Clive Cox |
WISE | 2 |
| 2009 | An effective model of using negative relevance feedback for information filteringabstractOver the years, people have often held the hypothesis that negative feedback should be very useful for largely improving the performance of information filtering systems; however, we have not obtained very effective models to support this hypothesis. This paper, proposes an effective model that use negative relevance feedback based on a pattern mining approach to improve extracted features. This study focuses on two main issues of using negative relevance feedback: the selection of constructive negative examples to reduce the space of negative examples; and the revision of existing features based on the selected negative examples. The former selects some offender documents, where offender documents are negative documents that are most likely to be classified in the positive group. The later groups the extracted features into three groups: the positive specific category, general category and negative specific category to easily update the weight. An iterative algorithm is also proposed to implement this approach on RCV1 data collections, and substantial experiments show that the proposed approach achieves encouraging performance. Abdulmohsen Algarni, Yuefeng Li 0001, Yue Xu 0001, Raymond Y. K. Lau |
CIKM | 3 |
| 2009 | Thai Word Segmentation with Hidden Markov Model and Decision Tree
Poramin Bheganan, Richi Nayak, Yue Xu 0001 |
PAKDD | 3 |
| 2009 | Personalized Recommender Systems Integrating Social Tags and Item TaxonomyabstractThe social tags in web 2.0 are becoming another important information source to profile users' interests and preferences to make personalized recommendations. To solve the problem of low information sharing caused by the free-style vocabulary of tags and the long tails of the distribution of tags and items, this paper proposes an approach to integrate the social tags given by users and the item taxonomy with standard vocabulary and hierarchical structure provided by experts to make personalized recommendations. The experimental results show that the proposed approach can effectively improve the information sharing and recommendation accuracy. Huizhi Liang 0001, Yue Xu 0001, Yuefeng Li 0001, Richi Nayak, Li-Tung Weng |
Web Intelligence | 2 |
| 2008 | A two-stage text mining model for information filteringabstractMismatch and overload are the two fundamental issues regarding the effectiveness of information filtering. Both term-based and pattern (phrase) based approaches have been employed to address these issues. However, they all suffer from some limitations with regard to effectiveness. This paper proposes a novel solution that includes two stages: an initial topic filtering stage followed by a stage involving pattern taxonomy mining. The objective of the first stage is to address mismatch by quickly filtering out probable irrelevant documents. The threshold used in the first stage is motivated theoretically. The objective of the second stage is to address overload by apply pattern mining techniques to rationalize the data relevance of the reduced document set after the first stage. Substantial experiments on RCV1 show that the proposed solution achieves encouraging performance. Yuefeng Li 0001, Xujuan Zhou, Peter Bruza, Yue Xu 0001, Raymond Y. K. Lau |
CIKM | 4 |
| 2008 | Deriving non-redundant approximate association rules from hierarchical datasetsabstractAssociation rule mining plays an important job in knowledge and information discovery. However, there are still shortcomings with the quality of the discovered rules and often the number of discovered rules is huge and contain redundancies, especially in the case of multi-level datasets. Previous work has shown that the mining of non-redundant rules is a promising approach to solving this problem, with work by [6,8,9,10] focusing on single level datasets. Recent work by Shaw et. al. [7] has extended the nonredundant approaches presented in [6,8,9] to include the elimination of redundant exact basis rules from multi-level datasets. Here we propose a continuation of the work in [7] that allows for the removal of hierarchically redundant approximate basis rules from multi-level datasets by using a dataset’s hierarchy or taxonomy. Gavin Shaw, Yue Xu 0001, Shlomo Geva |
CIKM | 2 |
| 2008 | A User Driven Data Mining Process Model and Learning System
Esther Ge, Richi Nayak, Yue Xu 0001, Yuefeng Li 0001 |
DASFAA | 3 |
| 2008 | Extracting Non-redundant Approximate Rules from Multi-level DatasetsabstractAssociation rule mining plays an important job in knowledge and information discovery. Often the number of the discovered rules is huge and many of them are redundant, especially for multi-level datasets. Previous work has shown that the mining of non-redundant rules is a promising approach to solving this problem, with work in focusing on single level datasets. Recent work by Shaw et. al. has extended the non-redundant approaches presented in to include the elimination of redundant exact basis rules from multi-level datasets. In this paper, we propose an extension to the work in to allow for the removal of hierarchically redundant approximate basis rules from multi-level datasets through the use of the datasetpsilas hierarchy or taxonomy. Experimentation shows our approach can effectively generate both multi-level and cross level non-redundant rule sets which are lossless. Gavin Shaw, Yue Xu 0001, Shlomo Geva |
ICTAI (2) | 2 |
| 2008 | Exploiting Item Taxonomy for Solving Cold-Start Problem in Recommendation MakingabstractRecommender systems' performance can be easily affected when there are no sufficient item preferences data provided by previous users, and it is commonly referred to as cold-start problem. This paper suggests another information source, item taxonomies, in addition to item preference data for assisting recommendation making. Item taxonomy information has been popularly applied in diverse ecommerce domains for product or content classification, and therefore can be easily obtained and adapted by recommender systems. In this paper, we investigate the implicit relations between users' item preferences and taxonomic preferences, suggest and verify that users who share similar item preferences may also share similar taxonomic preferences. Under this assumption, a novel recommendation technique is proposed that combines the users' item preferences and the additional taxonomic preferences together to make better quality recommendations as well as alleviate the cold-start problem. Empirical evaluations to this approach are conducted and the results show that the proposed technique outperforms other existing techniques in both recommendation quality and computation efficiency. Li-Tung Weng, Yue Xu 0001, Yuefeng Li 0001, Richi Nayak |
ICTAI (2) | 2 |
| 2008 | Concise representations for approximate association rulesabstractThe quality of association rule mining has drawn more and more attention recently. One problem with the quality of the discovered association rules is the huge size of the extracted rule set. Often for a dataset, a huge number of rules can be extracted, but many of them can be redundant to other rules and thus useless in practice. Mining non-redundant rules is a promising approach to solve this problem. In this paper, we firstly propose a definition for redundancy; then we propose a concise representation called reliable basis for representing non-redundant association rules for both exact rules and approximate rules. We prove that the redundancy elimination based on the reliable basis does not reduce the belief to the extracted rules. We also prove that all association rules can be deduced from the reliable basis. Therefore the reliable basis is a lossless representation of association rules. Experimental results show that the reliable basis significantly reduces the number of extracted rules. Yue Xu 0001, Yuefeng Li 0001, Gavin Shaw |
SMC | 1 |
| 2008 | Combining Trust and Reputation Management for Web-Based Services
Audun Jøsang, Touhid Bhuiyan, Yue Xu 0001, Clive Cox |
TrustBus | 3 |
| 2007 | Generating concise association rulesabstractAssociation rule mining has made many achievements in the area of knowledge discovery. However, the quality of the extracted association rules is a big concern. One problem with the quality of the extracted association rules is the huge size of the extracted rule set. As a matter of fact, very often tens of thousands of association rules are extracted among which many are redundant thus useless. Mining non-redundant rules is a promising approach to solve this problem. The Min-max exact basis proposed by Pasquier et al [Pasquier05] has showed exciting results by generating only non-redundant rules. In this paper, we first propose a relaxing definition for redundancy under which the Min-max exact basis still contains redundant rules; then we propose a condensed representation called Reliable exact basis for exact association rules. The rules in the Reliable exact basis are not only non-redundant but also more succinct than the rules in Min-max exact basis. We prove that the redundancy eliminated by the Reliable exact basis does not reduce the belief to the Reliable exact basis. The size of the Reliable exact basis is much smaller than that of the Min-max exact basis. Moreover, we prove that all exact association rules can be deduced from the Reliable exact basis. Therefore the Reliable exact basis is a lossless representation of exact association rules. Experimental results show that the Reliable exact basis significantly reduces the number of non-redundant rules. Yue Xu 0001, Yuefeng Li 0001 |
CIKM | 1 |
| 2007 | Granule Based Intertransaction Association Rule MiningabstractIntertransaction association rule mining is used to discover patterns between different transactions. It breaks the scope of association rule mining on the same transaction. Currently the FITI algorithm is the state of the art in intertransaction association rule mining. However, the FTTI introduces many unneeded combinations of items because the set of extended items is much larger than the set of items. Thus, we propose an alternative approach of granule based intertransaction association rule mining, where a granule is a group of transactions that meet a certain constraint. The experimental results show that this approach is promising in real-world industry. Wanzhong Yang, Yuefeng Li 0001, Yue Xu 0001 |
ICTAI (1) | 3 |
| 2007 | Collection Profiling for Collection Fusion in Distributed Information Retrieval Systems
Chengye Lu, Yue Xu 0001, Shlomo Geva |
KSEM | 2 |
| 2007 | Mining Fuzzy Domain Ontology from Textual DatabasesabstractOntology plays an essential role in the formalization of common information (e.g., products, services, relationships of businesses) for effective human-computer interactions. However, engineering of these ontologies turns out to be very labor intensive and time consuming. Although some text mining methods have been proposed for automatic or semi-automatic discovery of crisp ontologies, the robustness, accuracy, and computational efficiency of these methods need to be improved to support large scale ontology construction for real-world applications. This paper illustrates a novel fuzzy domain ontology mining algorithm for supporting real-world ontology engineering. In particular, contextual information of the knowledge sources is exploited for the extraction of high quality domain ontologies and the uncertainty embedded in the knowledge sources is modeled based on the notion of fuzzy sets. Empirical studies have confirmed that the proposed method can discover high quality fuzzy domain ontology which leads to significant improvement in information retrieval performance. Raymond Y. K. Lau, Yuefeng Li 0001, Yue Xu 0001 |
Web Intelligence | 3 |
| 2007 | Using Information Filtering in Web Data Mining ProcessabstractThe amount of Web information is growing rapidly, improving the efficiency and accuracy of Web information retrieval is uphill battle. There are two fundamental issues regarding the effectiveness of Web information gathering: information mismatch and overload. To tackle these difficult issues, an integrated information filtering and sophisticated data processing model has been presented in this paper. In the first phase of the proposed scheme, an information filter that based on user search intents was incorporated in Web search process to quickly filter out irrelevant data. In the second data processing phase, a pattern taxonomy model (PTM) was carried out using the reduced data. PTM rationalizes the data relevance by applying data mining techniques that involves more rigorous computations. Several experiments have been conducted and the results show that more effective and efficient access Web information has been achieved using the new scheme. Xujuan Zhou, Yuefeng Li 0001, Peter Bruza, Sheng-Tang Wu, Yue Xu 0001, Raymond Y. K. Lau |
Web Intelligence | 5 |
| 2007 | Mining Non-Redundant Association Rules Based on Concise BasesabstractAssociation rule mining has many achievements in the area of knowledge discovery. However, the quality of the extracted association rules has not drawn adequate attention from researchers in data mining community. One big concern with the quality of association rule mining is the size of the extracted rule set. As a matter of fact, very often tens of thousands of association rules are extracted among which many are redundant, thus useless. In this paper, we first analyze the redundancy problem in association rules and then propose a reliable exact association rule basis from which more concise nonredundant rules can be extracted. We prove that the redundancy eliminated using the proposed reliable association rule basis does not reduce the belief to the extracted rules. Moreover, this paper proposes a level wise approach for efficiently extracting closed itemsets and minimal generators — a key issue in closure based association rule mining. Yue Xu 0001, Yuefeng Li 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2006 | Multi-Tier Granule Mining for Representations of Multidimensional Association RulesabstractIt is a big challenge to promise the quality of multidimensional association mining. The essential issue is how to represent meaningful multidimensional association rules efficiently. Currently we have not found satisfactory approaches for solving this challenge because of the complicated correlation between attributes. Multi-tier granule mining is an initiative for solving this challenging issue. It divides attributes into some tiers and then compresses the large multidimensional database into granules at each tier. It also builds association mappings to illustrate the correlation between tiers. In this way, the meaningful association rules can be justified according to these association mappings. Yuefeng Li 0001, Wanzhong Yang, Yue Xu 0001 |
ICDM | 3 |
| 2006 | Deploying Approaches for Pattern Refinement in Text MiningabstractText mining is the technique that helps users find useful information from a large amount of digital text documents on the Web or databases. Instead of the keyword-based approach which is typically used in this field, the pattern-based model containing frequent sequential patterns is employed to perform the same concept of tasks. However, how to effectively use these discovered patterns is still a big challenge. In this study, we propose two approaches based on the use of pattern deploying strategies. The performance of the pattern deploying algorithms for text mining is investigated on the Reuters dataset RCVI and the results show that the effectiveness is improved by using our proposed pattern refinement approaches. Sheng-Tang Wu, Yuefeng Li 0001, Yue Xu 0001 |
ICDM | 3 |
| 2006 | Distributed Recommender Profiling and Selection with Gittins IndicesabstractMost existing recommender systems nowadays operate in a single organizational base, and very often they do not have sufficient resources to be used in order to generate quality recommendations. Therefore, it would be beneficial if recommender systems of different organizations can cooperate together to share their resources and recommendations. In this paper, we present a distributed recommender system model that consists of multiple recommender systems from different organizations. With the hope to provide better recommendation service to users, the recommender systems can improve their performances by sharing their recommendations cooperatively. A recommender selection technique based on the Gittins indices is presented in this paper, and it makes selections based on the stability, average performance and selection frequency of the recommenders Li-Tung Weng, Yue Xu 0001, Yuefeng Li 0001, Richi Nayak |
Web Intelligence | 2 |
| 2006 | Utilizing Search Intent in Topic Ontology-Based User Profile for Web MiningabstractIt is well known that taking the Web user profiles into account can enhance the effectiveness of Web mining systems. However, due to the dynamic and complex nature of Web users, automatically acquiring worthwhile user profiles was found to be very challenging. Ontology-based user profile can possess more accurate user information. This research emphasizes on acquiring search intentions information. This paper presents a new approach of developing user profile for Web searching. The model considers the user's search intentions by the process of PTM (Pattern-Taxonomy Model). Initial experiments show that the user profile based on search intention is more useful than the generic PTM user profile. Developing user profile that contains user search intentions is essential for effective Web search and retrieval. Xujuan Zhou, Sheng-Tang Wu, Yuefeng Li 0001, Yue Xu 0001, Raymond Y. K. Lau, Peter Bruza |
Web Intelligence | 4 |
| 2005 | An Effective Deploying Algorithm for Using Pattern-Taxonomy
Sheng-Tang Wu, Yuefeng Li 0001, Yue Xu 0001 |
iiWAS | 3 |
| 2004 | Automatic Pattern-Taxonomy Extraction for Web MiningabstractIn this paper, we propose a model for discovering frequent sequential patterns, phrases, which can be used as profile descriptors of documents. It is indubitable that we can obtain numerous phrases using data mining algorithms. However, it is difficult to use these phrases effectively for answering what users want. Therefore, we present a pattern taxonomy extraction model which performs the task of extracting descriptive frequent sequential patterns by pruning the meaningless ones. The model then is extended and tested by applying it to the information filtering system. The results of the experiment show that pattern-based methods outperform the keyword-based methods. The results also indicate that removal of meaningless patterns not only reduces the cost of computation but also improves the effectiveness of the system. Sheng-Tang Wu, Yuefeng Li 0001, Yue Xu 0001, Binh Pham 0001, Yi-Ping Phoebe Chen |
Web Intelligence | 3 |