EDBT 2026 Demo / reviewers in the wild / expert
Xiaoying Gao
dblp:79/86
· DBLP profile ↗
80ranked-venue papers
7as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 59 · 7 first-author · 22 since 2021Databases, data management, data science and information retrieval · 24 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Human-computer interaction and ubiquitous computing · 2Computer networks · 1Security and privacy · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Boundary-Aware Multi-Objective Genetic Programming Classifier with Stacking-Based IntegrationabstractGenetic Programming (GP) has been widely adopted for classifier construction, owing to its flexible representation and inherent feature construction capability. However, existing GP-based classifiers mainly focus on optimizing global performance metrics, while paying limited attention to regions near the decision boundary. Moreover, under multi-objective formulations, the final classifier is often constructed by selecting a single best solution from the evolved solution set or applying majority voting to it, which may not fully leverage the diverse decision behaviors of evolved classifiers. In this paper, we propose a boundary-aware multi-objective GP framework. The evolutionary search balances overall classification performance and model complexity, while a boundary-aware criterion is incorporated into the environmental selection stage as an auxiliary selection bias favoring classifiers with more balanced and well-distributed local decision behavior. After evolution, a selective stacking strategy is employed to select a compact subset of representative classifiers and integrate their raw outputs through a lightweight meta-classifier. Experiments on seven datasets show that incorporating the boundary-aware selection into GP evolution improves classification performance, while the selective stacking strategy yields further gains. The proposed method outperforms single-objective GP baselines and five multi-objective GP methods employing different integration strategies. Kaiwei Yuan, Xiaoying Gao, Jianbin Ma |
GECCO | 2 |
| 2026 | Conditional Information Extraction with Diffusion Model on Fact-Condition Star GraphabstractConditional Knowledge Graphs (CKGs) extend traditional knowledge graphs by incorporating conditional constraints, enabling a more accurate understanding of complex knowledge with conditional constraints for the semantic web. Conditional information extraction (CIE) aims to extract not only traditional fact triples but also their corresponding conditional qualifiers, forming quintuples that represents these constraints. Existing CIE methods typically treat conditional quintuples as flat structures, overlooking the hierarchical dependencies. Additionally, they often require exploring all possible mention combinations, leading to a large interaction space. These two issues hinder the extraction performance. To this end, we propose a Diffusion Model on Fact-condition Star Graph for CIE (Diff-CIE). We adapt a star graph structure where fact triples serve as central nodes and conditional tuples as leaf nodes, explicitly modeling the hierarchical dependencies. We then leverage the diffusion model to reformulate CIE as a progressive denoising process on these nodes, refining a fixed number of noised nodes into quintuples, thereby reducing the interaction space. Furthermore, to mitigate the inherent optimization instability in traditional diffusion-based information extraction methods, we introduce a deterministic in-order matching strategy to provide an auxiliary constraint. Extensive experiments on three datasets demonstrate that Diff-CIE consistently outperforms state-of-the-art baselines and has higher efficiency, achieving an improvement in F1 metric of over 1.19%, validating the effectiveness of our methods. Yunxiao Yang, Jianting Chen, Xiaoying Gao, Zaiyuan Di, Yang Xiang 0006 |
WWW | 3 |
| 2026 | A diffusion-driven multi-view mixed contrastive learning framework for bundle recommendation
Xiaoying Gao, Jianting Chen, Yunxiao Yang, Zaiyuan Di, Yang Xiang 0006 |
Expert Syst. Appl. | 1 |
| 2026 | Intent disentangling model with hypergraph for next POI recommendation
Xiaoying Gao, Ling Ding 0003, Jianting Chen, Yujian Mo, Yunxiao Yang, Zaiyuan Di, Zhihao Wang 0005, Yang Xiang 0006 |
Expert Syst. Appl. | 1 |
| 2026 | Knowledge adapting and soft retrieval: Leveraging large language models for uncertain knowledge graph reasoning
Yunxiao Yang, Jianting Chen, Xiaoying Gao, Zaiyuan Di, Yang Xiang 0006 |
Knowl. Based Syst. | 3 |
| 2026 | Developing distance-based genetic programming classifiers by reconstructing datasets for imbalanced binary classification
Wenyang Meng, Xiaoying Gao, Jianbin Ma |
Pattern Recognit. | 4 |
| 2025 | Multi-Objective Genetic Programming for Imbalanced Classification with Adaptive Thresholds and a New Fitness FunctionabstractGenetic programming (GP) is widely used for classifier construction due to its flexible representation and feature construction characteristics. Traditional GP methods, however, often rely on a fixed threshold, typically 0, which fails to reflect the true distribution of the data in imbalanced datasets. To overcome this, we propose a multi-objective GP method that adaptively adjusts the threshold during evolution using Youden's Index. This adaptive threshold adjustment allows the classifiers to better fit the data distribution. Additionally, we introduce a class separation metric, distt, aimed at enhancing the clarity of the classification boundaries and improving the generalization ability of the evolved classifiers. We use the multi-objective GP, along with the optimal threshold of each classifier, to jointly optimize the accuracy of the minority and majority classes, as well as the class separation metric distt, selecting the best classifier from the Pareto front for unseen data. Experiments on 7 imbalanced datasets demonstrate that our method outperforms single-objective GP with fixed thresholds and four GP-based algorithms, showcasing superior performance and improved classification clarity. Furthermore, our proposed clarity metric distt improves classification performance, ensuring better generalization and enhanced decision boundaries. Minghui Bai, Xiaoying Gao, Jiaxin Niu, Jianbin Ma |
GECCO | 2 |
| 2025 | Advancing Comprehensive Aspect-Based Sentiment Analysis with Generative Models
Bisma Ayaz, Xiaoying Gao, Bing Xue 0001 |
PAKDD (7) | 2 |
| 2025 | Advancing Rubric-Based Automated Essay Scoring with Multi-view BERT: A Case Study in New Zealand
Xiaoying Gao, Yi Mei 0001 |
PAKDD (7) | 2 |
| 2025 | LLM-Based Simulation Tool for Clinician-Patient Communication Training: A Dual-Mode AI Approach
Magezi Julius, Junhong Zhao, Xiaoying Gao, Jon Herries, Melita MacDonald, Brad Peckler |
PRICAI | 3 |
| 2025 | User group-enhanced user feature distribution transfer framework for non-overlapping cross-domain recommendations
Xiaoying Gao, Ling Ding 0003, Jianting Chen, Yunxiao Yang, Yang Xiang 0006 |
Knowl. Based Syst. | 1 |
| 2024 | Autonomous Aspect-Image Instruction a2II: Q-Former Guided Multimodal Sentiment ClassificationabstractMultimodal aspect-oriented sentiment classification (MABSC) task has garnered significant attention, which aims to identify the sentiment polarities of aspects by combining both language and vision information. However, the limited multimodal data in this task has become a big gap for the vision-language multimodal fusion. While large-scale vision-language pretrained models have been adapted to multiple tasks, their use for MABSC task is still in a nascent stage. In this work, we present an attempt to use the instruction tuning paradigm to MABSC task and leverage the ability of large vision-language models to alleviate the limitation in the fusion of textual and image modalities. To tackle the problem of potential irrelevance between aspects and images, we propose a plug-and-play selector to autonomously choose the most appropriate instruction from the instruction pool, thereby reducing the impact of irrelevant image noise on the final sentiment classification results. We conduct extensive experiments in various scenarios and our model achieves state-of-the-art performance on benchmark datasets, as well as in few-shot settings. Junjia Feng, Mingqian Lin, Lin Shang 0001, Xiaoying Gao |
LREC/COLING | 4 |
| 2024 | An Ontology-based Three-Stage Approach to Medical Text classification with Feature Selection by Particle Swarm OptimisationabstractThe document classification (DC) task assigns predefined classes to unlabeled documents using trained models. In the medical field, DC is crucial for tasks like categorizing risk factors and classifying electronic health records. This paper addresses challenges in medical document analysis, such as the prevalence of abbreviations and acronyms. Existing classification performance in medical documents is suboptimal. The paper introduces novel feature engineering methods leveraging domain-specific knowledge to enhance classification performance. Results indicate that the Three-Stage approach surpasses related works, showcasing improved medical document classification performance. Mahdi Abdollahi, Xiaoying Gao, Yi Mei 0001, Shameek Ghosh, Jinyan Li 0001, Michael Narag |
KES | 2 |
| 2024 | SeCor: Aligning Semantic and Collaborative Representations by Large Language Models for Next-Point-of-Interest RecommendationsabstractThe widespread adoption of location-based applications has created a growing demand for point-of-interest (POI) recommendation, which aims to predict a user’s next POI based on their historical check-in data and current location. However, existing methods often struggle to capture the intricate relationships within check-in data. This is largely due to their limitations in representing temporal and spatial information and underutilizing rich semantic features. While large language models (LLMs) offer powerful semantic comprehension to solve them, they are limited by hallucination and the inability to incorporate global collaborative information. To address these issues, we propose a novel method SeCor, which treats POI recommendation as a multi-modal task and integrates semantic and collaborative representations to form an efficient hybrid encoding. SeCor first employs a basic collaborative filtering model to mine interaction features. These embeddings, as one modal information, are fed into LLM to align with semantic representation, leading to efficient hybrid embeddings. To mitigate the hallucination, SeCor recommends based on the hybrid embeddings rather than directly using the LLM’s output text. Extensive experiments on three public real-world datasets show that SeCor outperforms all baselines, achieving improved recommendation performance by effectively integrating collaborative and semantic information through LLMs. Shirui Wang, Bohan Xie, Ling Ding 0003, Xiaoying Gao, Jianting Chen, Yang Xiang 0006 |
RecSys | 4 |
| 2024 | Dual De-confounded Causal Intervention method for knowledge graph error detection
Yunxiao Yang, Jianting Chen, Xiaoying Gao, Yang Xiang 0006 |
Knowl. Based Syst. | 3 |
| 2024 | Genetic Programming for Document Classification: A Transductive Transfer Learning SystemabstractDocument classification is a challenging task to the data being high-dimensional and sparse. Many transfer learning methods have been investigated for improving the classification performance by effectively transferring knowledge from a source domain to a target domain, which is similar to but different from the source domain. However, most of the existing methods cannot handle the case that the training data of the target domain does not have labels. In this study, we propose a transductive transfer learning system, utilizing solutions evolved by genetic programming (GP) on a source domain to automatically pseudolabel the training data in the target domain in order to train classifiers. Different from many other transfer learning techniques, the proposed system pseudolabels target-domain training data to retrains classifiers using all target-domain features. The proposed method is examined on nine transfer learning tasks, and the results show that the proposed transductive GP system has better prediction accuracy on the test data in the target domain than existing transfer learning approaches including subspace alignment-domain adaptation methods, feature-level-domain adaptation methods, and one latest pseudolabeling strategy-based method. Bing Xue 0001, Xiaoying Gao, Mengjie Zhang 0001 |
IEEE Trans. Cybern. | 3 |
| 2023 | Multi-head attention based candidate segment selection in QA over hybrid dataabstractQuestion Answering based on Tabular and Textual data is a novel task proposed in recent years in the field of QA. At present, most QA systems return answers from a single data form, such as knowledge graphs, tables, texts. However, hybrid data including structured and unstructured data is quite pervasive in real life instead of a single form. Recent research on TAT-QA mainly suffers from the higher error of extracting supporting evidences from both tabular and textual content. This paper aimed to address the problem of failure evidence extraction from more complex and realistic hybrid data. We first proposed two types of metrics to evaluate the performance of evidence extraction on hybrid data, i.e. wrong evidence ratio (WER) and missing evidence ratio (MER). Then we utilize a candidate extractor to obtain supporting evidence related to the question. Third, an origin selector is designed to determine from where the question’s answer comes. Finally, the loss of origin selector is fused to the final loss function, which can improve the evidence extraction performance. Experimental results on the TAT-QA dataset showed that our proposed model outperforms the best baseline in terms of F1, WER and MER, which proves the effectiveness of our model. Qian Chen 0023, Xiaoying Gao, Suge Wang |
Intell. Data Anal. | 2 |
| 2022 | Deep Structure-Aware Approach for QA Over Incomplete Knowledge Bases
Qian Chen 0023, Xiaoying Gao, Suge Wang |
NLPCC (1) | 2 |
| 2022 | A feature selection method with feature ranking using genetic programmingabstractFeature selection is a data processing method which aims to select effective feature subsets from original features. Feature selection based on evolutionary computation (EC) algorithms can often achieve better classification performance because of their global search ability. However, feature selection methods using EC cannot get rid of invalid features effectively. A small number of invalid features still exist till the termination of the algorithms. In this paper, a feature selection method using genetic programming (GP) combined with feature ranking (FRFS) is proposed. It is assumed that the more the original features appear in the GP individuals' terminal nodes, the more valuable these features are. To further decrease the number of selected features, FRFS using a multi-criteria fitness function which is named as MFRFS is investigated. Experiments on 15 datasets show that FRFS can obtain higher classification performance with smaller number of features compared with the feature selection method without feature ranking. MFRFS further reduces the number of features while maintaining the classification performance compared with FRFS. Comparisons with five benchmark techniques show that MFRFS can achieve better classification performance. Guopeng Liu, Jianbin Ma, Tongle Hu, Xiaoying Gao |
Connect. Sci. | 4 |
| 2022 | Token replacement-based data augmentation methods for hate speech detectionabstractAbstract Hate speech detection mostly involves the use of text data. This data, usually sourced from various social media platforms, have been known to be plagued with numerous issues that result in a reduction of its quality and hence, the quality of the trained models. Some of these issues are the lack of diversity and the diminutive class of interest in the dataset which results in overfitted models that do not generalize well on other or newly collected data. The different ways of handling these issues include augmenting the data with diverse samples, engineering non-redundant features or designing robust classification models. In this study, the focus is on the data augmentation aspect. Data augmentation is a popular method for improving the quality of existing datasets by generating synthetic samples that mimic the distribution of the original samples. There is a lack of extensive studies on how hate speech texts respond to varying textual data augmentation techniques and methods. Specifically, we provide further insight into the token replacement method of textual data augmentation by performing empirical studies that investigate which embedding method(s) is a robust source of synonym for replacement process, what effective method(s) can be used to select words to be replaced, and how to confirm if the label within each class is preserved. Our proposed methods, validated on two commonly used hate speech datasets affected by a known lack of diversity and diminutive class of interest issues, significantly improve classification performance and provides insights into token replacement methods. Kosisochukwu Judith Madukwe, Xiaoying Gao, Bing Xue 0001 |
World Wide Web | 2 |
| 2021 | Evolving Neural Networks for Text Classification using Genetic Algorithm-based ApproachesabstractConvolutional Neural Networks (CNNs) have been well-known for their promising performance in text classification and sentiment analysis because they can preserve the 1D spatial orientation of a document, where the sequence of words is essential. However, designing the network architecture of CNNs is by no means an easy task, since it requires domain knowledge from both the deep CNN and text classification areas, which are often not available and can increase operating costs for anyone wishing to implement this method. Furthermore, such domain knowledge is often different in different text classification problems. To resolve these issues, this paper proposes the use of Genetic Algorithm to automatically search for the optimal network architecture without requiring any intervention from experts. The proposed approach is applied on the IMDB dataset, and the experimental results show that it achieves competitive performance with the current state-of-the-art and manually-designed approaches in terms of accuracy, and also it requires only a few hours of training time. Hayden Andersen, Sean Stevenson, Tuan Ha, Xiaoying Gao, Bing Xue 0001 |
CEC | 4 |
| 2021 | Evolving Character-Level DenseNet Architectures Using Genetic Programming
Trevor Londt, Xiaoying Gao, Peter Andreae |
EvoApplications | 2 |
| 2021 | Fake News Detection Using Multiple-View Text Representation
Tuan Ha, Xiaoying Gao |
PRICAI (2) | 2 |
| 2021 | What Emotion Is Hate? Incorporating Emotion Information into the Hate Speech Detection Task
Kosisochukwu Judith Madukwe, Xiaoying Gao, Bing Xue 0001 |
PRICAI (2) | 2 |
| 2021 | Substituting clinical features using synthetic medical phrases: Medical text data augmentation techniques
Mahdi Abdollahi, Xiaoying Gao, Yi Mei 0001, Shameek Ghosh, Jinyan Li 0001, Michael Narag |
Artif. Intell. Medicine | 2 |
| 2021 | Output-based transfer learning in genetic programming for document classification
Bing Xue 0001, Xiaoying Gao, Mengjie Zhang 0001 |
Knowl. Based Syst. | 3 |
| 2021 | Sentiment Time Series Calibration for Event DetectionabstractEvent detection based on sentiment time series, which describe the trend of users' emotions or attitudes towards specific topics over time, has been widely applied in the analysis of social network or text mining. Most of the contributions directly generate time series sequences by classifiers. However, due to the missing corpus labels or the limited performance of the classifier, such generated sentiment time series may not correspond to the actual values, especially when the sentiment value changes drastically, called extreme value. We propose a new method to calibrate sentiment times series for event detection based on evaluation on a sampling dataset. Theoretical analysis of the calibration method is explicated, and it is proved that the sampling error of the performance indicators can be limited to a minimal range for extreme values, thus sentiment value error can be reduced. Experiments on simulated datasets and real-world datasets illustrate the effectiveness and robustness of our method. Lin Shang 0001, Xiaoying Gao |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Ontology-Guided Data Augmentation for Medical Document Classification
Mahdi Abdollahi, Xiaoying Gao, Yi Mei 0001, Shameek Ghosh, Jinyan Li 0001 |
AIME | 2 |
| 2020 | A filter-based feature construction and feature selection approach for classification using Genetic Programming
Jianbin Ma, Xiaoying Gao |
Knowl. Based Syst. | 2 |
| 2019 | An Ontology-based Two-Stage Approach to Medical Text Classification with Feature Selection by Particle Swarm OptimisationabstractDocument classification (DC) is the task of assigning pre-defined labels to unseen documents by utilizing a model trained on the available labeled documents. DC has attracted much attention in medical fields recently because many issues can be formulated as a classification problem. It can assist doctors in decision making and correct decisions can reduce the medical expenses. Medical documents have special attributes that distinguish them from other texts and make them difficult to analyze. For example, many acronyms and abbreviations, and short expressions make it more challenging to extract information. The classification accuracy of the current medical DC methods is not satisfactory. The goal of this work is to enhance the input feature sets of the DC method to improve the accuracy. To approach this goal, a novel two-stage approach is proposed. In the first stage, a domain-specific dictionary, namely the Unified Medical Language System (UMLS), is employed to extract the key features belonging to the most relevant concepts such as diseases or symptoms. In the second stage, PSO is applied to select more related features from the extracted features in the first stage. The performance of the proposed approach is evaluated on the 2010 Informatics for Integrating Biology and the Bedside (i2b2) data set which is a widely used medical text dataset. The experimental results show substantial improvement by the proposed method on the accuracy of classification. Mahdi Abdollahi, Xiaoying Gao, Yi Mei 0001, Shameek Ghosh, Jinyan Li 0001 |
CEC | 2 |
| 2019 | Genetic Programming based Transfer Learning for Document Classification with Self-taught and Ensemble LearningabstractDocument classification is a common but challenging task in text mining, since the feature set used is often high-dimensional and sparse. Transfer learning has been applied to improve the classification performance of a (target) domain by transferring knowledge from a previously learnt (source) domain. When there are no labels provided for documents in target domains, it is challenging to effectively transfer knowledge from source domains to target domains. In this paper, we develop a new Genetic Programming (GP) based transfer learning method for document classification, which utilises the evolved GP programs from the source domain to learn a set of weak GP classification models on the target domain with unlabelled documents, which is called self-taught learning. These weak classifiers are combined with the GP programs transferred from the source domain to predict the labels of test documents in the target domain. The experimental results show that the GP programs from source domains with their weak classifiers can effectively classify documents in the target domain. Bing Xue 0001, Xiaoying Gao, Mengjie Zhang 0001 |
CEC | 3 |
| 2019 | Stratifying Risk of Coronary Artery Disease Using Discriminative Knowledge-Guided Medical Concept Pairings from Clinical Notes
Mahdi Abdollahi, Xiaoying Gao, Yi Mei 0001, Shameek Ghosh, Jinyan Li 0001 |
PRICAI (3) | 2 |
| 2019 | DDI-PULearn: a positive-unlabeled learning method for large-scale prediction of drug-drug interactionsabstractBACKGROUND: Drug-drug interactions (DDIs) are a major concern in patients' medication. It's unfeasible to identify all potential DDIs using experimental methods which are time-consuming and expensive. Computational methods provide an effective strategy, however, facing challenges due to the lack of experimentally verified negative samples. RESULTS: To address this problem, we propose a novel positive-unlabeled learning method named DDI-PULearn for large-scale drug-drug-interaction predictions. DDI-PULearn first generates seeds of reliable negatives via OCSVM (one-class support vector machine) under a high-recall constraint and via the cosine-similarity based KNN (k-nearest neighbors) as well. Then trained with all the labeled positives (i.e., the validated DDIs) and the generated seed negatives, DDI-PULearn employs an iterative SVM to identify a set of entire reliable negatives from the unlabeled samples (i.e., the unobserved DDIs). Following that, DDI-PULearn represents all the labeled positives and the identified negatives as vectors of abundant drug properties by a similarity-based method. Finally, DDI-PULearn transforms these vectors into a lower-dimensional space via PCA (principal component analysis) and utilizes the compressed vectors as input for binary classifications. The performance of DDI-PULearn is evaluated on simulative prediction for 149,878 possible interactions between 548 drugs, comparing with two baseline methods and five state-of-the-art methods. Related experiment results show that the proposed method for the representation of DDIs characterizes them accurately. DDI-PULearn achieves superior performance owing to the identified reliable negatives, outperforming all other methods significantly. In addition, the predicted novel DDIs suggest that DDI-PULearn is capable to identify novel DDIs. CONCLUSIONS: The results demonstrate that positive-unlabeled learning paves a new way to tackle the problem caused by the lack of experimentally verified negatives in the computational prediction of DDIs. Yi Zheng 0002, Xiaocai Zhang, Zhixun Zhao, Xiaoying Gao, Jinyan Li 0001 |
BMC Bioinform. | 5 |
| 2019 | Old drug repositioning and new drug discovery through similarity learning from drug-target joint feature spacesabstractBACKGROUND: Detection of new drug-target interactions by computational algorithms is of crucial value to both old drug repositioning and new drug discovery. Existing machine-learning methods rely only on experimentally validated drug-target interactions (i.e., positive samples) for the predictions. Their performance is severely impeded by the lack of reliable negative samples. RESULTS: We propose a method to construct highly-reliable negative samples for drug target prediction by a pairwise drug-target similarity measurement and OCSVM with a high-recall constraint. On one hand, we measure the pairwise similarity between every two drug-target interactions by combining the chemical similarity between their drugs and the Gene Ontology-based similarity between their targets. Then we calculate the accumulative similarity with all known drug-target interactions for each unobserved drug-target interaction. On the other hand, we obtain the signed distance from OCSVM learned from the known interactions with high recall (≥0.95) for each unobserved drug-target interaction. After normalizing all accumulative similarities and signed distances to the range [0,1], we compute the score for each unobserved drug-target interaction via averaging its accumulative similarity and signed distance. Unobserved interactions with lower scores are preferentially served as reliable negative samples for the classification algorithms. The performance of the proposed method is evaluated on the interaction data between 1094 drugs and 1556 target proteins. Extensive comparison experiments using four classical classifiers and one domain predictive method demonstrate the superior performance of the proposed method. A better decision boundary has been learned from the constructed reliable negative samples. CONCLUSIONS: Proper construction of highly-reliable negative samples can help the classification models learn a clear decision boundary which contributes to the performance improvement. Yi Zheng 0002, Xiaocai Zhang, Zhixun Zhao, Xiaoying Gao, Jinyan Li 0001 |
BMC Bioinform. | 5 |
| 2018 | Particle Swarm Optimization Based Two-Stage Feature Selection in Text MiningabstractText mining is an important and popular data mining topic, where a fundamental objective is to enable users to extract informative data from text-based assets and perform related operations on the text, like retrieval, classification, and summarization. For text classification, one of the most important steps is feature selection, because not all the features in the text dataset are useful for classification. Irrelevant and redundant features should be removed to increase the accuracy and decrease the complexity and running time, but it is often an expensive process, and most existing methods using a simple filter to remove features, which might potentially loose some useful ones because of feature interactions. Furthermore, there is little research using particle swarm optimization (PSO) algorithms to select informative features for text classification. This paper presents an approach using a novel two-stage method for text feature selection, where with the features selected by four different filter ranking methods at the first stage, more irrelevant features are removed by PSO to compose the final feature subset. The proposed algorithm is compared with four traditional feature selection methods on the commonly used Reuter-21578 dataset. The experimental results show that the proposed two-stage method can substantially reduce the dimensionality of the feature space and improve the classification accuracy. Xiaohan Bai, Xiaoying Gao, Bing Xue 0001 |
CEC | 2 |
| 2018 | Predicting Drug Targets from Heterogeneous Spaces using Anchor Graph Hashing and Ensemble LearningabstractThe in silico prediction of potential drug-targetinteractions is of critical importance in drug research. Existing computational methods have achieved remarkable prediction accuracy, however usually obtain poor prediction efficiency due to computational problems. To improve the prediction efficiency, we propose to predict drug targets based on inte- gration of heterogeneous features with anchor graph hashing and ensemble learning. First, we encode each drug as a 5682- bit vector, and each target as a 4198-bit vector using their heterogeneous features respectively. Then, these vectors are embedded into low-dimensional Hamming Space using anchor graph hashing. Next, we append hashing bits of a target to hashing bits of a drug as a vector to represent the drug-target pair. Finally, vectors of positive samples composed of known drug-target pairs and randomly selected negative samples are used to train and evaluate the ensemble learning model. The performance of the proposed method is evaluated on simulative target prediction of 1094 drugs from DrugBank. Extensive comparison experiments demonstrate that the proposed method can achieve high prediction efficiency while preserving satisfactory accuracy. In fact, it is 99.3 times faster and only 0.001 less in AUC than the best literature method “Pairwise Kernel Method”. Yi Zheng 0002, Xiaocai Zhang, Xiaoying Gao, Jinyan Li 0001 |
IJCNN | 4 |
| 2017 | Cluster-than-Label: Semi-Supervised Approach for Domain AdaptationabstractThe performance of a conventional machine learning model trained on a source domain degrades poorly when they are tested on a different data distribution (target domain). These traditional models deal with this problem by training a new paradigm for the particular different data distribution (target domain). Therefore, training of a new paradigm for the individual data distribution is computationally expensive. This paper demonstrates that how to adapt to a new data distribution (target domain), utilising the model trained on source domain and avoiding the cost of re-training and the need for access to the source labelled data. In particular, we introduce an Efficient Semi-supervised Cluster-than-Label Cross-domain Adaptation Algorithm (SCTLCDA) to address the cross-domain adaptation classification problem in which we utilised both labelled and unlabelled data samples in the target domain, as well as completely unlabelled data samples in the source domain. Subsequently, we also describe that our proposed method can manage large datasets and easily lead to cross-domain adaptation problem. The effectiveness and performance of our method are confirmed by experiments on two real-world applications: Crossdomain sentiments and Web-Spam classification problem. Xiaoying Gao, Ian Welch |
AINA | 2 |
| 2016 | A Machine Learning Based Web Spam Filtering ApproachabstractWeb spam has the effect of polluting search engine results and decreasing the usefulness of search engines.Web spam can be classified according to the methods used to raise the web page's ranking by subverting web search engine's algorithms used to rank search results. The main types are: content spam, link spam and cloaking spam. There has been little or no work on automatically classifying web spam by type. This paper has two contributions, (i) we propose a Dual-Margin Multi-Class Hypersphere Support Vector Machine (DMMH- SVM) classifier approach to automatically classifying web spam by type, (ii) we introduce novel cloaking-based spam features which help our classifier model to achieve high precision and recall rate, thereby reducing the false positive rates. The effectiveness of the proposed model is justified analytically. Our experimental results demonstrated that DMMH-SVM outperforms existing algorithms with novel cloaking features. Xiaoying Gao, Ian Welch, Masood Mansoori |
AINA | 2 |
| 2016 | View-based text representationabstractDocument clustering is useful for many research areas such as Text Mining and Information Retrieval. Therefore, it is desirable to be able to cluster documents accurately. The clustering quality depends not only on the clustering algorithm used but also on the way text is represented in the algorithm. Text is typically represented using the All-Words Vector Space Model in text mining applications. However, this representation's effectiveness is limited when clustering by a specific criterion or view, such as geographical location, and also tends to be computationally expensive due to not reducing the feature space. This paper presents a new representation, the View-Based Vector Space Model, which models a document differently depending on a given view, and is computationally efficient. The basic representation requires predefined ontologies to be available. In order to improve the usefulness of this model, a new text representation learning method which uses Particle Swarm Optimisation to construct ontologies automatically given a set of documents is also presented. Chahine Koleejan, Xiaoying Gao |
CEC | 2 |
| 2016 | Novel Features for Web Spam DetectionabstractRecent research on web spam detection has shown promising results, and many new and efficient detection algorithms have been developed. While most research focuses on developing algorithms, our investigation shows that the features used in the algorithms are in fact very important, and different features can lead to very different results. This paper investigates three types of web spam, content-based, link-based and cloaking, and introduces new features for identifying the three types of spam. Our experimental results show that the introduction of new features significantly improves the detection performance. Xiaoying Gao, Ian Welch |
ICTAI | 2 |
| 2016 | Learning Under Data Shift for Domain Adaptation: A Model-Based Co-clustering Transfer Learning Solution
Xiaoying Gao, Ian Welch |
PKAW | 2 |
| 2016 | A Differential Evolution Approach to Feature Selection and Instance Selection
Jiaheng Wang 0004, Bing Xue 0001, Xiaoying Gao, Mengjie Zhang 0001 |
PRICAI | 3 |
| 2016 | Query aspects approach to web searchabstractThis paper introduces an aspect based model for analysing search queries, where queries are represented as aspects or concepts instead of “bag of words” or strings. A search query may consist of multiple aspects, while some aspects are covered well in the search results and some others are underrepresented. We think the underrepresented aspects are the main reason for the retrieval of irrelevant documents. This paper introduces novel algorithms that identify query aspects and identify underrepresented aspects. This model has many applications and this paper focuses on three of them: query difficulty prediction, query expansion and interactive query expansion. The main idea is that a hard query is a query with multiple aspects and some of the aspects are underrepresented, and a query can be improved or expanded by adding terms that are semantically related to the underrepresented aspects. Our experiments show that our aspect based methods significantly outperform existing methods. Daniel Crabtree, Xiaoying Gao, Peter Andreae |
Web Intell. | 2 |
| 2015 | Multi-objective multi-view clustering ensemble based on evolutionary approachabstractClustering ensembles is a clustering technique which derives a better clustering solution from a set of candidate clustering solutions. Clustering ensemble methods have to address two distinct but interlinked problems: Generating multiple candidate solutions from the data and producing a final clustering solution. Our recently proposed clustering ensembles method (MMOEA) based on NSGA-II used multiple views to address the first problem and a novel cluster oriented approach to address the second problem. MMOEA used a simple crossover method to explore the search space and three objective functions to determine the quality of a candidate clustering solution. The use of a simple crossover method led to slow convergence and using three objectives in NSGA-II framework is often discouraged. This paper presents a new clustering ensemble method, which introduces new ideas for crossover, mutation, tuning steps and two objective functions (instead of three) in an evolutionary process. The results show that our new method outperforms recent methods for clustering ensembles on different multi-view datasets. Xiaoying Gao, Peter Andreae |
CEC | 2 |
| 2015 | Multi-objective clustering ensemble for high-dimensional data based on Strength Pareto Evolutionary Algorithm (SPEA-II)abstractClustering is one of the fundamental data analysis techniques, which aims to find distinct groups of similar objects and discovers hidden structures in data. A recent clustering approach, clustering ensembles tries to derive an improved clustering solution based on previously generated different candidate clustering solutions. Clustering ensembles have two steps: generating multiple candidate clustering solutions from the data and forming a final clustering solution from previously generated candidate clustering solutions. A problem of the first step is the text representation, where word frequencies are often used as features. Other semantic information of the text such as topics, hypertext, etc are ignored. The problem for the second step is that the current popular median partition approach selects one clustering solution from previously generated candidate clustering solutions. A common clustering ensemble approach uses word frequencies as features to represent text data (documents). However, documents usually contain semantically rich information i.e. words, hypertext, titles, topics etc. The cluster ensemble approach ignores the semantic information of the documents and hence is prone to produce futile groupings of the documents. In this research work, we present a new multi-objective clustering ensemble method based on Strength Pareto Evolutionary Algorithm (SPEA-II). Our method utilizes the semantic information (rich features) to address the first problem of clustering ensembles. The cluster oriented evolutionary approach which derives the final clustering solution by selecting better quality clusters is in the second step of our method to address the second problem. The results show that our new method provides better results than other clustering ensemble methods. Xiaoying Gao, Peter Andreae |
DSAA | 2 |
| 2015 | A Soft Subspace Clustering Method for Text Data Using a Probability Based Feature Weighting Scheme
Xiaoying Gao, Peter Andreae |
WISE (2) | 2 |
| 2015 | DIKEA: Exploiting Wikipedia for keyphrase extractionabstractAutomatic keyphrase extraction is the challenging task of assigning keyphrases to documents to capture the main topics. It assists many research areas in the field of text mining – indexing, clustering, and summarisation. A landmark research KEA (Keyphrase Extraction Algorithm) formulated the probl em as a supervised machine learning problem and successfully applied a Naïve Bayes model to it. KEA showed great promise but its performance is not satisfactory. Its state-of-art extension KEA++ significantly improved its performance but relies on a domain specific vocabulary which is often not available or incomplete for other domains. We present a novel domain-independent system (DIKEA) which makes three main contributions to this field of research: utilising the largest online knowledge source available, Wikipedia, for keyphrase candidate selection; adding new features including a Wikipedia-based feature, link probability; and further boosting performance by using a multilayer perceptron network. Our experiments showed that DIKEA outperformed KEA++ while keeping the overall solution domain-independent. DIKEA was also tested on a benchmark dataset provided by a workshop on Semantic Evaluation (SemEval-2010), allowing comparisons with the 19 other related systems which participated. Our experiments show that DIKEA ranks first when considering only the top 5 keyphrases extracted from each document, and ranks second overall. David X. Wang, Xiaoying Gao, Peter Andreae |
Web Intell. | 2 |
| 2014 | Multi-view clustering of web documents using multi-objective genetic algorithmabstractClustering ensembles are a common approach to clustering problem, which combine a collection of clustering into a superior solution. The key issues are how to generate different candidate solutions and how to combine them. Common approach for generating candidate clustering solutions ignores the multiple representations of the data (i.e., multiple views) and the standard approach of simply selecting the best solution from candidate clustering solutions ignores the fact that there may be a set of clusters from different candidate clustering solutions which can form a better clustering solution. This paper presents a new clustering method that exploits multiple views to generate different clustering solutions and then selects a combination of clusters to form a final clustering solution. Our method is based on Nondominated Sorting Genetic Algorithm (NSGA-II), which is a multi-objective optimization approach. Our new method is compared with five existing algorithms on three data sets that have increasing difficulty. The results show that our method significantly outperforms other methods. Xiaoying Gao, Peter Andreae |
IEEE Congress on Evolutionary Computation | 2 |
| 2014 | Using Asymmetric Associations for Commonsense Causality Detection
Shahida Jabeen, Xiaoying Gao, Peter Andreae |
PRICAI | 2 |
| 2014 | Wallace: Incorporating Search into Chatting
Alexandre Sawczuk da Silva, Xiaoying Gao, Peter Andreae |
PRICAI | 2 |
| 2014 | Probabilistic Associations as a Proxy for Semantic Relatedness
Shahida Jabeen, Xiaoying Gao, Peter Andreae |
WISE (1) | 2 |
| 2014 | A Hybrid Model for Learning Semantic Relatedness Using Wikipedia-Based Features
Shahida Jabeen, Xiaoying Gao, Peter Andreae |
WISE (1) | 2 |
| 2013 | Detecting heap-spray attacks in drive-by downloads: Giving attackers a handabstractIn the anatomy of drive-by download attacks, one of the key steps is to place malicious code (shellcode) in the memory of the browser process in order to carry out a drive-by download attack. There are two common techniques to carry out this task: stack-based and heap-based injections. However, introduction of stack protection makes the stack-based injection harder to carry out successfully. The heap-based injections become common methods to deliver shellcode to the heap memory of the web browsers. This paper presents the role of heap-spray in drive-by download attacks. We propose a new detection mechanism which makes shellcode in heap-spray executed in order to detect drive-by download attack. The solution not only benefits detection of drive-by download attacks but also analysis of malware behavior. Van Lam Le, Ian Welch, Xiaoying Gao, Peter Komisarczuk |
LCN | 3 |
| 2013 | Directional Context Helps: Guiding Semantic Relatedness Computation by Asymmetric Word Associations
Shahida Jabeen, Xiaoying Gao, Peter Andreae |
WISE (1) | 2 |
| 2013 | Exploiting User Queries for Search Result Clustering
Xiaoying Gao, Peter Andreae |
WISE (1) | 2 |
| 2013 | Query directed clustering
Daniel Crabtree, Xiaoying Gao, Peter Andreae |
Knowl. Inf. Syst. | 2 |
| 2012 | Query Directed Web Page Clustering Using Suffix Tree and Wikipedia Links
John Park, Xiaoying Gao, Peter Andreae |
ADMA | 2 |
| 2012 | Harnessing Wikipedia Semantics for Computing Contextual Relatedness
Shahida Jabeen, Xiaoying Gao, Peter Andreae |
PRICAI | 2 |
| 2012 | Automatic Keyword Extraction from Single-Sentence Natural Language Queries
David X. Wang, Xiaoying Gao, Peter Andreae |
PRICAI | 2 |
| 2012 | A Novel Scoring Model to Detect Potential Malicious Web PagesabstractMalicious web pages have embedded within them active contents that exploit vulnerabilities in users' browsers and plug-ins in order to compromise the users' machines. Approaches from research into identifying malicious web pages can be classified into two groups depending upon the types of web page features used: either run-time features based upon observing what happens when the web page is loaded (slow but accurate) or static features based upon the content, structure or property of the web page (fast but inaccurate). Hybrid approaches combine the best of both to provide scalable systems with good accuracy by using the static feature based approach as a pre-filter for the run-time feature based approach. One of critical challenges for such hybrid approaches is to build effective pre-filter which has a capability to make the trade-off between reducing number of web pages passed through to the run-time feature detector and misidentifying malicious web pages as benign. This paper presents a novel scoring model to filter potential malicious web pages by using static features from various sources of information about malicious web pages, finding suitable algorithms to score maliciousness of each source of information, and finally finding the best ways to combine scores from different sources of information in order to achieve the best accuracy. The result shows that our novel scoring model can combine knowledge from various sources of information about web pages very effectively in order to filter potential malicious web pages. Van Lam Le, Ian Welch, Xiaoying Gao, Peter Komisarczuk |
TrustCom | 3 |
| 2011 | Improving Suffix Tree Clustering with New Ranking and Similarity Measures
Phiradit Worawitphinyo, Xiaoying Gao, Shahida Jabeen |
ADMA (2) | 2 |
| 2011 | Two-Stage Classification Model to Detect Malicious Web PagesabstractMalicious web pages are an emerging security concern on the Internet due to their popularity and their potential serious impacts. Detecting and analyzing them is very costly because of their qualities and complexities. There has been some research approaches carried out in order to detect them. The approaches can be classified into two main groups based on their used analysis features: static feature based and run-time feature based approaches. While static feature based approach shows it strengthens as light-weight system, run-time feature based approach has better performance in term of detection accuracy. This paper presents a novel two-stage classification model to detect malicious web pages. Our approach divided detection process into two stages: Estimating maliciousness of web pages and then identifying malicious web pages. Static features are light-weight but less valuable so they are used to identify potential malicious web pages in the first stage. Only potential malicious web pages are forwarded to the second stage for further investigation. On the other hand, run-time features are costly but more valuable so they are used in the final stage to identify malicious web pages. Van Lam Le, Ian Welch, Xiaoying Gao, Peter Komisarczuk |
AINA | 3 |
| 2011 | Phoneme Based Representation for Vietnamese Web Page ClassificationabstractThis paper proposes a novel text representation for Web pages written in Vietnamese. This representation is based on an analysis of Vietnamese documents at phonetic level in which each document will be represented as a bag of phonemes. It is designed to capture sound-based information in documents and to be helpful for resolving some non-topic text classification problems including automatic Vietnamese language identification of a document, ancient Vietnamese document detection, author identification, and poem identification. We apply some typical machine learning methods including NB, KNN and SVMs to build text classifiers. The experimental results show a significant improvement in terms of effectiveness and efficiency compared to the traditional syllable based representation in most cases. Giang-Son Nguyen, Xiaoying Gao, Peter Andreae |
Web Intelligence | 2 |
| 2010 | Improving AbraQ: An Automatic Query Expansion AlgorithmabstractOur previous research has developed AbraQ, an innovative automatic query expansion algorithm that automatically adds a term to a search query to improve the search results. AbraQ differs from other relevance feedback approaches in that it works independently of the quality of the original search result, which means it works well for hard search tasks when there are not any relevant documents retrieved for the original query. Our experiments showed that it significantly improved precision for hard search tasks with multi-aspect queries, while other query expansion techniques often improve recall with no positive effects on precision. This paper further introduces an improved version called AbraQ2, which changes the way in which aspect vocabularies are constructed, and introduces a new algorithm for automatic relevance judgments. Our experiments show that these improvements help to find better queries that return more relevant documents to the user. Glen Robertson, Xiaoying Gao |
Web Intelligence | 2 |
| 2010 | Knowledge acquisition method from domain text based on theme logic model and artificial neural network
Yunpeng Wu, Xuening Liu, Xiaoying Gao |
Expert Syst. Appl. | 4 |
| 2007 | Exploiting underrepresented query aspects for automatic query expansionabstractUsers attempt to express their search goals through web search queries. When a search goal has multiple components or aspects, documents that represent all the aspects are likely to be more relevant than those that only represent some aspects. Current web search engines often produce result sets whose top ranking documents represent only a subset of the query aspects. By expanding the query using the right keywords, the search engine can find documents that represent more query aspects and performance improves. This paper describes AbraQ, an approach for automatically finding the right keywords to expand the query. AbraQ identifies the aspects in the query, identifies which aspects are underrepresented in the result set of the original query, and finally, for any particularly underrepresented aspect, identifies keywords that would enhance that aspect's representation and automatically expands the query using the best one. The paper presents experiments that show AbraQ significantly increases the precision of hard queries, whereas traditional automatic query expansion techniques have not improved precision. AbraQ also compared favourably against a range of interactive query expansion techniques that require user involvement including clustering, web-log analysis, relevance feedback, and pseudo relevance feedback. Daniel Crabtree, Peter Andreae, Xiaoying Gao |
KDD | 3 |
| 2007 | Automatic Data Record Detection in Web Pages
Xiaoying Gao, Le Phong Bao Vuong, Mengjie Zhang 0001 |
KSEM | 1 |
| 2007 | QC4 - A Clustering Evaluation Method
Daniel Crabtree, Peter Andreae, Xiaoying Gao |
PAKDD | 3 |
| 2007 | Understanding Query Aspects with applications to Interactive Query ExpansionabstractFor many hard queries, users spend a lot of time refining their queries to find relevant documents. Many methods help by suggesting refinements, but it is hard for users to choose the best refinement, as the best refinements are often quite obscure. This paper presents Qasp, an approach that overcomes the limitations of other refinement approaches by using query aspects to find different refinements of ambiguous queries. Qasp clusters the refinements so that descriptive refinements occur together with more obscure and potentially better performing refinements, thereby explaining the effect of refinements to the user. Experiments are presented that show Qasp significantly increases the precision of hard queries. The experiments also show that Qasp's clustering method does find meaningful groups of refinements that help users choose good refinements, which would otherwise be overlooked. Daniel Crabtree, Peter Andreae, Xiaoying Gao |
Web Intelligence | 3 |
| 2007 | A New Crossover Operator in Genetic Programming for Object ClassificationabstractThe crossover operator has been considered "the centre of the storm" in genetic programming (GP). However, many existing GP approaches to object recognition suggest that the standard GP crossover is not sufficiently powerful in producing good child programs due to the totally random choice of the crossover points. To deal with this problem, this paper introduces an approach with a new crossover operator in GP for object recognition, particularly object classification. In this approach, a local hill-climbing search is used in constructing good building blocks, a weight called looseness is introduced to identify the good building blocks in individual programs, and the looseness values are used as heuristics in choosing appropriate crossover points to preserve good building blocks. This approach is examined and compared with the standard crossover operator and the headless chicken crossover (HCC) method on a sequence of object classification problems. The results suggest that this approach outperforms the HCC, the standard crossover, and the standard crossover operator with hill climbing on all of these problems in terms of the classification accuracy. Although this approach spends a bit longer time than the standard crossover operator, it significantly improves the system efficiency over the HCC method. Mengjie Zhang 0001, Xiaoying Gao, Weijun Lou |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2006 | Looseness Controlled Crossover in GP for Object RecognitionabstractThis paper describes an approach to improving the crossover operator in genetic programming for object recognition particularly object classification problems. In this approach, instead of randomly choosing the crossover points as in the standard crossover operator, we use a measure called looseness to guide the selection of crossover points. Rather than using the genetic beam search only, this approach uses a hybrid beam-hill climbing search scheme in the evolutionary process. This approach is examined and compared with the standard crossover operator and the headless chicken crossover method on a sequence of object classification problems. The results suggest that this approach outperforms both the headless chicken crossover and the standard crossover on all of these problems. Mengjie Zhang 0001, Xiaoying Gao, Weijun Lou |
IEEE Congress on Evolutionary Computation | 2 |
| 2006 | Modelling Citation Networks for Improving Scientific Paper Classification Performance
Mengjie Zhang 0001, Xiaoying Gao, Minh Duc Cao, Yuejin Ma |
PRICAI | 2 |
| 2006 | Investigation of Brood Size in GP with Brood Recombination Crossover for Object Recognition
Mengjie Zhang 0001, Xiaoying Gao, Weijun Lou, Dongping Qian |
PRICAI | 2 |
| 2006 | Query Directed Web Page ClusteringabstractWeb page clustering methods categorize and organize search results into semantically meaningful clusters that assist users with search refinement; but finding clusters that are semantically meaningful to users is difficult. In this paper, we describe a new Web page clustering algorithm, QDC, which uses the user's query as part of a reliable measure of cluster quality. The new algorithm has five key innovations: a new query directed cluster quality guide that uses the relationship between clusters and the query, an improved cluster merging method that generates semantically coherent clusters by using cluster description similarity in additional to cluster overlap, a new cluster splitting method that fixes the cluster chaining or cluster drifting problem, an improved heuristic for cluster selection that uses the query directed cluster quality guide, and a new method of improving clusters by ranking the pages by relevance to the cluster. We evaluate QDC by comparing its clustering performance against that of four other algorithms on eight data sets (four use full text data and four use snippet data) by using eleven different external evaluation measurements. We also evaluate QDC by informally analysing its real world usability and performance through comparison with six other algorithms on four data sets. QDC provides a substantial performance improvement over other Web page clustering algorithms Daniel Crabtree, Peter Andreae, Xiaoying Gao |
Web Intelligence | 3 |
| 2006 | Data Extraction from Semi-structured Web Pages by ClusteringabstractThis paper introduces an approach to the use of clustering for data extraction from semi-structured Web pages. A variant Hierarchical Agglomerative Clustering (HAC) algorithm K-neighbours-HAC is developed which uses the similarities of the data format (HTML tags) and the data content (text string values) to group similar text tokens into clusters. Using these clusters, similar text tokens are identi- fied as data fields and extracted as target information. The approach is examined and compared with a number of existing information extraction systems on two different sets of web pages and the results suggest that the new approach is effective for web information extraction and that it outperforms all of the existing approaches on these web sites. Le Phong Bao Vuong, Xiaoying Gao, Mengjie Zhang 0001 |
Web Intelligence | 2 |
| 2005 | Improving Web Clustering by Cluster SelectionabstractWeb page clustering is a technology that puts semantically related Web pages into groups and is useful for categorizing, organizing, and refining search results. When clustering using only textual information, suffix tree clustering (STC) outperforms other clustering algorithms by making use of phrases and allowing clusters to overlap. One problem of STC and other similar algorithms is how to select a small set of clusters to display to the user from a very large set of generated clusters. The cluster selection method used in STC is flawed in that it does not handle overlapping clusters appropriately. This paper introduces a new cluster scoring function and a new cluster selection algorithm to overcome the problems with overlapping clusters, which are combined with STC to make a new clustering algorithm ESTC. This paper's experiments show that ESTC significantly outperforms STC and that even with less data ESTC performs similarly to a commercial clustering search engine. Daniel Crabtree, Xiaoying Gao, Peter Andreae |
Web Intelligence | 2 |
| 2005 | Standardized Evaluation Method for Web Clustering ResultsabstractWeb clustering assists users of a search engine by presenting search results as clusters of related pages. Many clustering algorithms with different characteristics have been developed: but the lack of a standardized Web clustering evaluation method that can evaluate clusterings with different characteristics has prevented effective comparison of algorithms. The paper solves this by introducing a new structure for defining general ideal clusterings and new measurements for evaluating clusterings with different characteristics by comparing them against the general ideal clustering. Daniel Crabtree, Xiaoying Gao, Peter Andreae |
Web Intelligence | 2 |
| 2004 | Approximately Repetitive Structure Detection for Wrapper Induction
Xiaoying Gao, Peter Andreae, Richard Collins |
PRICAI | 1 |
| 2004 | Automatic Pattern Construction For Web Information ExtractionabstractThis paper describes a domain independent approach for automatically constructing information extraction patterns for semi-structured web pages. Given a randomly chosen page from a web site of similarly structured pages, the system identifies a region of the page that has a regular "tabular" structure, and then infers an extraction pattern that will match the "rows" of the region and identify the data elements. The approach was tested on three corpora containing a series of tabular web sites from different domains and achieved a success rate of at least 80%. A significant strength of the system is that it can infer extraction patterns from a single training page and does not require any manual labeling of the training page. Xiaoying Gao, Mengjie Zhang 0001, Peter Andreae |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 1 |
| 2003 | Learning Information Extraction Patterns from Tabular Web Pages without Manual LabellingabstractWe describe a domain independent approach to automatically constructing information extraction patterns for semistructured Web pages. The approach was tested on three corpora containing a series of tabular Web sites from different domains and achieved a success rate of at least 80%. A significant strength of the system is that it can infer extraction patterns from a single training page and does not require any manual labeling of the training page. Xiaoying Gao, Mengjie Zhang 0001, Peter Andreae |
Web Intelligence | 1 |