Guangtao Wang

dblp:26/11029 · DBLP profile ↗
← Back
35ranked-venue papers
11as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 6 first-author · 9 since 2021Databases, data management, data science and information retrieval · 10 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2
YearPublicationVenuePosition
2026 IFHAGrec: Instruction-Finetuned Heterogeneous-Aware Graph Neural Network for Temporally Weighted Recommendation Model
Hongrui Wang 0004, Shanchuan Yu, Xiaoyan Zhu 0003, Guangtao Wang, Jiayin Wang 0002, Jiaxuan Li 0001, Jindong Jiang
KSEM (1)4
2024 Graph-enhanced and collaborative attention networks for session-based recommendation
Xiaoyan Zhu 0003, Yu Zhang 0203, Jiayin Wang 0002, Guangtao Wang
Knowl. Based Syst.4
2023 Dynamic ensemble learning for multi-label classification
Xiaoyan Zhu 0003, Jiaxuan Li 0001, Jingtao Ren, Jiayin Wang 0002, Guangtao Wang
Inf. Sci.5
2023 EfficientFace: an efficient deep network with feature enhancement for accurate face detection
Guangtao Wang, Jun Li 0033, Zhijian Wu, Jifeng Shen, Wankou Yang
Multim. Syst.1
2023 Meta-HGT: Metapath-aware HyperGraph Transformer for heterogeneous information network embedding
Lingyun Song, Guangtao Wang, Xuequn Shang 0001
Neural Networks3
2023 Exploring Interactive and Contrastive Relations for Nested Named Entity Recognition
abstract
Nested named entities (nested NEs) refer to the situation where one named entity is included or nested within another named entity, which cannot be recognized by the traditional sequence labeling methods. Recently, span-based methods have become the mainstream methods for nested Named Entity Recognition (nested NER). The fundamental concept behind this method is to enumerate nearly all potential spans as entity mentions and subsequently classify them. However, span-based methods independently classify spans without considering the semantic relations among them, which negatively impacts the span representation. To address the issue, we propose a novel deep learning architecture for nested NER that explores interactive and contrastive relations among spans. Specifically, we design a scale transformation mechanism that embeds geometric information into span representations, which enhances the model's ability to encode interactive relations between spans. Additionally, we introduce a supervised contrastive learning loss that pulls apart highly overlapping spans in the embedding space to encode the contrastive relations. Experiments show that our method achieves state-of-the-art or competitive performance on three publicly nested NER datasets, thus validating its effectiveness.
Yuefei Wu, Guangtao Wang, Yanping Chen 0010, Wei Wu 0069, Zai Zhang 0002, Bin Shi 0003, Bo Dong 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Improving Time Sensitivity for Question Answering over Temporal Knowledge Graphs
abstract
Question answering over temporal knowledge graphs (KGs) efficiently uses facts contained in a temporal KG, which records entity relations and when they occur in time, to answer natural language questions (e.g., "Who was the president of the US before Obama?").These questions often involve three time-related challenges that previous work fail to adequately address: 1) questions often do not specify exact timestamps of interest (e.g., "Obama" instead of 2000); 2) subtle lexical differences in time relations (e.g., "before" vs "after"); 3) off-the-shelf temporal KG embeddings that previous work builds on ignore the temporal order of timestamps, which is crucial for answering temporal-order related questions.In this paper, we propose a time-sensitive question answering (TSQA) framework to tackle these problems.TSQA features a timestamp estimation module to infer the unwritten timestamp from the question.We also employ a time-sensitive KG encoder to inject ordering information into the temporal KG embeddings that TSQA is based on.With the help of techniques to reduce the search space for potential answers, TSQA significantly outperforms the previous state of the art on a new benchmark for question answering over temporal KGs, especially achieving a 32% (absolute) error reduction on complex questions that require multiple steps of reasoning over facts in the temporal KG.
Guangtao Wang, Peng Qi 0003, Jing Huang 0019
ACL (1)2
2022 A Synopsis Based Approach for Itemset Frequency Estimation over Massive Multi-Transaction Stream
abstract
The streams where multiple transactions are associated with the same key are prevalent in practice, e.g., a customer has multiple shopping records arriving at different time. Itemset frequency estimation on such streams is very challenging since sampling based methods, such as the popularly used reservoir sampling, cannot be used. In this article, we propose a novel k -Minimum Value (KMV) synopsis based method to estimate the frequency of itemsets over multi-transaction streams. First, we extract the KMV synopses for each item from the stream. Then, we propose a novel estimator to estimate the frequency of an itemset over the KMV synopses. Comparing to the existing estimator, our method is not only more accurate and efficient to calculate but also follows the downward-closure property. These properties enable the incorporation of our new estimator with existing frequent itemset mining (FIM) algorithm (e.g., FP-Growth) to mine frequent itemsets over multi-transaction streams. To demonstrate this, we implement a KMV synopsis based FIM algorithm by integrating our estimator into existing FIM algorithms, and we prove it is capable of guaranteeing the accuracy of FIM with a bounded size of KMV synopsis. Experimental results on massive streams show our estimator can significantly improve on the accuracy for both estimating itemset frequency and FIM compared to the existing estimators.
Guangtao Wang, Gao Cong, Ying Zhang 0001, Zhen Hai, Jieping Ye
ACM Trans. Knowl. Discov. Data1
2021 Multi-hop Attention Graph Neural Networks
abstract
Self-attention mechanism in graph neural networks (GNNs) led to state-of-the-art performance on many graph representation learning tasks. Currently, at every layer, attention is computed between connected pairs of nodes and depends solely on the representation of the two nodes. However, such attention mechanism does not account for nodes that are not directly connected but provide important network context. Here we propose Multi-hop Attention Graph Neural Network (MAGNA), a principled way to incorporate multi-hop context information into every layer of attention computation. MAGNA diffuses the attention scores across the network, which increases the receptive field for every layer of the GNN. Unlike previous approaches, MAGNA uses a diffusion prior on attention values, to efficiently account for all paths between the pair of disconnected nodes. We demonstrate in theory and experiments that MAGNA captures large-scale structural information in every layer, and has a low-pass effect that eliminates noisy high-frequency information from graph data. Experimental results on node classification as well as the knowledge graph completion benchmarks show that MAGNA achieves state-of-the-art results: MAGNA achieves up to 5.7% relative error reduction over the previous state-of-the-art on Cora, Citeseer, and Pubmed. MAGNA also obtains the best performance on a large-scale Open Graph Benchmark dataset. On knowledge graph completion MAGNA advances state-of-the-art on WN18RR and FB15k-237 across four different performance metrics.
Guangtao Wang, Rex Ying, Jing Huang 0019, Jure Leskovec
IJCAI1
2021 Inductive Learning on Commonsense Knowledge Graph Completion
abstract
Commonsense knowledge graph (CKG) is a special type of knowledge graph (KG), where entities are composed of free-form text. Existing CKG completion methods focus on transductive learning setting, where all the entities are present during training. Here, we propose the first inductive learning setting for CKG completion, where unseen entities may appear at test time. We emphasize that the inductive learning setting is crucial for CKGs, because unseen entities are frequently introduced due to the fact that CKGs are dynamic and highly sparse. We propose InductivE as the first framework targeted at the inductive CKG completion task. InductivE first ensures the inductive learning capability by directly computing entity embeddings from raw entity attributes. Second, a graph neural network with novel densification process is proposed to further enhance unseen entity representation with neighboring structural information. Experimental results show that InductivE performs especially well on inductive scenarios where it achieves above 48% improvement over previous methods while also outperforms state-of-the-art baselines in transductive settings.
Bin Wang 0040, Guangtao Wang, Jing Huang 0019, Jiaxuan You, Jure Leskovec, C.-C. Jay Kuo
IJCNN2
2021 Graph Ensemble Learning over Multiple Dependency Trees for Aspect-level Sentiment Classification
abstract
Xiaochen Hou, Peng Qi, Guangtao Wang, Rex Ying, Jing Huang, Xiaodong He, Bowen Zhou. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Xiaochen Hou, Peng Qi 0003, Guangtao Wang, Rex Ying, Jing Huang 0019, Xiaodong He 0001, Bowen Zhou 0001
NAACL-HLT3
2021 Ensemble of ML-KNN for classification algorithm recommendation
Xiaoyan Zhu 0003, Chenzhen Ying, Jiayin Wang 0002, Jiaxuan Li 0001, Xin Lai 0003, Guangtao Wang
Knowl. Based Syst.6
2020 Select, Answer and Explain: Interpretable Multi-Hop Reading Comprehension over Multiple Documents
abstract
Interpretable multi-hop reading comprehension (RC) over multiple documents is a challenging problem because it demands reasoning over multiple information sources and explaining the answer prediction by providing supporting evidences. In this paper, we propose an effective and interpretable Select, Answer and Explain (SAE) system to solve the multi-document RC problem. Our system first filters out answer-unrelated documents and thus reduce the amount of distraction information. This is achieved by a document classifier trained with a novel pairwise learning-to-rank loss. The selected answer-related documents are then input to a model to jointly predict the answer and supporting sentences. The model is optimized with a multi-task learning objective on both token level for answer prediction and sentence level for supporting sentences prediction, together with an attention-based interaction between these two tasks. Evaluated on HotpotQA, a challenging multi-hop RC data set, the proposed SAE system achieves top competitive performance in distractor setting compared to other existing systems on the leaderboard.
Kevin Huang 0002, Guangtao Wang, Jing Huang 0019, Xiaodong He 0001, Bowen Zhou 0001
AAAI3
2020 Orthogonal Relation Transforms with Graph Context Modeling for Knowledge Graph Embedding
abstract
Distance-based knowledge graph embeddings have shown substantial improvement on the knowledge graph link prediction task, from TransE to the latest state-of-the-art RotatE.However, complex relations such as N-to-1, 1-to-N and N-to-N still remain challenging to predict.In this work, we propose a novel distance-based approach for knowledge graph link prediction.First we extend the RotatE from 2D complex domain to high dimensional space with orthogonal transforms to model relations.The orthogonal transform embedding for relations keeps the capability for modeling symmetric/anti-symmetric, inverse and compositional relations while achieves better modeling capacity.Second, the graph context is integrated into distance scoring functions directly.Specifically, graph context is explicitly modeled via two directed context representations.Each node embedding in knowledge graph is augmented with two context representations, which are computed from the neighboring outgoing and incoming nodes/edges respectively.The proposed approach improves prediction accuracy on the difficult N-to-1, 1-to-N and N-to-N cases.Our experimental results show that it achieves state-of-the-art results on two common benchmarks FB15k-237 and WNRR-18, especially on FB15k-237 which has many high in-degree nodes.Code available at https://github. com/JD-AI-Research-Silicon-Valley/ KGEmbedding-OTE.
Yun Tang 0002, Jing Huang 0019, Guangtao Wang, Xiaodong He 0001, Bowen Zhou 0001
ACL3
2020 Improving Neural Language Generation with Spectrum Control
Lingxiao Wang 0001, Jing Huang 0019, Kevin Huang 0002, Ziniu Hu, Guangtao Wang, Quanquan Gu
ICLR5
2019 Multi-hop Reading Comprehension across Multiple Documents by Reasoning over Heterogeneous Graphs
abstract
Multi-hop reading comprehension (RC) across documents poses new challenge over single-document RC because it requires reasoning over multiple documents to reach the final answer. In this paper, we propose a new model to tackle the multi-hop RC problem. We introduce a heterogeneous graph with different types of nodes and edges, which is named as Heterogeneous Document-Entity (HDE) graph. The advantage of HDE graph is that it contains different granularity levels of information including candidates, documents and entities in specific document contexts. Our proposed model can do reasoning over the HDE graph with nodes representation initialized with co-attention and self-attention based context encoders. We employ Graph Neural Networks (GNN) based message passing algorithms to accumulate evidences on the proposed HDE graph. Evaluated on the blind test set of the Qangaroo WikiHop data set, our HDE graph based single model delivers competitive result, and the ensemble model achieves the state-of-the-art performance.
Guangtao Wang, Jing Huang 0019, Yun Tang 0002, Xiaodong He 0001, Bowen Zhou 0001
ACL (1)2
2019 A new unsupervised feature selection algorithm using similarity-based feature clustering
abstract
Abstract Unsupervised feature selection is an important problem, especially for high‐dimensional data. However, until now, it has been scarcely studied and the existing algorithms cannot provide satisfying performance. Thus, in this paper, we propose a new unsupervised feature selection algorithm using similarity‐based feature clustering, Feature Selection‐based Feature Clustering (FSFC). FSFC removes redundant features according to the results of feature clustering based on feature similarity. First, it clusters the features according to their similarity. A new feature clustering algorithm is proposed, which overcomes the shortcomings of K‐means. Second, it selects a representative feature from each cluster, which contains most interesting information of features in the cluster. The efficiency and effectiveness of FSFC are tested upon real‐world data sets and compared with two representative unsupervised feature selection algorithms, Feature Selection Using Similarity (FSUS) and Multi‐Cluster‐based Feature Selection (MCFS) in terms of runtime, feature compression ratio, and the clustering results of K‐means. The results show that FSFC can not only reduce the feature space in less time, but also significantly improve the clustering performance of K‐means.
Xiaoyan Zhu 0003, Yu Wang 0069, Yingbin Li, Yonghui Tan, Guangtao Wang, Qinbao Song
Comput. Intell.5
2018 Margin Based PU Learning
abstract
The PU learning problem concerns about learning from positive and unlabeled data. A popular heuristic is to iteratively enlarge training set based on some margin-based criterion. However, little theoretical analysis has been conducted to support the success of these heuristic methods. In this work, we show that not all margin-based heuristic rules are able to improve the learned classifiers iteratively. We find that a so-called large positive margin oracle is necessary to guarantee the success of PU learning. Under this oracle, a provable positive-margin based PU learning algorithm is proposed for linear regression and classification under the truncated Gaussian distributions. The proposed algorithm is able to reduce the recovering error geometrically proportional to the positive margin. Extensive experiments on real-world datasets verify our theory and the state-of-the-art performance of the proposed PU learning algorithm.
Tieliang Gong, Guangtao Wang, Jieping Ye, Zongben Xu
AAAI2
2018 A new classification algorithm recommendation method based on link prediction
Xiaoyan Zhu 0003, Xiaomei Yang, Chenzhen Ying, Guangtao Wang
Knowl. Based Syst.4
2017 Functional Annotation of Human Protein Coding Isoforms via Non-convex Multi-Instance Learning
abstract
Functional annotation of human genes is fundamentally important for understanding the molecular basis of various genetic diseases. A major challenge in determining the functions of human genes lies in the functional diversity of proteins, that is, a gene can perform different functions as it may consist of multiple protein coding isoforms (PCIs). Therefore, differentiating functions of PCIs can significantly deepen our understanding of the functions of genes. However, due to the lack of isoform-level gold-standards (ground-truth annotation), many existing functional annotation approaches are developed at gene-level. In this paper, we propose a novel approach to differentiate the functions of PCIs by integrating sparse simplex projection---that is, a nonconvex sparsity-inducing regularizer---with the framework of multi-instance learning (MIL). Specifically, we label the genes that are annotated to the function under consideration as positive bags and the genes without the function as negative bags. Then, by sparse projections onto simplex, we learn a mapping that embeds the original bag space to a discriminative feature space. Our framework is flexible to incorporate various smooth and non-smooth loss functions such as logistic loss and hinge loss. To solve the resulting highly nontrivial non-convex and non-smooth optimization problem, we further develop an efficient block coordinate descent algorithm. Extensive experiments on human genome data demonstrate that the proposed approaches significantly outperform the state-of-the-art methods in terms of functional annotation accuracy of human PCIs and efficiency.
Tingjin Luo, Shang Qiu, Dongyun Yi, Guangtao Wang, Jieping Ye, Jie Wang 0005
KDD6
2017 Unsupervised Embedding for Latent Similarity by Modeling Heterogeneous MOOC Data
Zhuoxuan Jiang, Shanshan Feng 0001, Weizheng Chen, Guangtao Wang, Xiaoming Li 0001
PAKDD (2)4
2017 Predicting Bugs in Software Code Changes Using Isolation Forest
abstract
Identifying bug immediately when it is introduced can help improve the validity and effectiveness of bug fixing. Predicting bugs in software code changes makes such identification possible. Buggy changes, changes that introduce bugs into source code, can be viewed as anomalies relative to clean changes for that they are rare and irregular. Thus, anomaly detection techniques can be applied to buggy change prediction. Isolation Forest, which detects anomalies based on the hypothesis that the anomalies have the shortest average path length on the constructed random forest, has exhibited its good performance on anomaly detection compared to other anomaly detection methods. In this paper, we adopt it in predicting bugs in software code changes. Empirical study with eight practical open source projects are conducted to validate the effective of Isolation Forest in bug prediction in software code changes. Results of the empirical study show that compared to traditional classification methods used in literature, Isolation Forest can achieve better clean precision, buggy recall, buggy F-measure, AUC and Gmean.
Yueyang He, Guangtao Wang, Heli Sun, Yong Wang 0076
QRS3
2017 LinkLPA: A Link-Based Label Propagation Algorithm for Overlapping Community Detection in Networks
abstract
Community detection is an important methodology for understanding the intrinsic structure and function of complex networks. Because overlapping community is one of the characteristics of real‐world networks and should be considered for community detection, in this article, we propose an algorithm, called link‐based label propagation algorithm (LinkLPA), to detect overlapping communities. Because the link partition is conceptually natural for the problem of overlapping community detection, LinkLPA first transforms node partition problem into link partition problem and employs a new label propagation algorithm with preference on links instead of nodes to detect communities due to the simplicity and efficiency of label propagation algorithm. Then the proposed LinkLPA performs a postprocessing to refine the detected overlapping communities by avoiding over‐overlapping and incorrect partition of weak ties. Experimental results on a large number of real‐world and synthetic networks show that the proposed method achieves high accuracy on detecting overlapping communities in networks.
Heli Sun, Guangtao Wang, Xiaolin Jia, Qinbao Song
Comput. Intell.4
2016 Single Classifier Selection for Ensemble Learning
Guangtao Wang, Xiaomei Yang
ADMA1
2016 A machine learning based software process model recommendation method
Qinbao Song, Xiaoyan Zhu 0003, Guangtao Wang, Heli Sun, Chenhao Xue, Baowen Xu
J. Syst. Softw.3
2016 Automatic Clustering via Outward Statistical Testing on Density Metrics
abstract
Clustering is one of the research hotspots in the field of data mining and has extensive applications in practice. Recently, Rodriguez and Laio [1] published a clustering algorithm on Science that identifies the clustering centers in an intuitive way and clusters objects efficiently and effectively. However, the algorithm is sensitive to a preassigned parameter and suffers from the identification of the “ideal” number of clusters. To overcome these shortages, this paper proposes a new clustering algorithm that can detect the clustering centers automatically via statistical testing. Specifically, the proposed algorithm first defines a new metric to measure the density of an object that is more robust to the preassigned parameter, further generates a metric to evaluate the centrality of each object. Afterwards, it identifies the objects with extremely large centrality metrics as the clustering centers via an outward statistical testing method. Finally, it groups the remaining objects into clusters containing their nearest neighbors with higher density. Extensive experiments are conducted over different kinds of clustering data sets to evaluate the performance of the proposed algorithm and compare with the algorithm in Science. The results show the effectiveness and robustness of the proposed algorithm.
Guangtao Wang, Qinbao Song
IEEE Trans. Knowl. Data Eng.1
2015 An improved data characterization method and its application in classification algorithm recommendation
Guangtao Wang, Qinbao Song
Appl. Intell.1
2015 A dissimilarity-based imbalance data classification algorithm
Qinbao Song, Guangtao Wang, Liang He 0006, Xiaolin Jia
Appl. Intell.3
2014 A Generic Multilabel Learning-Based Classification Algorithm Recommendation Method
abstract
As more and more classification algorithms continue to be developed, recommending appropriate algorithms to a given classification problem is increasingly important. This article first distinguishes the algorithm recommendation methods by two dimensions: (1) meta-features, which are a set of measures used to characterize the learning problems, and (2) meta-target, which represents the relative performance of the classification algorithms on the learning problem. In contrast to the existing algorithm recommendation methods whose meta-target is usually in the form of either the ranking of candidate algorithms or a single algorithm, this article proposes a new and natural multilabel form to describe the meta-target. This is due to the fact that there would be multiple algorithms being appropriate for a given problem in practice. Furthermore, a novel multilabel learning-based generic algorithm recommendation method is proposed, which views the algorithm recommendation as a multilabel learning problem and solves the problem by the mature multilabel learning algorithms. To evaluate the proposed multilabel learning-based recommendation method, extensive experiments with 13 well-known classification algorithms, two kinds of meta-targets such as algorithm ranking and single algorithm, and five different kinds of meta-features are conducted on 1,090 benchmark learning problems. The results show the effectiveness of our proposed multilabel learning-based recommendation method.
Guangtao Wang, Qinbao Song
ACM Trans. Knowl. Discov. Data1
2013 A novel feature subset selection algorithm based on association rule mining
abstract
In this paper, a novel feature selection algorithm FEAST is proposed based on association rule mining. The proposed algorithm first mines association rules from a data set; then, it identifies the relevant and interactive feature values with the constraint association rules whose consequent is the target concept, detects and eliminates the redundant feature values with the constraint association rules whose consequent and antecedent are both of single feature value. Finally, it obtains the feature subset by mapping the feature values to the corresponding features. As the support and confidence thresholds are two important parameters in association rule mining and play a vital role in FEAST, a partial least square regression (PLSR) based threshold prediction method is presented as well. The effectiveness of FEAST is tested on both synthetic and real world data sets, and the classification results of five different types of classifiers with seven representative feature selection algorithms are compared. The results on the synthetic data sets show that FEAST can effectively identify irrelevant and redundant features while reserving interactive ones. The results on the real world data sets show that FEAST outperforms other feature selection algorithms in terms of classification accuracies. In addition, the PLSR based threshold prediction method is performed on the real world data sets, and the results show it works well in recommending proper support and confidence thresholds for FEAST.
Guangtao Wang, Qinbao Song
Intell. Data Anal.1
2013 A Feature Subset Selection Algorithm Automatic Recommendation Method
abstract
Many feature subset selection (FSS) algorithms have been proposed, but not all of them are appropriate for a given feature selection problem. At the same time, so far there is rarely a good way to choose appropriate FSS algorithms for the problem at hand. Thus, FSS algorithm automatic recommendation is very important and practically useful. In this paper, a meta learning based FSS algorithm automatic recommendation method is presented. The proposed method first identifies the data sets that are most similar to the one at hand by the k-nearest neighbor classification algorithm, and the distances among these data sets are calculated based on the commonly-used data set characteristics. Then, it ranks all the candidate FSS algorithms according to their performance on these similar data sets, and chooses the algorithms with best performance as the appropriate ones. The performance of the candidate FSS algorithms is evaluated by a multi-criteria metric that takes into account not only the classification accuracy over the selected features, but also the runtime of feature selection and the number of selected features. The proposed recommendation method is extensively tested on 115 real world data sets with 22 well-known and frequently-used different FSS algorithms for five representative classifiers. The results show the effectiveness of our proposed FSS algorithm recommendation method.
Guangtao Wang, Qinbao Song, Heli Sun, Baowen Xu, Yuming Zhou
J. Artif. Intell. Res.1
2013 Selecting feature subset for high dimensional data via the propositional FOIL rules
Guangtao Wang, Qinbao Song, Baowen Xu, Yuming Zhou
Pattern Recognit.1
2013 A Fast Clustering-Based Feature Subset Selection Algorithm for High-Dimensional Data
abstract
Feature selection involves identifying a subset of the most useful features that produces compatible results as the original entire set of features. A feature selection algorithm may be evaluated from both the efficiency and effectiveness points of view. While the efficiency concerns the time required to find a subset of features, the effectiveness is related to the quality of the subset of features. Based on these criteria, a fast clustering-based feature selection algorithm (FAST) is proposed and experimentally evaluated in this paper. The FAST algorithm works in two steps. In the first step, features are divided into clusters by using graph-theoretic clustering methods. In the second step, the most representative feature that is strongly related to target classes is selected from each cluster to form a subset of features. Features in different clusters are relatively independent, the clustering-based strategy of FAST has a high probability of producing a subset of useful and independent features. To ensure the efficiency of FAST, we adopt the efficient minimum-spanning tree (MST) clustering method. The efficiency and effectiveness of the FAST algorithm are evaluated through an empirical study. Extensive experiments are carried out to compare FAST and several representative feature selection algorithms, namely, FCBF, ReliefF, CFS, Consist, and FOCUS-SF, with respect to four types of well-known classifiers, namely, the probability-based Naive Bayes, the tree-based C4.5, the instance-based IB1, and the rule-based RIPPER before and after feature selection. The results, on 35 publicly available real-world high-dimensional image, microarray, and text data, demonstrate that the FAST not only produces smaller subsets of features but also improves the performances of the four types of classifiers.
Qinbao Song, Jingjie Ni, Guangtao Wang
IEEE Trans. Knowl. Data Eng.3
2012 Selecting Feature Subset via Constraint Association Rules
Guangtao Wang, Qinbao Song
PAKDD (2)1
2012 Automatic recommendation of classification algorithms based on data set characteristics
Qinbao Song, Guangtao Wang
Pattern Recognit.2