EDBT 2026 Demo / reviewers in the wild / expert
Xiehua Yu
dblp:136/7204
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2024
0000-0001-5896-1654ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Online streaming feature selection based on hierarchical structure informationabstractSummary Hierarchical classification learning aims to exploit the hierarchical relationship between data categories. The high dimensionality and dynamic of the data feature space are the main challenges of this research. Hierarchical feature selection uses a hierarchical structure to divide large‐scale tasks into multiple small tasks, which can more effectively improve the training speed and prediction accuracy of classification models. To present, existing online feature selection methods ignore the hierarchical structure of data. In addition, the dependency relationships in the hierarchical structure can serve as auxiliary knowledge to aid feature selection. Based on this, this paper proposes an online streaming hierarchical feature selection method based on kernelized fuzzy rough sets (OFS‐HNFRS). First, we use the prior knowledge of the hierarchical structure to divide the sample set into multiple subsets. Second, the dependency relationship of hierarchical structure is extended to kernelized fuzzy rough sets, and hierarchical category dependency based on kernelized fuzzy rough sets is defined. Finally, a new online feature selection framework is proposed, which is used to evaluate the relevance, significance, and redundancy of features. We verify the effectiveness of the proposed algorithm on six hierarchical datasets and eight flat datasets. Shuxian Lin, Chenxi Wang 0002, Xiehua Yu, Huirong Fang, Yaojin Lin |
Concurr. Comput. Pract. Exp. | 3 |
| 2023 | Label distribution learning with high-order label correlationsabstractSummary Label distribution learning (LDL) is an emerging learning paradigm, which can be used to solve the label ambiguity problem. In spite of the recent great progress in LDL algorithms considering label correlations, the majority of existing methods only measure pairwise label correlations through the commonly used similarity metric, which is incapable of accurately reflecting the complex relationship between labels. To solve this problem, a novel label distribution learning method—based on high‐order label correlations (LDL‐HLC) is proposed. By virtue of the ‐regularization sparse reconstruction of the label space, the high‐order label correlations matrix is firstly obtained. Then, a new regular term can be constructed to fit the final prediction label distribution via the correction matrix. Furthermore, efficient classification performance and complete feature selection are guaranteed by common features learning via ‐regularization. Finally, the performance and effectiveness of the proposed algorithm are well illustrated through extensive experiments on 14 label distribution datasets and comparisons with some existing algorithms. Yulin Li 0002, Yaojin Lin, Xiehua Yu, Lei Guo 0020, Shaozi Li |
Concurr. Comput. Pract. Exp. | 3 |
| 2023 | Multi-label feature selection based on relative entropy and fuzzy neighborhood mutual discrimination indexabstractAbstract Multi‐label feature selection eliminates irrelevant and redundant features, and then improves the performance of multi‐label classification models. Most multi‐label feature selection algorithms assume that the training set contains logical labels, which means that labels are equally important for instances. However, in practical applications, there are different importances with respect to labels. To solve the problem, a multi‐label feature selection method based on relative entropy and fuzzy neighborhood mutual discriminant index is proposed. Firstly, logical labels are converted to label distribution through label enhancement. Secondly, the neighborhood and relative entropy are introduced into the label distribution, the label neighborhood similarity matrix is constructed to describe the similarity of samples under label space. Finally, the fuzzy neighborhood mutual discrimination index is used to combine the candidate features with the label neighborhood similarity matrix, which is used to judge the distinguishing ability of the candidate features. Comprehensive experiment of eight multi‐label datasets shows that the proposed algorithm has better classification performance than other compared algorithms. Chenxi Wang 0002, Chen E, Mengli Ren, Lei Guo 0020, Xiehua Yu, Yaojin Lin, Shaozi Li |
Concurr. Comput. Pract. Exp. | 5 |
| 2023 | Online feature selection for hierarchical classification learning based on improved ReliefFabstractAbstract In hierarchical classification learning, the feature space of data has high dimensionality and is unknown with emergent features. To solve the above problems, we propose an online hierarchical feature selection algorithm based on adaptive ReliefF. Firstly, ReliefF is adaptively improved via using the density information of instances around the target sample, making it unnecessary to prespecify parameters. Secondly, the hierarchical relationship between classes is used, and a new method for calculating the feature weight of hierarchical data is defined. Then, an online correlation analysis method based on feature interaction is designed. Finally, the adaptive ReliefF algorithm is improved based on feature redundancy, and the feature weight is scaled by the correlation between features in order to achieve the dynamic updating of feature redundancy. A large number of experiments verify the effectiveness of the proposed algorithm. Chenxi Wang 0002, Mengli Ren, Chen E, Lei Guo 0020, Xiehua Yu, Yaojin Lin, Shaozi Li |
Concurr. Comput. Pract. Exp. | 5 |
| 2022 | Multilabel causal variable discovery in multisourceabstractAbstract Multilabel causal feature selection, as a well‐known and effective approach in dealing with high‐dimensional multilabel data, is a popular topic. Amount of causal feature selection algorithms have achieved a great deal of success in classification and prediction tasks. However, the descriptive information of data is collected from different data sources in many practical applications. While few researches focus on the causal variable discovery in multisource environments due to the complex causal relationships. To address these problems, we propose a causal feature selection framework in multisource environments to solve the above problems. Firstly, we mine the causal mechanism with respect to the class attribute under the assumption that only a single data source is included. Secondly, by utilizing the concept of causal invariance in causal inference, we formulate the problem of causal feature selection with multiple data sources as a search problem for an invariant set across data sources. In addition, we give the upper and lower bounds of the causal invariant set. Finally, we design a novel multisource multilabel causal feature selection (MMCFS) algorithm. To verify the effectiveness of the proposed algorithm, we compare it with 12 feature selection methods on synthetic datasets. Experiment results show that the classification performance of MMCFS achieves highly competitive performance against other comparing algorithms. Yun-an Wang, Yaojin Lin, Xiehua Yu, Zhisen Wei, Shaozi Li |
Concurr. Comput. Pract. Exp. | 3 |
| 2022 | Neighborhood rough set based multi-label feature selection with label correlationabstractSummary Neighborhood rough set (NRS) is considered as an effective tool for feature selection and has been widely used in processing high‐dimensional data. However, most of the existing methods are difficult to deal with multi‐label data and are lack of considering label correlation (LC), which is an important issue in multi‐label learning. Therefore, in this article, we introduce a new NRS model with considering LC. First, we explore LC by calculating the similarity relation between labels and divide the related labels into several label subsets. Then, a new neighborhood relation is proposed, which can solve the problem of neighborhood granularity selection by using the nearest neighbor information distribution of instances under the related labels. On this basis, the NRS model is reconstructed by embedding LC information, and the related properties of the model are discussed. Moreover, we design a new feature significance function to evaluate the quality of features, which can well capture the specific relationship between features and labels. Finally, a greedy forward feature selection algorithm is designed. Extensive experiments which are conducted on different types of datasets verify the effectiveness of the proposed algorithm. Yilin Wu 0001, Xiehua Yu, Yaojin Lin, Shaozi Li |
Concurr. Comput. Pract. Exp. | 3 |