VLDB 2026 Research / reviewers in the wild / expert
Wenhao Shu
dblp:52/5115
· DBLP profile ↗
49ranked-venue papers
26as first author
23since 2021 · last 2027
0000-0003-2422-6760ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 24 first-author · 21 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | A two-stage multi-objective enhanced NSGA-II method for multi-label feature selection
Wenhao Shu, Linjun Zhu, Wenbin Qian |
Expert Syst. Appl. | 1 |
| 2026 | Semi-supervised outlier detection for partially labeled numerical data using generalized multigranulation fuzzy neighborhood rough set
Wenhao Shu, Yueming Jiang, Wenbin Qian |
Appl. Intell. | 1 |
| 2026 | Correction to: Semi-supervised outlier detection for partially labeled numerical data using generalized multigranulation fuzzy neighborhood rough set
Wenhao Shu, Yueming Jiang, Wenbin Qian |
Appl. Intell. | 1 |
| 2026 | An accelerator and feature selection using fuzzy information granularity to partially labeled data
Zhenchao Yan, Songlin He, Jianhui Yu, Wenhao Shu, Chase Qishi Wu |
Appl. Intell. | 4 |
| 2026 | LIMFS: Label interaction-aware multi-label feature selection
Wenhao Shu, Duoqian Miao 0001 |
Expert Syst. Appl. | 4 |
| 2026 | Granular Ball-Guided Disambiguation for Partial Multilabel Feature Selection via Maximum Consistency Minimum UncertaintyabstractPartial multilabel feature selection (PMLFS) is a prevalent subject that aims to enhance the performance of multilabel learning (MLL) in the context of noisy labels. In PMLFS, a crucial aspect is handling the false positive labels hidden in the candidate label set, as the imprecise annotations could mislead the feature selection process. However, many existing approaches for partial label disambiguation rely on topology information and tend to be error-prone. Besides, feature selection frameworks are often built upon a linear regression model, leading to a reliance on the classifier and a deficiency in exploring local structures. Focusing on the issues above, this article proposes a novel two-stage PMLFS method, resorting to the ideology of granular computing. In the first stage, a label disambiguation method is developed using label-specific information. Specifically, a specific granular ball computing model is designed to characterize the distribution of datapoints labeled differently, and therefore, using the affinity relationships among samples and balls, the label-specific information concealed in the data distribution can be captured for label disambiguation. In the second stage, a filter-based feature selection method that explores the local structure of samples is presented. This method relies on a devised fuzzy decision neighborhood rough set (FDNRS) to capture more detailed membership information by maximizing the neighborhood consistency of samples' related labels. Simultaneously, the feature selection method minimizes the uncertainty derived from unrelated labels. Extensive experiments on 12 datasets in terms of four evaluation metrics demonstrated the effectiveness of the proposed approach. Fankang Xu, Wenbin Qian, Wenhao Shu, Weiping Ding 0001, Shuyin Xia |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Leveraging hyper-interval granules labeling and local mixed neighborhood entropy for semi-supervised feature selection
Wenhao Shu, Guojing Liao, Wenbin Qian |
Neurocomputing | 1 |
| 2025 | Partial Multilabel Learning Using Noise-Tolerant Broad Learning System With Label Enhancement and Dimensionality ReductionabstractPartial multilabel learning (PML) addresses the issue of noisy supervision, which contains an overcomplete set of candidate labels for each instance with only a valid subset of training data. Using label enhancement techniques, researchers have computed the probability of a label being ground truth. However, enhancing labels in the noisy label space makes it impossible for the existing partial multilabel label enhancement methods to achieve satisfactory results. Besides, few methods simultaneously involve the ambiguity problem, the feature space's redundancy, and the model's efficiency in PML. To address these issues, this article presents a novel joint partial multilabel framework using broad learning systems (namely BLS-PML) with three innovative mechanisms: 1) a trustworthy label space is reconstructed through a novel label enhancement method to avoid the bias caused by noisy labels; 2) a low-dimensional feature space is obtained by a confidence-based dimensionality reduction method to reduce the effect of redundancy in the feature space; and 3) a noise-tolerant BLS is proposed by adding a dimensionality reduction layer and a trustworthy label layer to deal with PML problem. We evaluated it on six real-world and seven synthetic datasets, using eight state-of-the-art partial multilabel algorithms as baselines and six evaluation metrics. Out of 144 experimental scenarios, our method significantly outperforms the baselines by about 80%, demonstrating its robustness and effectiveness in handling partial multilabel tasks. Wenbin Qian, Yanqiang Tu, Wenhao Shu, Yiu-Ming Cheung |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Label Disambiguation-Based Feature Selection for Partial Multi-label Learning
Fankang Xu, Wenbin Qian, Xingxing Cai, Wenhao Shu, Yiu-Ming Cheung, Weiping Ding 0001 |
ICPR (7) | 4 |
| 2024 | Semi-supervised feature selection based on discernibility matrix and mutual information
Wenbin Qian, Lijuan Wan, Wenhao Shu |
Appl. Intell. | 3 |
| 2024 | Multi-label feature selection for missing labels by granular-ball based mutual information
Wenhao Shu, Yichen Hu, Wenbin Qian |
Appl. Intell. | 1 |
| 2024 | Label distribution feature selection based on label-specific features
Wenhao Shu, Qiang Xia 0005, Wenbin Qian |
Appl. Intell. | 1 |
| 2024 | Neighborhood multigranulation rough sets for cost-sensitive feature selection on hybrid data
Wenhao Shu, Qiang Xia 0005, Wenbin Qian |
Neurocomputing | 1 |
| 2024 | Neighborhood relation-based incremental label propagation algorithm for partially labeled hybrid data
Wenhao Shu, Dongtao Cao, Wenbin Qian |
Mach. Learn. | 1 |
| 2023 | Neighbourhood discernibility degree-based semisupervised feature selection for partially labelled mixed-type data with granular ball
Wenhao Shu, Jianhui Yu, Wenbin Qian |
Appl. Intell. | 1 |
| 2023 | Information gain-based semi-supervised feature selection for hybrid data
Wenhao Shu, Zhenchao Yan, Jianhui Yu, Wenbin Qian |
Appl. Intell. | 1 |
| 2023 | Semi-supervised feature selection for partially labeled mixed-type data based on multi-criteria measure approach
Wenhao Shu, Jianhui Yu, Zhenchao Yan, Wenbin Qian |
Int. J. Approx. Reason. | 1 |
| 2023 | Multi-label feature selection based on rough granular-ball and label distribution
Wenbin Qian, Fankang Xu, Wenhao Shu, Weiping Ding 0001 |
Inf. Sci. | 4 |
| 2023 | Partial multi-label learning via three-way decision-based tri-training
Wenbin Qian, Yanqiang Tu, Wenhao Shu |
Knowl. Based Syst. | 4 |
| 2022 | Incremental neighborhood entropy-based feature selection for mixed-type data under the variation of feature set
Wenhao Shu, Wenbin Qian, Yonghong Xie |
Appl. Intell. | 1 |
| 2022 | Information granularity-based incremental feature selection for partially labeled hybrid dataabstractFeature selection can reduce the dimensionality of data effectively. Most of the existing feature selection approaches using rough sets focus on the static single type data. However, in many real-world applications, data sets are the hybrid data including symbolic, numerical and missing features. Meanwhile, an object set in the hybrid data often changes dynamically with time. For the hybrid data, since acquiring all the decision labels of them is expensive and time-consuming, only small portion of the decision labels for the hybrid data is obtained. Therefore, in this paper, incremental feature selection algorithms based on information granularity are developed for dynamic partially labeled hybrid data with the variation of an object set. At first, the information granularity is given to measure the feature significance for partially labeled hybrid data. Then, incremental mechanisms of information granularity are proposed with the variation of an object set. On this basis, incremental feature selection algorithms with the variation of a single object and group of objects are proposed, respectively. Finally, extensive experimental results on different UCI data sets demonstrate that compared with the non-incremental feature selection algorithms, incremental feature selection algorithms can select a subset of features in shorter time without losing the classification accuracy, especially when the group of objects changes dynamically, the group incremental feature selection algorithm is more efficient. Wenhao Shu, Zhenchao Yan, Jianhui Yu, Wenbin Qian |
Intell. Data Anal. | 1 |
| 2022 | Feature selection for label distribution learning via feature similarity and label correlation
Wenbin Qian, Yinsong Xiong, Wenhao Shu |
Inf. Sci. | 4 |
| 2021 | Cost-sensitive feature selection on multi-label data via neighborhood granularity and label enhancement
Xuandong Long, Wenbin Qian, Wenhao Shu |
Appl. Intell. | 4 |
| 2020 | Mutual information-based label distribution feature selection for multi-label learning
Wenbin Qian, Wenhao Shu |
Knowl. Based Syst. | 4 |
| 2020 | Incremental feature selection for dynamic hybrid data using neighborhood rough set
Wenhao Shu, Wenbin Qian, Yonghong Xie |
Knowl. Based Syst. | 1 |
| 2019 | An Efficient Uncertainty Measure-based Attribute Reduction Approach for Interval-valued Data with Missing ValuesabstractAttribute reduction plays an important role in knowledge discovery and data mining. Confronted with data characterized by the interval and missing values in many data analysis tasks, it is interesting to research the attribute reduction for interval-valued data with missing values. Uncertainty measures can supply efficient viewpoints, which help us to disclose the substantive characteristics of such data. Therefore, this paper addresses the attribute reduction problem based on uncertainty measure for interval-valued data with missing values. At first, an uncertainty measure is provided for measuring candidate attributes, and then an efficient attribute reduction algorithm is developed for the interval-valued data with missing values. To improve the efficiency of attribute reduction, the objects that fall within the positive region are deleted from the whole object set in the process of selecting attributes. Finally, experimental results demonstrate that the proposed algorithm can find a subset of attributes in much shorter time than existing attribute reduction algorithms without losing the classification performance. Wenhao Shu, Wenbin Qian, Yonghong Xie, Zhaoping Tang |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 1 |
| 2019 | Incremental approaches for feature selection from dynamic data with the variation of multiple objects
Wenhao Shu, Wenbin Qian, Yonghong Xie |
Knowl. Based Syst. | 1 |
| 2017 | Feature Selection based on Discernibility Function in Incomplete Data with Fuzzy DecisionabstractRough set theory has been applied successfully in knowledge discovery, computational intelligence and decision analysis. It can only deal with features of symbolic type for complete data. However, incomplete data with fuzzy decision under a preference-ordered relation is common in real-world applications. In this paper, we propose a feature selection framework for such data by combining the dominance-based rough sets. At first, the judgment theorems are established by the Boolean reasoning techniques. Then, the discernibility matrix and the discernibility function approach are proposed to find all subsets of features. In addition, an efficient feature selection algorithm to find a feature subset is proposed. Finally, the experimental results show that, in most cases for different data sets, the proposed algorithm is effective and efficient for feature selection from the incomplete data with fuzzy decision. Wenbin Qian, Wenhao Shu |
ICTAI | 2 |
| 2016 | Improving matrix approximation for recommendation via a clustering-based reconstructive method
Ke Ji, Runyuan Sun, Wenhao Shu |
Neurocomputing | 4 |
| 2016 | Multi-criteria feature selection on cost-sensitive data with missing values
Wenhao Shu, Hong Shen 0001 |
Pattern Recognit. | 1 |
| 2015 | Cost-Sensitive Feature Selection on Heterogeneous Data
Wenbin Qian, Wenhao Shu |
PAKDD (2) | 2 |
| 2015 | An incremental approach to attribute reduction from dynamic incomplete decision systems in rough set theory
Wenhao Shu, Wenbin Qian |
Data Knowl. Eng. | 1 |
| 2015 | Mutual information criterion for feature selection from incomplete data
Wenbin Qian, Wenhao Shu |
Neurocomputing | 2 |
| 2015 | Next-song recommendation with temporal dynamics
Ke Ji, Runyuan Sun, Wenhao Shu |
Knowl. Based Syst. | 3 |
| 2014 | A Consistency-Based Dimensionality Reduction Algorithm in Incomplete Data
Wenbin Qian, Wenhao Shu |
APWeb | 2 |
| 2014 | Mutual Information-Based Feature Selection from Set-Valued DataabstractIn many machine learning and data mining applications, it may happen that the data acquired for classification analysis are set-valued, i.e., The feature values of an object set are set-valued, which can be used to characterize uncertain information in decision making tasks. Set-valued data are the generalized models of single-valued data. Some mutual information-based feature selection algorithms have been extensively studied, but less effort has been made to investigate the feature selection issue with the mutual information analysis in set-valued data. Just owing to these, mutual information is firstly introduced in the set-valued data in this paper. Unlike the traditional computations, the mutual information is estimated on the unmarked objects. Correspondingly, a feature selection algorithm based on mutual information is developed, which is implemented in a dwindling universe to quicken the feature selection process. Compared with the state-of-the-art methods, the experimental results on different data sets demonstrate the efficiency and effectiveness of the proposed algorithm in set-valued data. Wenhao Shu, Wenbin Qian |
ICTAI | 1 |
| 2014 | A New Evaluation Function for Entropy-Based Feature Selection from Incomplete Data
Wenhao Shu, Hong Shen 0001, Yingpeng Sang, Yidong Li, Jun Wu 0007 |
PAKDD (2) | 1 |
| 2014 | Updating attribute reduction in incomplete decision systems with the variation of attribute set
Wenhao Shu, Hong Shen 0001 |
Int. J. Approx. Reason. | 1 |
| 2014 | A fast approach to attribute reduction from perspective of attribute measures in incomplete decision systems
Wenhao Shu, Wenbin Qian |
Knowl. Based Syst. | 1 |
| 2014 | Incremental feature selection based on rough set in dynamic incomplete data
Wenhao Shu, Hong Shen 0001 |
Pattern Recognit. | 1 |
| 2013 | A rough-set based incremental approach for updating attribute reduction under dynamic incomplete decision systemsabstractEfficient attribute reduction in large-scale incomplete decision systems is a challenging problem. The computation of tolerance classes induced by the condition attributes in the incomplete decision system is a key part among all existing attribute reduction algorithms. Moreover, updating attribute reduction for dynamically-increasing decision systems has attracted much attention, in view of that incremental attribute reduction algorithms in a dynamic incomplete decision system have not yet been sufficiently discussed so far. In this paper, we first introduce a simpler way of computing tolerance classes than the classical method. Then we present an incremental attribute reduction algorithm to compute an attribute reduct for a dynamically-increasing incomplete decision system. Compared with the non-incremental algorithms, our incremental attribute reduction algorithm can compute a new attribute reduct in much shorter time. Experiments on four data sets downloaded from UCI show that the feasibility and effectiveness of the proposed incremental algorithm. Wenhao Shu, Hong Shen 0001 |
FUZZ-IEEE | 1 |
| 2010 | A method for fuzzy risk analysis based on the new similarity of trapezoidal fuzzy numbers
Zhangyan Xu, Shichao Zhang 0001, Wenbin Qian, Wenhao Shu |
Expert Syst. Appl. | 4 |
| 2005 | A multiple classifier approach to detect Chinese character recognition errors
Kei Yuen Hung, Robert Wing Pong Luk, Daniel S. Yeung, Korris Fu-Lai Chung, Wenhao Shu |
Pattern Recognit. | 5 |
| 2004 | Feature Selection For Chinese Character Recognition Based On Inductive LearningabstractFeature selection is a difficult but important issue in the field of machine learning and pattern recognition. In this paper, features for Chinese character recognition are selected by using inductive learning algorithms. The existing inductive learning method based on extension matrix requires precise consistency between positive example and negative example sets, which is very difficult to maintain in most practical cases. The traditional decision tree algorithm ID3 considers only the performance of the discriminating power while selecting features. However, in actual practice the consideration of the associated cost of feature extraction may become a significant concern. In addressing these problems we propose a modified extension matrix approach to select feature subset from the training example set with noises. A decision tree algorithm based on information gain and cost evaluation is also proposed to facilitate cost consideration. The comparative experiments show that the proposed algorithms perform better than the existing inductive learning algorithms to a certain extent. Guoliang Qian, Daniel S. Yeung, Eric C. C. Tsang, Wenhao Shu |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2003 | An Approach To Natural Stroke Extraction For Off-Line Loosely-Constrained Handwritten Chinese CharactersabstractThis paper proposes a new approach to extracting natural strokes from the skeletons of loosely-constrained, off-line handwritten Chinese characters. It admits the output substrokes from a previously proposed fuzzy substroke extractor as its inputs. By identifying a number of expected ambiguities which include mutual similarities, unstable touches and joint/cross distortions, fuzzy stroke models are constructed and a "hit-all" fuzzy stroke matching strategy is pursued. Fuzzy partitioning technique is used to generate a ranked list of consistent stroke sets from the set of fuzzy strokes being identified. With this approach, a maximum of 20 distinct natural stroke classes can be extracted from each input character, together with an estimate on the actual count of strokes which compose the character. Our system offers a number of performance tuning capabilities such as the computation of the fuzzy scores of each extracted stroke, the adjustment on the fuzzy stroke model parameters, and the potential of incorporating one's personal writing styles into our methodology. Daniel S. Yeung, Hank-Shun Fong, Eric C. C. Tsang, Wenhao Shu, Xiaolong Wang 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2000 | Detection of Language (Model) ErrorsabstractThe bigram language models are popular, in much language processing applications, in both Indo-European and Asian languages. However, when the language model for Chinese is applied in a novel domain, the accuracy is reduced significantly, from 96% to 78% in our evaluation. We apply pattern recognition techniques (i.e. Bayesian, decision tree and neural network classifiers) to discover language model errors. We have examined 2 general types of features: model-based and language-specific features. In our evaluation, Bayesian classifiers produce the best recall performance of 80% but the precision is low (60%). Neural network produced good recall (75%) and precision (80%) but both Bayesian and Neural network have low skip ratio (65%). The decision tree classifier produced the best precision (81%) and skip ratio (76%) but its recall is the lowest (73%). Kei Yuen Hung, Robert Wing Pong Luk, Daniel S. Yeung, Korris Fu-Lai Chung, Wenhao Shu |
EMNLP | 5 |
| 2000 | An extension matrix approach to Chinese character recognitionabstractOptical character recognition (OCR) provides a solution to acquire, archive and retrieve a large amount of paper-based information which is still commonly used in our daily life. The process of a classical optical character recognition system consists of a series of stages, such as format analysis, text segmentation, feature extraction and classification. This paper focuses on the last two stages, and two contributions can be claimed: first, rapid transformed stroke density features (SDF) are used for preliminary classification and outline primitive structural features for final classification. Second, the original extension matrix algorithm is improved by heuristic path searching on the basis of information entropy as well as Laplace error rate evaluation function. Our experimental results prove that the rapid transformed SDFs are insensitive to image translation or rotation, and that the improved extension matrix algorithm outperforms other inductive approaches based on AE1 and AQ15. The excellent performance with respect to a large data set also indicates our proposed approach is effective and efficient. Wenhao Shu, Daming Shi 0001, Guoliang Qian, Fusi Wang |
SMC | 1 |
| 1998 | Feature selection for handwritten Chinese character recognition based on genetic algorithmsabstractFeature selection is of great importance in recognition system design because it directly affects the overall performance of the recognition system. Feature selection can be considered as a problem of global combinatorial optimization. It is a very time-consuming task to search the most suitable features amongst a huge number of possible feature combinations, therefore, an effective and efficient search technique is desired. In this paper, we use genetic algorithms (GA) to design a feature selection approach for handwritten Chinese character recognition. Four contributions are claimed: First, the general transformed divergence among classes, which is derived from Mahalanobis distances, is proposed to be the fitness function in the feature selection based on GA; Second, a special crossover operator other than traditional one is given; Third, a special criterion of terminating selections is inferred from the criterion of minimum error probability in a Bayes classifier; Fourth, we compare our method with the feature selection based on branch-and-bound algorithm (BAB), which is often used to reduce the calculation of feature selection via exhaustive search. The analyses of the experimental results can be proceeded that traditional GA is an ergodic Markov chain, while, BAB is a depth first heuristic algorithm for exhaustive search. We conclude that the GA-based method proposed in this paper is promising to solve the feature selection problems in a multidimensional space. Daming Shi 0001, Wenhao Shu |
SMC | 2 |
| 1988 | An accurate method for recognition of printed Chinese charactersabstractAn accurate method for recognition of printed Chinese characters is proposed. An improved peripheral method and the Walsh transform of both horizontal and vertical projection functions of a Chinese character are used for feature extraction. Study shows that this method has the advantages of fast processing speed, accurate recognition rate, and strong resistance to noise.> Wenhao Shu, Guo-wei Cui, Ri-hua Zhao |
ICPR | 1 |