Wenhao Shu

dblp:52/5115 · DBLP profile ↗
← Back
49ranked-venue papers
26as first author
23since 2021 · last 2027
0000-0003-2422-6760ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 42 · 24 first-author · 21 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author
YearPublicationVenuePosition
2027 A two-stage multi-objective enhanced NSGA-II method for multi-label feature selection
Wenhao Shu, Linjun Zhu, Wenbin Qian
Expert Syst. Appl.1
2026 Semi-supervised outlier detection for partially labeled numerical data using generalized multigranulation fuzzy neighborhood rough set
Wenhao Shu, Yueming Jiang, Wenbin Qian
Appl. Intell.1
2026 Correction to: Semi-supervised outlier detection for partially labeled numerical data using generalized multigranulation fuzzy neighborhood rough set
Wenhao Shu, Yueming Jiang, Wenbin Qian
Appl. Intell.1
2026 An accelerator and feature selection using fuzzy information granularity to partially labeled data
Zhenchao Yan, Songlin He, Jianhui Yu, Wenhao Shu, Chase Qishi Wu
Appl. Intell.4
2026 LIMFS: Label interaction-aware multi-label feature selection
Wenhao Shu, Duoqian Miao 0001
Expert Syst. Appl.4
2026 Granular Ball-Guided Disambiguation for Partial Multilabel Feature Selection via Maximum Consistency Minimum Uncertainty
abstract
Partial multilabel feature selection (PMLFS) is a prevalent subject that aims to enhance the performance of multilabel learning (MLL) in the context of noisy labels. In PMLFS, a crucial aspect is handling the false positive labels hidden in the candidate label set, as the imprecise annotations could mislead the feature selection process. However, many existing approaches for partial label disambiguation rely on topology information and tend to be error-prone. Besides, feature selection frameworks are often built upon a linear regression model, leading to a reliance on the classifier and a deficiency in exploring local structures. Focusing on the issues above, this article proposes a novel two-stage PMLFS method, resorting to the ideology of granular computing. In the first stage, a label disambiguation method is developed using label-specific information. Specifically, a specific granular ball computing model is designed to characterize the distribution of datapoints labeled differently, and therefore, using the affinity relationships among samples and balls, the label-specific information concealed in the data distribution can be captured for label disambiguation. In the second stage, a filter-based feature selection method that explores the local structure of samples is presented. This method relies on a devised fuzzy decision neighborhood rough set (FDNRS) to capture more detailed membership information by maximizing the neighborhood consistency of samples' related labels. Simultaneously, the feature selection method minimizes the uncertainty derived from unrelated labels. Extensive experiments on 12 datasets in terms of four evaluation metrics demonstrated the effectiveness of the proposed approach.
Fankang Xu, Wenbin Qian, Wenhao Shu, Weiping Ding 0001, Shuyin Xia
IEEE Trans. Neural Networks Learn. Syst.3
2025 Leveraging hyper-interval granules labeling and local mixed neighborhood entropy for semi-supervised feature selection
Wenhao Shu, Guojing Liao, Wenbin Qian
Neurocomputing1
2025 Partial Multilabel Learning Using Noise-Tolerant Broad Learning System With Label Enhancement and Dimensionality Reduction
abstract
Partial multilabel learning (PML) addresses the issue of noisy supervision, which contains an overcomplete set of candidate labels for each instance with only a valid subset of training data. Using label enhancement techniques, researchers have computed the probability of a label being ground truth. However, enhancing labels in the noisy label space makes it impossible for the existing partial multilabel label enhancement methods to achieve satisfactory results. Besides, few methods simultaneously involve the ambiguity problem, the feature space's redundancy, and the model's efficiency in PML. To address these issues, this article presents a novel joint partial multilabel framework using broad learning systems (namely BLS-PML) with three innovative mechanisms: 1) a trustworthy label space is reconstructed through a novel label enhancement method to avoid the bias caused by noisy labels; 2) a low-dimensional feature space is obtained by a confidence-based dimensionality reduction method to reduce the effect of redundancy in the feature space; and 3) a noise-tolerant BLS is proposed by adding a dimensionality reduction layer and a trustworthy label layer to deal with PML problem. We evaluated it on six real-world and seven synthetic datasets, using eight state-of-the-art partial multilabel algorithms as baselines and six evaluation metrics. Out of 144 experimental scenarios, our method significantly outperforms the baselines by about 80%, demonstrating its robustness and effectiveness in handling partial multilabel tasks.
Wenbin Qian, Yanqiang Tu, Wenhao Shu, Yiu-Ming Cheung
IEEE Trans. Neural Networks Learn. Syst.4
2024 Label Disambiguation-Based Feature Selection for Partial Multi-label Learning
Fankang Xu, Wenbin Qian, Xingxing Cai, Wenhao Shu, Yiu-Ming Cheung, Weiping Ding 0001
ICPR (7)4
2024 Semi-supervised feature selection based on discernibility matrix and mutual information
Wenbin Qian, Lijuan Wan, Wenhao Shu
Appl. Intell.3
2024 Multi-label feature selection for missing labels by granular-ball based mutual information
Wenhao Shu, Yichen Hu, Wenbin Qian
Appl. Intell.1
2024 Label distribution feature selection based on label-specific features
Wenhao Shu, Qiang Xia 0005, Wenbin Qian
Appl. Intell.1
2024 Neighborhood multigranulation rough sets for cost-sensitive feature selection on hybrid data
Wenhao Shu, Qiang Xia 0005, Wenbin Qian
Neurocomputing1
2024 Neighborhood relation-based incremental label propagation algorithm for partially labeled hybrid data
Wenhao Shu, Dongtao Cao, Wenbin Qian
Mach. Learn.1
2023 Neighbourhood discernibility degree-based semisupervised feature selection for partially labelled mixed-type data with granular ball
Wenhao Shu, Jianhui Yu, Wenbin Qian
Appl. Intell.1
2023 Information gain-based semi-supervised feature selection for hybrid data
Wenhao Shu, Zhenchao Yan, Jianhui Yu, Wenbin Qian
Appl. Intell.1
2023 Semi-supervised feature selection for partially labeled mixed-type data based on multi-criteria measure approach
Wenhao Shu, Jianhui Yu, Zhenchao Yan, Wenbin Qian
Int. J. Approx. Reason.1
2023 Multi-label feature selection based on rough granular-ball and label distribution
Wenbin Qian, Fankang Xu, Wenhao Shu, Weiping Ding 0001
Inf. Sci.4
2023 Partial multi-label learning via three-way decision-based tri-training
Wenbin Qian, Yanqiang Tu, Wenhao Shu
Knowl. Based Syst.4
2022 Incremental neighborhood entropy-based feature selection for mixed-type data under the variation of feature set
Wenhao Shu, Wenbin Qian, Yonghong Xie
Appl. Intell.1
2022 Information granularity-based incremental feature selection for partially labeled hybrid data
abstract
Feature selection can reduce the dimensionality of data effectively. Most of the existing feature selection approaches using rough sets focus on the static single type data. However, in many real-world applications, data sets are the hybrid data including symbolic, numerical and missing features. Meanwhile, an object set in the hybrid data often changes dynamically with time. For the hybrid data, since acquiring all the decision labels of them is expensive and time-consuming, only small portion of the decision labels for the hybrid data is obtained. Therefore, in this paper, incremental feature selection algorithms based on information granularity are developed for dynamic partially labeled hybrid data with the variation of an object set. At first, the information granularity is given to measure the feature significance for partially labeled hybrid data. Then, incremental mechanisms of information granularity are proposed with the variation of an object set. On this basis, incremental feature selection algorithms with the variation of a single object and group of objects are proposed, respectively. Finally, extensive experimental results on different UCI data sets demonstrate that compared with the non-incremental feature selection algorithms, incremental feature selection algorithms can select a subset of features in shorter time without losing the classification accuracy, especially when the group of objects changes dynamically, the group incremental feature selection algorithm is more efficient.
Wenhao Shu, Zhenchao Yan, Jianhui Yu, Wenbin Qian
Intell. Data Anal.1
2022 Feature selection for label distribution learning via feature similarity and label correlation
Wenbin Qian, Yinsong Xiong, Wenhao Shu
Inf. Sci.4
2021 Cost-sensitive feature selection on multi-label data via neighborhood granularity and label enhancement
Xuandong Long, Wenbin Qian, Wenhao Shu
Appl. Intell.4
2020 Mutual information-based label distribution feature selection for multi-label learning
Wenbin Qian, Wenhao Shu
Knowl. Based Syst.4
2020 Incremental feature selection for dynamic hybrid data using neighborhood rough set
Wenhao Shu, Wenbin Qian, Yonghong Xie
Knowl. Based Syst.1
2019 An Efficient Uncertainty Measure-based Attribute Reduction Approach for Interval-valued Data with Missing Values
abstract
Attribute reduction plays an important role in knowledge discovery and data mining. Confronted with data characterized by the interval and missing values in many data analysis tasks, it is interesting to research the attribute reduction for interval-valued data with missing values. Uncertainty measures can supply efficient viewpoints, which help us to disclose the substantive characteristics of such data. Therefore, this paper addresses the attribute reduction problem based on uncertainty measure for interval-valued data with missing values. At first, an uncertainty measure is provided for measuring candidate attributes, and then an efficient attribute reduction algorithm is developed for the interval-valued data with missing values. To improve the efficiency of attribute reduction, the objects that fall within the positive region are deleted from the whole object set in the process of selecting attributes. Finally, experimental results demonstrate that the proposed algorithm can find a subset of attributes in much shorter time than existing attribute reduction algorithms without losing the classification performance.
Wenhao Shu, Wenbin Qian, Yonghong Xie, Zhaoping Tang
Int. J. Uncertain. Fuzziness Knowl. Based Syst.1
2019 Incremental approaches for feature selection from dynamic data with the variation of multiple objects
Wenhao Shu, Wenbin Qian, Yonghong Xie
Knowl. Based Syst.1
2017 Feature Selection based on Discernibility Function in Incomplete Data with Fuzzy Decision
abstract
Rough set theory has been applied successfully in knowledge discovery, computational intelligence and decision analysis. It can only deal with features of symbolic type for complete data. However, incomplete data with fuzzy decision under a preference-ordered relation is common in real-world applications. In this paper, we propose a feature selection framework for such data by combining the dominance-based rough sets. At first, the judgment theorems are established by the Boolean reasoning techniques. Then, the discernibility matrix and the discernibility function approach are proposed to find all subsets of features. In addition, an efficient feature selection algorithm to find a feature subset is proposed. Finally, the experimental results show that, in most cases for different data sets, the proposed algorithm is effective and efficient for feature selection from the incomplete data with fuzzy decision.
Wenbin Qian, Wenhao Shu
ICTAI2
2016 Improving matrix approximation for recommendation via a clustering-based reconstructive method
Ke Ji, Runyuan Sun, Wenhao Shu
Neurocomputing4
2016 Multi-criteria feature selection on cost-sensitive data with missing values
Wenhao Shu, Hong Shen 0001
Pattern Recognit.1
2015 Cost-Sensitive Feature Selection on Heterogeneous Data
Wenbin Qian, Wenhao Shu
PAKDD (2)2
2015 An incremental approach to attribute reduction from dynamic incomplete decision systems in rough set theory
Wenhao Shu, Wenbin Qian
Data Knowl. Eng.1
2015 Mutual information criterion for feature selection from incomplete data
Wenbin Qian, Wenhao Shu
Neurocomputing2
2015 Next-song recommendation with temporal dynamics
Ke Ji, Runyuan Sun, Wenhao Shu
Knowl. Based Syst.3
2014 A Consistency-Based Dimensionality Reduction Algorithm in Incomplete Data
Wenbin Qian, Wenhao Shu
APWeb2
2014 Mutual Information-Based Feature Selection from Set-Valued Data
abstract
In many machine learning and data mining applications, it may happen that the data acquired for classification analysis are set-valued, i.e., The feature values of an object set are set-valued, which can be used to characterize uncertain information in decision making tasks. Set-valued data are the generalized models of single-valued data. Some mutual information-based feature selection algorithms have been extensively studied, but less effort has been made to investigate the feature selection issue with the mutual information analysis in set-valued data. Just owing to these, mutual information is firstly introduced in the set-valued data in this paper. Unlike the traditional computations, the mutual information is estimated on the unmarked objects. Correspondingly, a feature selection algorithm based on mutual information is developed, which is implemented in a dwindling universe to quicken the feature selection process. Compared with the state-of-the-art methods, the experimental results on different data sets demonstrate the efficiency and effectiveness of the proposed algorithm in set-valued data.
Wenhao Shu, Wenbin Qian
ICTAI1
2014 A New Evaluation Function for Entropy-Based Feature Selection from Incomplete Data
Wenhao Shu, Hong Shen 0001, Yingpeng Sang, Yidong Li, Jun Wu 0007
PAKDD (2)1
2014 Updating attribute reduction in incomplete decision systems with the variation of attribute set
Wenhao Shu, Hong Shen 0001
Int. J. Approx. Reason.1
2014 A fast approach to attribute reduction from perspective of attribute measures in incomplete decision systems
Wenhao Shu, Wenbin Qian
Knowl. Based Syst.1
2014 Incremental feature selection based on rough set in dynamic incomplete data
Wenhao Shu, Hong Shen 0001
Pattern Recognit.1
2013 A rough-set based incremental approach for updating attribute reduction under dynamic incomplete decision systems
abstract
Efficient attribute reduction in large-scale incomplete decision systems is a challenging problem. The computation of tolerance classes induced by the condition attributes in the incomplete decision system is a key part among all existing attribute reduction algorithms. Moreover, updating attribute reduction for dynamically-increasing decision systems has attracted much attention, in view of that incremental attribute reduction algorithms in a dynamic incomplete decision system have not yet been sufficiently discussed so far. In this paper, we first introduce a simpler way of computing tolerance classes than the classical method. Then we present an incremental attribute reduction algorithm to compute an attribute reduct for a dynamically-increasing incomplete decision system. Compared with the non-incremental algorithms, our incremental attribute reduction algorithm can compute a new attribute reduct in much shorter time. Experiments on four data sets downloaded from UCI show that the feasibility and effectiveness of the proposed incremental algorithm.
Wenhao Shu, Hong Shen 0001
FUZZ-IEEE1
2010 A method for fuzzy risk analysis based on the new similarity of trapezoidal fuzzy numbers
Zhangyan Xu, Shichao Zhang 0001, Wenbin Qian, Wenhao Shu
Expert Syst. Appl.4
2005 A multiple classifier approach to detect Chinese character recognition errors
Kei Yuen Hung, Robert Wing Pong Luk, Daniel S. Yeung, Korris Fu-Lai Chung, Wenhao Shu
Pattern Recognit.5
2004 Feature Selection For Chinese Character Recognition Based On Inductive Learning
abstract
Feature selection is a difficult but important issue in the field of machine learning and pattern recognition. In this paper, features for Chinese character recognition are selected by using inductive learning algorithms. The existing inductive learning method based on extension matrix requires precise consistency between positive example and negative example sets, which is very difficult to maintain in most practical cases. The traditional decision tree algorithm ID3 considers only the performance of the discriminating power while selecting features. However, in actual practice the consideration of the associated cost of feature extraction may become a significant concern. In addressing these problems we propose a modified extension matrix approach to select feature subset from the training example set with noises. A decision tree algorithm based on information gain and cost evaluation is also proposed to facilitate cost consideration. The comparative experiments show that the proposed algorithms perform better than the existing inductive learning algorithms to a certain extent.
Guoliang Qian, Daniel S. Yeung, Eric C. C. Tsang, Wenhao Shu
Int. J. Pattern Recognit. Artif. Intell.4
2003 An Approach To Natural Stroke Extraction For Off-Line Loosely-Constrained Handwritten Chinese Characters
abstract
This paper proposes a new approach to extracting natural strokes from the skeletons of loosely-constrained, off-line handwritten Chinese characters. It admits the output substrokes from a previously proposed fuzzy substroke extractor as its inputs. By identifying a number of expected ambiguities which include mutual similarities, unstable touches and joint/cross distortions, fuzzy stroke models are constructed and a "hit-all" fuzzy stroke matching strategy is pursued. Fuzzy partitioning technique is used to generate a ranked list of consistent stroke sets from the set of fuzzy strokes being identified. With this approach, a maximum of 20 distinct natural stroke classes can be extracted from each input character, together with an estimate on the actual count of strokes which compose the character. Our system offers a number of performance tuning capabilities such as the computation of the fuzzy scores of each extracted stroke, the adjustment on the fuzzy stroke model parameters, and the potential of incorporating one's personal writing styles into our methodology.
Daniel S. Yeung, Hank-Shun Fong, Eric C. C. Tsang, Wenhao Shu, Xiaolong Wang 0001
Int. J. Pattern Recognit. Artif. Intell.4
2000 Detection of Language (Model) Errors
abstract
The bigram language models are popular, in much language processing applications, in both Indo-European and Asian languages. However, when the language model for Chinese is applied in a novel domain, the accuracy is reduced significantly, from 96% to 78% in our evaluation. We apply pattern recognition techniques (i.e. Bayesian, decision tree and neural network classifiers) to discover language model errors. We have examined 2 general types of features: model-based and language-specific features. In our evaluation, Bayesian classifiers produce the best recall performance of 80% but the precision is low (60%). Neural network produced good recall (75%) and precision (80%) but both Bayesian and Neural network have low skip ratio (65%). The decision tree classifier produced the best precision (81%) and skip ratio (76%) but its recall is the lowest (73%).
Kei Yuen Hung, Robert Wing Pong Luk, Daniel S. Yeung, Korris Fu-Lai Chung, Wenhao Shu
EMNLP5
2000 An extension matrix approach to Chinese character recognition
abstract
Optical character recognition (OCR) provides a solution to acquire, archive and retrieve a large amount of paper-based information which is still commonly used in our daily life. The process of a classical optical character recognition system consists of a series of stages, such as format analysis, text segmentation, feature extraction and classification. This paper focuses on the last two stages, and two contributions can be claimed: first, rapid transformed stroke density features (SDF) are used for preliminary classification and outline primitive structural features for final classification. Second, the original extension matrix algorithm is improved by heuristic path searching on the basis of information entropy as well as Laplace error rate evaluation function. Our experimental results prove that the rapid transformed SDFs are insensitive to image translation or rotation, and that the improved extension matrix algorithm outperforms other inductive approaches based on AE1 and AQ15. The excellent performance with respect to a large data set also indicates our proposed approach is effective and efficient.
Wenhao Shu, Daming Shi 0001, Guoliang Qian, Fusi Wang
SMC1
1998 Feature selection for handwritten Chinese character recognition based on genetic algorithms
abstract
Feature selection is of great importance in recognition system design because it directly affects the overall performance of the recognition system. Feature selection can be considered as a problem of global combinatorial optimization. It is a very time-consuming task to search the most suitable features amongst a huge number of possible feature combinations, therefore, an effective and efficient search technique is desired. In this paper, we use genetic algorithms (GA) to design a feature selection approach for handwritten Chinese character recognition. Four contributions are claimed: First, the general transformed divergence among classes, which is derived from Mahalanobis distances, is proposed to be the fitness function in the feature selection based on GA; Second, a special crossover operator other than traditional one is given; Third, a special criterion of terminating selections is inferred from the criterion of minimum error probability in a Bayes classifier; Fourth, we compare our method with the feature selection based on branch-and-bound algorithm (BAB), which is often used to reduce the calculation of feature selection via exhaustive search. The analyses of the experimental results can be proceeded that traditional GA is an ergodic Markov chain, while, BAB is a depth first heuristic algorithm for exhaustive search. We conclude that the GA-based method proposed in this paper is promising to solve the feature selection problems in a multidimensional space.
Daming Shi 0001, Wenhao Shu
SMC2
1988 An accurate method for recognition of printed Chinese characters
abstract
An accurate method for recognition of printed Chinese characters is proposed. An improved peripheral method and the Walsh transform of both horizontal and vertical projection functions of a Chinese character are used for feature extraction. Study shows that this method has the advantages of fast processing speed, accurate recognition rate, and strong resistance to noise.>
Wenhao Shu, Guo-wei Cui, Ri-hua Zhao
ICPR1