Huan Zhang 0007

dblp:23/1797-7 · DBLP profile ↗
← Back
14ranked-venue papers
10as first author
13since 2021 · last 2026
0000-0002-9914-6602ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 A general dual-view framework for instance weighted naive Bayes
Huan Zhang 0007, Kexin Meng, Pei Lv, Shuo He 0002, Mingliang Xu 0001
Pattern Recognit.1
2025 Adaptive Non-disjoint Discretization for Tree Augmented Naive Bayes
Pei Lv, Huan Zhang 0007
ADMA (4)5
2025 KA2Net: Kolmogorov-Arnold Attention Network for PCB Tiny Defect Detection
Qi Wang 0137, Zhihong Xie, Yongping Zeng, Huan Zhang 0007
ICIC (8)4
2025 Instance Correlation Graph-based Naive Bayes
abstract
Due to its simplicity, effectiveness and robustness, naive Bayes (NB) has continued to be one of the top 10 data mining algorithms. To improve its performance, a large number of improved algorithms have been proposed in the last few decades. However, in addition to Gaussian naive Bayes (GNB), there is little work on numerical attributes. At the same time, none of them takes into account the correlations among instances. To fill this gap, we propose a novel algorithm called instance correlation graph-based naive Bayes (ICGNB). Specifically, it first uses original attributes to construct an instance correlation graph (ICG) to represent the correlations among instances. Then, it employs a variational graph auto-encoder (VGAE) to generate new attributes from the constructed ICG and uses them to augment original attributes. Finally, it weights each augmented attribute to alleviate the attribute redundancy and builds GNB on the weighted attributes. The experimental results on tens of datasets show that ICGNB significantly outperforms its deserved competitors.Our codes and datasets are available at https://github.com/jiangliangxiao/ICGNB.
Liangxiao Jiang, Wenjun Zhang 0012, Liangjun Yu, Huan Zhang 0007
ICML5
2025 EMAWNB: Enhanced Multi-view Attribute Weighted Naive Bayes
Guanzhi Liu, Kexin Meng, Pei Lv, Huan Zhang 0007
PAKDD (3)4
2025 Dual-View Learning from Crowds
abstract
Crowdsourcing services provide a fast and cheap way to obtain substantial labeled data by employing crowd workers on the Internet. In crowdsourcing learning, two-stage methods have been widely used, which first infer the integrated label for each instance and then build a learning model using instances with their integrated labels. However, existing two-stage methods mainly focus on how to infer more accurate integrated labels, after that, most of them directly regard the integrated labels as class labels to build a learning model, which loses the detailed worker labeling information in multiple noisy labels and thus results in sub-optimal model accuracy. To solve this problem, in this study, we take the multiple noisy labels of each instance as its attribute value vector to construct another view in addition to the original attribute view, and propose a novel two-stage method called dual-view learning from crowds (DVLFC). In DVLFC, we first pick out workers with sufficient number of labels and augment the multiple noisy label set for each instance, then we build a supervised learning model in each view and at last we fuse their class-membership probabilities to get the final classification result. Extensive experiments on both real-world and artificial crowdsourced datasets prove the effectiveness of DVLFC.
Huan Zhang 0007, Liangxiao Jiang, Wenjun Zhang 0012, Geoffrey I. Webb
ACM Trans. Knowl. Discov. Data1
2023 Rigorous non-disjoint discretization for naive Bayes
abstract
Naive Bayes is a classical machine learning algorithm for which discretization is commonly used to transform quantitative attributes into qualitative attributes. Of numerous discretization methods, Non-Disjoint Discretization (NDD) proposes a novel perspective by forming overlapping intervals and always locating a value toward the middle of an interval. However, existing approaches to NDD fail to adequately consider the effect of multiple occurrences of a single value — a commonly occurring circumstance in practice. By necessity, all occurrences of a single value fall within the same interval. As a result, it is often not possible to discretize an attribute into intervals containing equal numbers of training instances. Current methods address this issue in an ad hoc manner, reducing the specificity of the resulting atomic intervals. In this study, we propose a non-disjoint discretization method for NB, called Rigorous Non-Disjoint Discretization (RNDD), that handles multiple occurrences of a single value in a systematic manner. Our extensive experimental results suggest that RNDD significantly outperforms NDD along with all other existing state-of-the-art competitors.
Huan Zhang 0007, Liangxiao Jiang, Geoffrey I. Webb
Pattern Recognit.1
2023 Multi-View Attribute Weighted Naive Bayes
abstract
Naive Bayes (NB) continues to be one of the top 10 data mining algorithms due to its simplicity, efficiency and efficacy. Numerous enhancements have been proposed to weaken its attribute conditional independence assumption. However, all of them only focus on the raw attribute view, which is hard to reflect all the data characteristics in real-world applications. To portray data characteristics more comprehensively, in this study, we construct two label views from the raw attributes and propose a novel model called multi-view attribute weighted naive Bayes (MAWNB). In MAWNB, we first build multiple super-parent one-dependence estimators (SPODEs) as well as random trees (RTs), then we utilize each of them to classify each training instance in turn and use all their predicted class labels to construct two label views. Next, to avoid attribute redundancy, we optimize the weight of each attribute value for each class by minimizing the negative conditional log-likelihood (CLL) in each view. Finally, the estimated class-membership probabilities by three views are fused to predict the class label for each test instance. Extensive experiments show that MAWNB significantly outperforms NB and all the other existing state-of-the-art competitors.
Huan Zhang 0007, Liangxiao Jiang, Wenjun Zhang 0012, Chaoqun Li 0001
IEEE Trans. Knowl. Data Eng.1
2022 Attribute augmented and weighted naive Bayes
Huan Zhang 0007, Liangxiao Jiang, Chaoqun Li 0001
Sci. China Inf. Sci.1
2022 Fine tuning attribute weighted naive Bayes
Huan Zhang 0007, Liangxiao Jiang
Neurocomputing1
2021 CS-ResNet: Cost-sensitive residual convolutional neural network for PCB cosmetic defect detection
abstract
In the printed circuit board (PCB) industry, cosmetic defect detection is an essential process to ensure product quality. However, existing PCB cosmetic defect detection approaches have a high false alarm rate, which lead to expensive labor costs of manual confirmation. To solve this problem, some traditional machine learning-based approaches have been proposed, but they just utilize hand-crafted features to build classifiers and thus are rough and sub-optimal. Recently, due to its powerful capability in automatic feature extraction, convolutional neural network (CNN) has been widely used in PCB cosmetic defect detection. However, few of them pay attention to the imbalanced class distribution as well as the different misclassification costs of real and pseudo defects, both of which are common problems in the PCB industry. To this end, in this study, we propose a novel model called cost-sensitive residual convolutional neural network (CS-ResNet) by adding a cost-sensitive adjustment layer in the standard ResNet. Specifically, we assign larger weights to minority real defects based on the class-imbalance degree and then optimize CS-ResNet by minimizing the weighted cross-entropy loss function. We conducted a series of experiments by comparing CS-ResNet with the standard ResNet, state-of-the-art CNN-based approach Auto-VRS and traditional machine learning-based approach HOG-SVM on a real-world PCB cosmetic defect dataset. Experimental results show that CS-ResNet achieves the highest Sensitivity (0.89), G-mean (0.91) and the lowest misclassification costs.
Huan Zhang 0007, Liangxiao Jiang, Chaoqun Li 0001
Expert Syst. Appl.1
2021 Collaboratively weighted naive Bayes
Huan Zhang 0007, Liangxiao Jiang, Chaoqun Li 0001
Knowl. Inf. Syst.1
2021 Attribute and instance weighted naive Bayes
Huan Zhang 0007, Liangxiao Jiang, Liangjun Yu
Pattern Recognit.1
2020 Class-specific attribute value weighting for Naive Bayes
Huan Zhang 0007, Liangxiao Jiang, Liangjun Yu
Inf. Sci.1