Tzu-Tsung Wong

dblp:67/123 · DBLP profile ↗
← Back
21ranked-venue papers
18as first author
4since 2021 · last 2026
0000-0001-8132-0214ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 10 first-author · 1 since 2021Databases, data management, data science and information retrieval · 8 · 8 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Multi-interval monotonic data transformation for privacy-preserving data mining in rule-based classification
abstract
Several methods have been proposed for privacy-preserving data mining, and classification is one of the major tasks in data mining. Previous privacy-preserving methods attempted to improve classification accuracy, while model utilization is also a critical issue in classification tasks. An individual can be identified from the combination of attributes, and this study considers both the achievement of privacy preserving and the utilization of classification models to design a novel method for transforming continuous or numeric attributes. To ensure the monotonic property of continuous attributes, multi-interval monotonic functions with randomly chosen parameters are employed for data transformation. Original and transformed data were analyzed by decision tree learning and RIPPER algorithm. By comparing with a perturbation method and a transformation method proposed by previous studies, the experimental results on ten data sets showed that only our method can achieve both privacy preserving and model utilization. The privacy of individuals can be enhanced by increasing the number of intervals for monotonic transformation.
Tzu-Tsung Wong, Jialin Su
Knowl. Inf. Syst.1
2025 A consistency analysis on four evaluation metrics for classifying imbalanced data
Tzu-Tsung Wong, Pei-Chen Chung
Knowl. Inf. Syst.1
2022 Linear Approximation of F-Measure for the Performance Evaluation of Classification Algorithms on Imbalanced Data Sets
abstract
Accuracy is a popular measure for evaluating the performance of classification algorithms tested on ordinary data sets. When a data set is imbalanced,F-measure will be a better choice than accuracy for this purpose. SinceF-measure is calculated as the harmonic mean of recall and precision, it is difficult to find the sampling distribution ofF-measure for evaluating classification algorithms. Since the values of recall and precision are dependent, their joint distribution is assumed to follow a bivariate normal distribution in this study. When the evaluation method isk-fold cross validation, a linear approximation approach is proposed to derive the sampling distribution ofF-measure. This approach is used to design methods for comparing the performance of two classification algorithms tested on single or multiple imbalanced data sets. The methods are tested on ten imbalanced data sets to demonstrate their effectiveness. The weight of recall provided by this linear approximation approach can help us to analyze the characteristics of classification algorithms.
Tzu-Tsung Wong
IEEE Trans. Knowl. Data Eng.1
2021 Multinomial naïve Bayesian classifier with generalized Dirichlet priors for high-dimensional imbalanced data
Tzu-Tsung Wong, Hsing-Chen Tsai
Knowl. Based Syst.1
2020 Hybrid classification algorithms based on instance filtering
Tzu-Tsung Wong, Nai-Yu Yang, Guo-Hong Chen
Inf. Sci.1
2020 Reliable Accuracy Estimates from k-Fold Cross Validation
abstract
It is popular to evaluate the performance of classification algorithms by k-fold cross validation. A reliable accuracy estimate will have a relatively small variance, and several studies therefore suggested to repeatedly perform k-fold cross validation. Most of them did not consider the correlation among the replications of k-fold cross validation, and hence the variance could be underestimated. The purpose of this study is to explore whether k-fold cross validation should be repeatedly performed for obtaining reliable accuracy estimates. The dependency relationships between the predictions of the same instance in two replications of k-fold cross validation are first analyzed for k-nearest neighbors with k = 1. Then, statistical methods are proposed to test the strength of the dependency level between the accuracy estimates resulting from two replications of k-fold cross validation. The experimental results on 20 data sets show that the accuracy estimates obtained from various replications of k-fold cross validation are generally highly correlated, and the correlation will be higher as the number of folds increases. The k-fold cross validation with a large number of folds and a small number of replications should be adopted for performance evaluation of classification algorithms.
Tzu-Tsung Wong, Po-Yang Yeh
IEEE Trans. Knowl. Data Eng.1
2017 Parametric methods for comparing the performance of two classification algorithms evaluated by k-fold cross validation on multiple data sets
Tzu-Tsung Wong
Pattern Recognit.1
2017 Dependency Analysis of Accuracy Estimates in k-Fold Cross Validation
abstract
A standard procedure for evaluating the performance of classification algorithms is k-fold cross validation. Since the training sets for any pair of iterations in k-fold cross validation are overlapping when the number of folds is larger than two, the resulting accuracy estimates are considered to be dependent. In this paper, the overlapping of training sets is shown to be irrelevant in determining whether two fold accuracies are dependent or not. Then a statistical method is proposed to test the appropriateness of assuming independence for the accuracy estimates in k-fold cross validation. This method is applied on 20 data sets, and the experimental results suggest that it is generally appropriate to assume that the fold accuracies are independent. The cross validation of non-overlapping training sets can make fold accuracies to be dependent. However, this dependence almost has no impact on estimating the sample variance of fold accuracies, and hence they can generally be assumed to be independent.
Tzu-Tsung Wong, Nai-Yu Yang
IEEE Trans. Knowl. Data Eng.1
2016 An efficient parameter estimation method for generalized Dirichlet priors in naïve Bayesian classifiers with multinomial models
Tzu-Tsung Wong, Chao-Rui Liu
Pattern Recognit.1
2015 Performance evaluation of classification algorithms by k-fold and leave-one-out cross validation
Tzu-Tsung Wong
Pattern Recognit.1
2014 Generalized Dirichlet priors for Naïve Bayesian classifiers with multinomial models in document classification
Tzu-Tsung Wong
Data Min. Knowl. Discov.1
2013 Naïve Bayesian Classifiers with Multinomial Models for rRNA Taxonomic Assignment
abstract
The introduction of next-generation sequencing in ecological studies has created a major revolution in microbial and fungal ecology. Direct sequencing of hypervariable regions from ribosomal RNA genes can provide rapid and inexpensive analysis for ecological communities. To get deep understanding from these rRNA fragments, the Ribosomal Database Project developed the "RDP Classifier” utilizing 8-mer nucleotide frequencies with Bayesian theorem to obtain taxonomy affiliation. The classifier is computationally efficient and works well with massive short sequences. However, the binary model employed in the RDP classifier does not consider the repetitive 8-mers in each reference sequence. Previous studies have pointed out that multinomial model usually results a better performance than binary model. In this study, we present the naïve Bayesian classifiers with multinomial models that take repetitive 8-mers into account for classifying microbial 16S and fungal 28S rRNA sequences. The results obtained from the multinomial approach were compared with those obtained from the binomial RDP classifier by 250-bp, 400-bp, 800-bp, and full-length reads to demonstrate that the multinomial approach can generally achieve a higher predictive accuracy in most hypervariable regions.
Kuan-Liang Liu, Tzu-Tsung Wong
IEEE ACM Trans. Comput. Biol. Bioinform.2
2012 A hybrid discretization method for naïve Bayesian classifiers
Tzu-Tsung Wong
Pattern Recognit.1
2011 A gene selection method for microarray data based on risk genes
Tzu-Tsung Wong, Ding-Qun Chen
Expert Syst. Appl.1
2011 Individual attribute prior setting methods for naïve Bayesian classifiers
Tzu-Tsung Wong, Liang-Hao Chang
Pattern Recognit.1
2010 A Probabilistic mechanism based on clustering analysis and distance measure for subset gene selection
Tzu-Tsung Wong, Kuan-Liang Liu
Expert Syst. Appl.1
2009 Alternative prior assumptions for improving the performance of naïve Bayesian classifiers
Tzu-Tsung Wong
Data Min. Knowl. Discov.1
2008 Two-stage classification methods for microarray data
Tzu-Tsung Wong, Ching-Han Hsu
Expert Syst. Appl.1
2005 Mining negative contrast sets from data with discrete attributes
Tzu-Tsung Wong, Kuo-Lung Tseng
Expert Syst. Appl.1
2003 Implications of the Dirichlet Assumption for Discretization of Continuous Variables in Naive Bayesian Classifiers
Chun-Nan Hsu, Hung-Ju Huang, Tzu-Tsung Wong
Mach. Learn.3
2000 Why Discretization Works for Naive Bayesian Classifiers
Chun-Nan Hsu, Hung-Ju Huang, Tzu-Tsung Wong
ICML3