Chih-Fong Tsai

dblp:69/6140 · DBLP profile ↗
← Back
63ranked-venue papers
31as first author
13since 2021 · last 2026
0000-0002-5991-2253ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 39 · 19 first-author · 10 since 2021Databases, data management, data science and information retrieval · 12 · 8 first-author · 1 since 2021Software engineering, systems software and programming languages · 5 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Ensemble Instance Selection for Medical Datasets: Single-Stage Parallel and Two-Stage Serial Ensemble Approaches
abstract
ABSTRACT Instance selection is an important data preprocessing step in data mining that removes noisy data from training datasets to build more effective models and reduces dataset size to improve computational efficiency during model training. Traditional instance selection algorithms have varying strengths and weaknesses, making it difficult for any single method to consistently achieve optimal performance across diverse datasets. Therefore, in this paper, two novel ensemble instance selection methods are proposed, namely Single‐Stage Parallel Ensemble (SSPE) and Two‐Stage Serial Ensemble (TSSE), which are based on combining multiple instance selection algorithms in the parallel and serial manners. The proposed ensemble approaches aim to improve data reduction and classification performance. Extensive experiments are conducted on 17 medical datasets of varying sizes, evaluating four instance selection algorithms, CNN, ENN, IPF, and GA, alongside three classifiers: KNN, SVM, and RF. Results demonstrate that certain SSPE combinations, particularly the union of ENN and IPF, outperform single baseline algorithms in classification accuracy and AUC while maintaining effective data reduction. Although TSSE achieves higher reduction rates, its classification performance is inferior to SSPE. Overall, the proposed methods serve as effective preprocessing tools for medical data and provide a strong baseline for future ensemble instance selection research.
Min-Wei Huang, Chih-Fong Tsai, Wun-Sin Wu
Concurr. Comput. Pract. Exp.2
2025 Dimensionality Reduction Strategies for Classification: ML Versus DL Approaches and Their Combinations
abstract
ABSTRACT Dimensionality reduction plays a vital role in enhancing the performance of data classification tasks by reducing the complexity of the feature space. This study examines the effectiveness of integrating dimensionality reduction techniques with classification algorithms across four strategic configurations: (1) machine learning (ML)‐based dimensionality reduction with ML classifiers, (2) deep learning (DL)‐based dimensionality reduction with DL classifiers, and two heterogeneous combinations that mix ML and DL methods. Using 20 benchmark datasets from diverse domains, with feature dimensions ranging from 44 to 19,993, we systematically evaluate and compare these configurations. The dimensionality reduction methods include three ML‐based feature selection techniques, Genetic Algorithm (GA), Information Gain (IG), and the C4.5 decision tree, and four DL‐based feature extraction approaches, Autoencoder (AE), Sparse Autoencoder (SAE), Denoising Autoencoder (DAE), and Variational Autoencoder (VAE). For classification, Support Vector Machine (SVM) and k‐Nearest Neighbours (KNN) are used as ML classifiers, while Multilayer Perceptron (MLP) and Deep Belief Network (DBN) serve as DL classifiers. Experimental results show that SAE consistently produces the most compact feature sets and improves classification performance, with the SAE + MLP combination achieving the best overall results. Furthermore, we explore ensemble dimensionality reduction strategies that integrate multiple algorithms. Although the best ensemble approach slightly outperforms the SAE + MLP model, the observed performance improvements are not statistically significant, that is, 0.839 versus 0.836 for AUC rates. In addition, SAE achieves a significantly higher dimensionality reduction rate compared to the best ensemble method (63% vs. 18%).
Chihli Hung, Chih-Fong Tsai, Ming-hui Wu
Expert Syst. J. Knowl. Eng.2
2024 Hybrid Technical-Visual Features for Stock Prediction
Chih-Fong Tsai, Ya-Han Hu, Ming-Chang Wang, Kang Ernest Liu
AINA (3)1
2024 The effect of review visibility and diagnosticity on review helpfulness - An accessibility-diagnosticity theory perspective
Chih-Fong Tsai, Ya-Han Hu, Chen-Wei Hu
Decis. Support Syst.2
2024 Majority re-sampling via sub-class clustering for imbalanced datasets
abstract
Many real world domain problem datasets are class imbalanced where the number of data in a given class is much less than in the other classes. In related literatures, under- and over-sampling techniques are widely used techniques to re-balance the class imbalanced datasets. However, their limitations include the risk of removing representative majority class data samples and the overfitting problem because of generating a large number of synthetic minority class data samples. Therefore, a novel approach, namely Majority Re-sampling visa Sub-class Clustering (MRSC) is introduced. It uses a clustering algorithm to group the majority class data into several clusters, i.e. sub-classes. Then, a new training set containing multiple sub-classes and a minority class is produced, after which the classifier is trained using this new multi-class dataset which has a lower imbalance ratio than the original dataset. The experimental results obtained using 44 two-class imbalanced datasets show that MRSC combined with the k-NN classifiers, including single and ensemble classifiers, significantly outperforms the other classifiers as well as seven state-of-the-art re-sampling approaches. Moreover, for the clustering algorithms based on affinity propagation and k-means, very similar results can be produced, without significant differences in performance, which indicate the stability of MRSC.
Shih-Wen Ke, Chih-Fong Tsai, Yi-Ying Pan, Wei-Chao Lin
J. Exp. Theor. Artif. Intell.2
2024 Missing value imputation and the effect of feature normalisation on financial distress prediction
abstract
In this paper, we focus on comparing the imputation performance of different deep and machine learning techniques on nine related datasets containing different missing rates ranging from 10% to 50%. Moreover, since each feature value ranges differently, such as total liability/equity ratio and earnings per share, the effect of feature normalisation on the imputation results is also examined to see whether normalising the feature values after missing value imputation can improve the prediction model performance. The experimental results show that the deep neural network technique does not necessarily perform better than the traditional machine learning technique for missing value imputation. In particular, the random forest imputation model performs the best, whereas the k-nearest neighbour method is the second best imputation model in terms of the AUC rates and type II errors. The performance in improvement of prediction models after performing feature normalisation is heavily dependent on the chosen classification technique. There is thus no need to consider the normalisation step when the random forest classifier is used. It is found that the deep neural network and support vector machine classifiers can significantly outperform those without feature normalisation.
Kuen-Liang Sue, Chih-Fong Tsai, Hau-Min Tsau
J. Exp. Theor. Artif. Intell.2
2023 Instance selection using one-versus-all and one-versus-one decomposition approaches in multiclass classification datasets
abstract
Abstract Instance is important in data analysis and mining; it filters out unrepresentative, redundant, or noisy data from a given training set to obtain effective model learning. Various instance selection algorithms are proposed in the literature, and their potential and applicability in data cleaning and preprocessing steps are demonstrated. For multiclass classification datasets, the existing instance selection algorithms must deal with all the instances across the different classes simultaneously to produce a reduced training set. Generally, every multiclass classification dataset can be regarded as a complex domain problem, which can be effectively solved using the divide‐and‐conquer principle. In this study, the one‐versus‐all (OVA) and one‐versus‐one (OVO) decomposition approaches were used to decompose a multiclass dataset into multiple binary class datasets. These approaches have been widely employed when constructing the classifier but have never been considered in instance selection. The results of instance selection performance obtained with the OVA, OVO, and baseline approaches were assessed and compared for 20 different domain multiclass datasets as the first study and five medical domain datasets as the validation study. Furthermore, three instance selection algorithms were compared, including IB3, DROP3, and GA. The results demonstrate that using the OVO approach to perform instance selection can make the support vector machine (SVM) and k‐nearest neighbour (k‐NN) classifiers perform significantly better than the OVA and baseline approaches in terms of the area under the ROC curve (AUC) rate, regardless of the instance selection algorithm used. Moreover, the OVO approach can provide reasonably good data reduction rates and processing times, which are all better than those of the OVA approach.
Ching-Lin Fang, Ming-Chang Wang, Chih-Fong Tsai, Wei-Chao Lin, Pei-Qi Liao
Expert Syst. J. Knowl. Eng.3
2023 An adaptive growing grid model for a non-stationary environment
Chihli Hung, Stefan Wermter, Yu-Liang Chi, Chih-Fong Tsai
Neurocomputing4
2022 A dynamic time warping approach for handling class imbalanced medical datasets with missing values: A case study of protein localization site prediction
Ling-Chien Hung, Ya-Han Hu, Chih-Fong Tsai, Min-Wei Huang
Expert Syst. Appl.3
2022 An investigation of solutions for handling incomplete online review datasets with missing values
abstract
Online review helpfulness prediction is an important research issue in electronic commerce and data mining. However, the collected datasets used for the analysis and prediction of the helpfulness of online reviews often contain some missing attribute values, such as reviewer background and rating information. In related literatures, many studies have either used the case deletion approach to remove the data containing missing values or considered the imputation of missing values by the mean/mode method. However, none of them consider the direct handling approach without missing value imputation for online review datasets by decision tree-related techniques. Therefore, in this paper, we investigate the suitability of different types of approaches to solve the incomplete dataset problem of online reviews. Specifically, for missing value imputation, several supervised learning techniques including MICE, KNN, SVM, and CART are examined. Moreover, for the direct handling approach without missing value imputation, CART is also performed for this task. The experimental results based on the TripAdvisor dataset for review helpfulness prediction show that the approach where incomplete online review datasets are handled directly without imputation by CART significantly outperforms the other approaches, including case deletion and missing value imputation approaches.
Ya-Han Hu, Chih-Fong Tsai
J. Exp. Theor. Artif. Intell.2
2022 Empirical comparison of supervised learning techniques for missing value imputation
Chih-Fong Tsai, Ya-Han Hu
Knowl. Inf. Syst.1
2022 Deep learning for missing value imputation of continuous data and the effect of data discretization
Wei-Chao Lin, Chih-Fong Tsai, Jia Rong Zhong
Knowl. Based Syst.2
2021 Towards missing electric power data imputation for energy management systems
Ming-Chang Wang, Chih-Fong Tsai, Wei-Chao Lin
Expert Syst. Appl.2
2020 Ensemble feature selection in high dimension, low sample size datasets: Parallel and serial combination approaches
Chih-Fong Tsai, Ya-Ting Sung
Knowl. Based Syst.1
2019 Feature selection in single and ensemble learning-based bankruptcy prediction models
abstract
Abstract Feature selection is an important data preprocessing step for the construction of an effective bankruptcy prediction model. The prediction performance can be affected by the employed feature selection and classification techniques. However, there have been very few studies of bankruptcy prediction that identify the best combination of feature selection and classification techniques. In this study, two types of feature selection methods, including filter‐ and wrapper‐based methods, are considered, and two types of classification techniques, including statistical and machine learning techniques, are employed in the development of the prediction methods. In addition, bagging and boosting ensemble classifiers are also constructed for comparison. The experimental results based on three related datasets that contain different numbers of input features show that the genetic algorithm as the wrapper‐based feature selection method performs better than the filter‐based one by information gain. It is also shown that the lowest prediction error rates for the three datasets are provided by combining the genetic algorithm with the naïve Bayes and support vector machine classifiers without bagging and boosting.
Wei-Chao Lin, Yu-Hsin Lu, Chih-Fong Tsai
Expert Syst. J. Knowl. Eng.3
2019 The optimal combination of feature selection and data discretization: An empirical study
Chih-Fong Tsai
Inf. Sci.1
2019 Under-sampling class imbalanced datasets by combining clustering analysis and instance selection
Chih-Fong Tsai, Wei-Chao Lin, Ya-Han Hu, Guan-Ting Yao
Inf. Sci.1
2018 A novel classifier ensemble approach for financial distress prediction
Deron Liang, Chih-Fong Tsai, An-Jie Dai, William Eberle
Knowl. Inf. Syst.2
2018 A class center based approach for missing value imputation
Chih-Fong Tsai, Miao-Ling Li, Wei-Chao Lin
Knowl. Based Syst.1
2017 Soft estimation by hierarchical classification and regression
Shih-Wen Ke, Wei-Chao Lin, Chih-Fong Tsai, Ya-Han Hu
Neurocomputing3
2017 Clustering-based undersampling in class-imbalanced data
Wei-Chao Lin, Chih-Fong Tsai, Ya-Han Hu, Jing-Shang Jhang
Inf. Sci.2
2016 Data preprocessing issues for incomplete medical datasets
abstract
Abstract While there is an ample amount of medical information available for data mining, many of the datasets are unfortunately incomplete – missing relevant values needed by many machine learning algorithms. Several approaches have been proposed for the imputation of missing values, using various reasoning steps to provide estimations from the observed data. One of the important steps in data mining is data preprocessing, where unrepresentative data is filtered out of the data to be mined. However, none of the related studies about missing value imputation consider performing a data preprocessing step before imputation. Therefore, the aim of this study is to examine the effect of two preprocessing steps, feature and instance selection, on missing value imputation. Specifically, eight different medical‐related datasets are used, containing categorical, numerical and mixed types of data. Our experimental results show that imputation after instance selection can produce better classification performance than imputation alone. In addition, we will demonstrate that imputation after feature selection does not have a positive impact on the imputation result.
Min-Wei Huang, Wei-Chao Lin, Chih-Wen Chen, Shih-Wen Ke, Chih-Fong Tsai, William Eberle
Expert Syst. J. Knowl. Eng.5
2016 Intangible assets evaluation: The machine learning perspective
Chih-Fong Tsai, Yu-Hsin Lu, Yu-Chung Hung, David C. Yen
Neurocomputing1
2016 Keypoint selection for efficient bag-of-words feature generation and effective image classification
Wei-Chao Lin, Chih-Fong Tsai, Zong-Yao Chen, Shih-Wen Ke
Inf. Sci.2
2016 Combining instance selection for better missing value imputation
Chih-Fong Tsai, Fu-Yu Chang
J. Syst. Softw.1
2016 Big data mining with parallel computing: A comparison of distributed and MapReduce methodologies
Chih-Fong Tsai, Wei-Chao Lin, Shih-Wen Ke
J. Syst. Softw.1
2016 3D model retrieval by sample based alignment
Zong-Yao Chen, Wei-Chao Lin, Chih-Fong Tsai, Shih-Wen Ke
J. Vis. Commun. Image Represent.3
2015 The Identification of Prolonged Length of Stay for Surgery Patients
abstract
When the hospitalization periods of an unexpectedly high number of patients are extended, the income of a hospital is substantially affected and the rate of hospital bed occupancy increases. Because none of the currently available score systems can be used to evaluate the possibility that patients who require urgent surgery in a single department prolong their length of stay (LOS), this study attempts to build a prolonged LOS prediction model utilizing a number of supervised learning techniques. This study involved analyzing the complete historical medical records and lab data of 897 clinical cases in which surgeries were performed by general surgery physicians. These clinical cases were divided into an urgent operation (UO) group comprising 462 cases and a non-UO group comprising 434 cases to develop a prolonged LOS prediction model by using several supervised learning techniques. The results indicated that the random forest method constituted the most accurate and stable prediction model. This study demonstrated that supervised learning techniques can be used to analyze patient medical records to accurately predict a prolonged LOS, thus, supervised learning techniques can serve as valuable reference tools for patient prognoses. The developed prediction models can facilitate the decision making of physicians when patients require surgery and increase patient safety.
Mao-Te Chuang, Ya-Han Hu, Chih-Fong Tsai, Chia-Lun Lo, Wei-Chao Lin
SMC3
2015 Customer segmentation issues and strategies for an automobile dealership with two clustering techniques
abstract
Abstract Companies can use customer segmentation to group customers with similar characteristics together and identify the differences between groups to develop marketing strategies. This study investigates the problem of customer segmentation in relation to automotive customer relationship management and presents a real case study of an automobile dealer in Taiwan. Although several past studies have adopted different clustering techniques with which to group customer attributes, few have simultaneously considered customer transaction behaviour and customer satisfaction variables. In addition, most previous work has used only a single clustering method for customer segmentation, which results in unreliable results and leads to inadequate marketing decisions. Therefore, in this study, we consider two clustering techniques, k‐means and expectation maximization, and compare their results for correctness. The experimental results show that four customer groups are identified with both clustering methods: loyal, potential, VIP and churn customer groups. Based on the segmentation results, several customized marketing strategies aimed at each of the four customer groups are suggested to improve the quality of services for effective customer relationship management.
Chih-Fong Tsai, Ya-Han Hu, Yu-Hsin Lu
Expert Syst. J. Knowl. Eng.1
2015 The effect of low-level image features on pseudo relevance feedback
Wei-Chao Lin, Zong-Yao Chen, Shih-Wen Ke, Chih-Fong Tsai, Wei-Yang Lin
Neurocomputing4
2015 Factors affecting rocchio-based pseudorelevance feedback in image retrieval
abstract
Pseudorelevance feedback (PRF) was proposed to solve the limitation of relevance feedback (RF), which is based on the user‐in‐the‐loop process. In PRF, the top‐k retrieved images are regarded as PRF. Although the PRF set contains noise, PRF has proven effective for automatically improving the overall retrieval result. To implement PRF, the Rocchio algorithm has been considered as a reasonable and well‐established baseline. However, the performance of Rocchio‐based PRF is subject to various representation choices (or factors). In this article, we examine these factors that affect the performance of Rocchio‐based PRF, including image‐feature representation, the number of top‐ranked images, the weighting parameters of Rocchio, and similarity measure. We offer practical insights on how to optimize the performance of Rocchio‐based PRF by choosing appropriate representation choices. Our extensive experiments on NUS‐WIDE‐LITE and Caltech 101 + Corel 5000 data sets show that the optimal feature representation is color moment + wavelet texture in terms of retrieval efficiency and effectiveness. Other representation choices are that using top‐20 ranked images as pseudopositive and pseudonegative feedback sets with the equal weight (i.e., 0.5) by the correlation and cosine distance functions can produce the optimal retrieval result.
Chih-Fong Tsai, Ya-Han Hu, Zong-Yao Chen
J. Assoc. Inf. Sci. Technol.1
2015 Learning to detect representative data for large scale instance selection
Wei-Chao Lin, Chih-Fong Tsai, Shih-Wen Ke, Chia-Wen Hung, William Eberle
J. Syst. Softw.2
2015 A privilege-based visual secret sharing model
Young-Chang Hou, Zen-Yu Quan, Chih-Fong Tsai
J. Vis. Commun. Image Represent.3
2015 The effect of feature selection on financial distress prediction
Deron Liang, Chih-Fong Tsai, Hsin-Ting Wu
Knowl. Based Syst.2
2015 CANN: An intrusion detection system based on combining cluster centers and nearest neighbors
Wei-Chao Lin, Shih-Wen Ke, Chih-Fong Tsai
Knowl. Based Syst.3
2015 Image retargeting using RGB-D camera
Wei-Yang Lin, Chih-Fong Tsai, Pei-Chen Wu, Bo-Rong Chen
Multim. Tools Appl.2
2015 Instance selection by genetic-based biological algorithm
Zong-Yao Chen, Chih-Fong Tsai, William Eberle, Wei-Chao Lin, Shih-Wen Ke
Soft Comput.2
2014 Towards high dimensional instance selection: An evolutionary approach
Chih-Fong Tsai, Zong-Yao Chen
Decis. Support Syst.1
2014 Evolutionary instance selection for text classification
Chih-Fong Tsai, Zong-Yao Chen, Shih-Wen Ke
J. Syst. Softw.1
2013 SVOIS: Support Vector Oriented Instance Selection for text classification
Chih-Fong Tsai
Inf. Syst.1
2013 Block-based progressive visual secret sharing
Young-Chang Hou, Zen-Yu Quan, Chih-Fong Tsai, A.-Yu Tseng
Inf. Sci.3
2013 Genetic algorithms in feature and instance selection
Chih-Fong Tsai, William Eberle, Chi-Yuan Chu
Knowl. Based Syst.1
2012 Simple instance selection for bankruptcy prediction
Chih-Fong Tsai, Kai-Chun Cheng
Knowl. Based Syst.1
2012 Determinants of intangible assets value: The data mining approach
Chih-Fong Tsai, Yu-Hsin Lu, David C. Yen
Knowl. Based Syst.1
2012 Machine Learning in Financial Crisis Prediction: A Survey
abstract
For financial institutions, the ability to predict or forecast business failures is crucial, as incorrect decisions can have direct financial consequences. Bankruptcy prediction and credit scoring are the two major research problems in the accounting and finance domain. In the literature, a number of models have been developed to predict whether borrowers are in danger of bankruptcy and whether they should be considered a good or bad credit risk. Since the 1990s, machine-learning techniques, such as neural networks and decision trees, have been studied extensively as tools for bankruptcy prediction and credit score modeling. This paper reviews 130 related journal papers from the period between 1995 and 2010, focusing on the development of state-of-the-art machine-learning techniques, including hybrid and ensemble classifiers. Related studies are compared in terms of classifier design, datasets, baselines, and other experimental factors. This paper presents the current achievements and limitations associated with the development of bankruptcy-prediction and credit-scoring models employing machine learning. We also provide suggestions for future research.
Wei-Yang Lin, Ya-Han Hu, Chih-Fong Tsai
IEEE Trans. Syst. Man Cybern. Part C3
2010 Combining multiple feature selection methods for stock prediction: Union, intersection, and multi-intersection approaches
Chih-Fong Tsai, Yu-Chieh Hsiao
Decis. Support Syst.1
2010 Variable selection by association rules for customer churn prediction of multimedia on demand
Chih-Fong Tsai, Mao-Yuan Chen
Expert Syst. Appl.1
2010 A triangle area based nearest neighbors approach to intrusion detection
Chih-Fong Tsai, Chia-Ying Lin
Pattern Recognit.1
2009 Earnings management prediction: A pilot study of combining neural networks and decision trees
Chih-Fong Tsai, Yen-Jiun Chiou
Expert Syst. Appl.1
2009 Intrusion detection by machine learning: A review
Chih-Fong Tsai, Yu-Feng Hsu, Chia-Ying Lin, Wei-Yang Lin
Expert Syst. Appl.1
2009 Customer churn prediction by hybrid neural networks
Chih-Fong Tsai, Yu-Hsin Lu
Expert Syst. Appl.1
2009 Component-based software version management based on a Component-Interface Dependency Matrix
Shi-Ming Huang, Chih-Fong Tsai, Po-Chun Huang
J. Syst. Softw.2
2009 Feature selection in bankruptcy prediction
Chih-Fong Tsai
Knowl. Based Syst.1
2008 Financial decision support using neural networks and support vector machines
abstract
Abstract:Bankruptcy prediction and credit scoring are the two important problems facing financial decision support. The multilayer perceptron (MLP) network has shown its applicability to these problems and its performance is usually superior to those of other traditional statistical models. Support vector machines (SVMs) are the core machine learning techniques and have been used to compare with MLP as the benchmark. However, the performance of SVMs is not fully understood in the literature because an insufficient number of data sets is considered and different kernel functions are used to train the SVMs. In this paper, four public data sets are used. In particular, three different sizes of training and testing data in each of the four data sets are considered (i.e. 3:7, 1:1 and 7:3) in order to examine and fully understand the performance of SVMs. For SVM model construction, the linear, radial basis function and polynomial kernel functions are used to construct the SVMs. Using MLP as the benchmark, the SVM classifier only performs better in one of the four data sets. On the other hand, the prediction results of the MLP and SVM classifiers are not significantly different for the three different sizes of training and testing data.
Chih-Fong Tsai
Expert Syst. J. Knowl. Eng.1
2008 The development of audit detection risk assessment system: Using the fuzzy theory and audit risk model
She-I Chang, Chih-Fong Tsai, Dong-Her Shih, Chia-Ling Hwang
Expert Syst. Appl.2
2008 A hybrid financial analysis model for business failure prediction
Shi-Ming Huang, Chih-Fong Tsai, David C. Yen, Yin-Lin Cheng
Expert Syst. Appl.2
2008 Market segmentation based on hierarchical self-organizing map for markets of multimedia on demand
Chihli Hung, Chih-Fong Tsai
Expert Syst. Appl.2
2008 Using neural network ensembles for bankruptcy prediction and credit scoring
Chih-Fong Tsai, Jhen-Wei Wu
Expert Syst. Appl.1
2007 Image mining by spectral features: A case study of scenery image classification
Chih-Fong Tsai
Expert Syst. Appl.1
2006 Qualitative evaluation of automatic assignment of keywords to images
Chih-Fong Tsai, Kenneth McGarry, John Tait
Inf. Process. Manag.1
2006 CLAIRE: A modular support vector image indexing and classification system
abstract
Many users of image retrieval systems would prefer to express initial queries using keywords. However, manual keyword indexing is very time-consuming. Therefore, a content-based image retrieval system which can automatically assign keywords to images would be very attractive. Unfortunately, it has proved very challenging to build such systems, except where either the image domain is restricted or the keywords relate only to low-level concepts such as color. This article presents a novel image indexing and classification system, called CLAIRE (CLAssifying Images for REtrieval), composed of one image processing module and three modules of support vector machines for color, texture, and high-level concept classification for keyword assignment. The experimental prototype system described here assigns up to five keywords selected from a controlled vocabulary of 60 terms to each image. The system is trained offline by 1639 examples from the Corel stock photo library. For evaluation, five judges reviewed a sample of 800 unknown images to identify which automatically assigned keywords were actually relevant to the image. The system proved to have an 80% probability to assign at least one relevant keyword to an image.
Chih-Fong Tsai, Kenneth McGarry, John Tait
ACM Trans. Inf. Syst.1
2005 Training support vector machines based on stacked generalization for image classification
Chih-Fong Tsai
Neurocomputing1
2003 Image classification using hybrid neural networks
abstract
Use of semantic content is one of the major issues which needs to be addressed for improving image retrieval effectiveness. We present a new approach to classify images based on the combination of image processing techniques and hybrid neural networks. Multiple keywords are assigned to an image to represent its main contents, i.e. semantic content. Images are divided into a number of regions and colour and texture features are extracted. The first classifier, a self-organising map (SOM) clusters similar images based on the extracted features. Then, regions of the representative images of these clusters were labeled and used to train the second classifier, composed of several support vector machines (SVMs). Initial experiments on the accuracy of keyword assignment for a small vocabulary are reported.
Chih-Fong Tsai, Kenneth McGarry, John Tait
SIGIR1