Shinn-Ying Ho

dblp:18/2715 · DBLP profile ↗
← Back
74ranked-venue papers
22as first author
3since 2021 · last 2025
0000-0002-1901-8353ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 33 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 31 · 13 first-authorHuman-computer interaction and ubiquitous computing · 7 · 4 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
5 papers
Bioinformatics and computational biology · 100%
Artificial intelligence
1 paper
3D vision · 100%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › biological network › network biology › network inference
gene regulatory network inference
0.722020
GREMA: modelling of emulated gene regulatory networks with confidence levels based on evolutionary intelligence to cope with the underdetermined problem · Bioinform. 2020
GeNOSA: inferring and experimentally supporting quantitative gene regulatory networks in prokaryotes · Bioinform. 2015
Bioinformatics and computational biology › systems biology
parameter estimation
0.412020
GREMA: modelling of emulated gene regulatory networks with confidence levels based on evolutionary intelligence to cope with the underdetermined problem · Bioinform. 2020
Bioinformatics and computational biology
negative sampling
0.312017
ESA-UbiSite: accurate prediction of human ubiquitination sites by identifying a set of effective negatives · Bioinform. 2017
Bioinformatics and computational biology › proteomics
post-translational modification prediction
0.312017
ESA-UbiSite: accurate prediction of human ubiquitination sites by identifying a set of effective negatives · Bioinform. 2017
Bioinformatics and computational biology › protein sequence analysis › protein sequence annotation
ubiquitination site prediction
0.312017
ESA-UbiSite: accurate prediction of human ubiquitination sites by identifying a set of effective negatives · Bioinform. 2017
Bioinformatics and computational biology › immunoinformatics
epitope prediction
0.112007
POPI: predicting immunogenicity of MHC class I binding peptides by mining informative physicochemical properties · Bioinform. 2007
Bioinformatics and computational biology
immunoinformatics
0.112007
POPI: predicting immunogenicity of MHC class I binding peptides by mining informative physicochemical properties · Bioinform. 2007
Bioinformatics and computational biology › structural biology › protein structure and function › protein stability prediction
mutation-induced stability change
0.112007
iPTREE-STAB: interpretable decision tree based method for predicting protein stability changes upon mutations · Bioinform. 2007
Bioinformatics and computational biology › structural biology › protein structure and function
protein stability prediction
0.112007
iPTREE-STAB: interpretable decision tree based method for predicting protein stability changes upon mutations · Bioinform. 2007

Methods — techniques the papers use, named apart from their topics

genetic algorithm · 0.5s-system · 0.4hill function · 0.4evolutionary intelligence · 0.4support vector machine · 0.4physicochemical properties · 0.3evolutionary screening algorithm · 0.3network component analysis · 0.2global optimization · 0.2adaptive boosting · 0.1sensitivity analysis · 0.0projective geometry · 0.0
YearPublicationVenuePosition
2025 miSAM: Robust MicroRNA Expression Estimation Using Bi-Objective Evolutionary Learning Algorithm in Cancer Transcriptomics
abstract
MicroRNAs (miRNAs) play crucial regulatory roles in cancer biology, but accurately quantifying their expression remains a significant unmet challenge in single-cell and spatial transcriptomics. Current sequencing technologies predominantly capture polyadenylated messenger RNAs (mRNAs), rendering them incapable of directly profiling miRNAs, which lack poly(A) tails. To address this gap, we thus propose miSAM, a novel computational framework for estimation of miRNA expressions based on a bi-objective combinatorial genetic algorithm in conjunction with support vector regression. The miSAM jointly optimizes the selection of a minimal subset of mRNAs, called signatures, while maximizing the Spearman correlation coefficient (SCC) between inferred and actual miRNA expression levels. Evaluation on The Cancer Genome Atlas (TCGA)-BRCA dataset, comprising 1,095 breast cancer samples with expression profiles of 1,881 miRNAs and 19,937 mRNAs, demonstrates the effectiveness of miSAM. Using the top five prognostic miRNA biomarkers for breast cancer, miSAM achieved a mean SCC of 0.633 on the test set while utilizing only 32.6 mRNAs in average, significantly outperforming baseline approaches including: (1) differential expression filtering-based SVR using 887 mRNAs (SCC = 0.562), and (2) LASSO-based SVR using 147.4 mRNAs (SCC = 0.470). miSAM also outperformed XGBoost, which yielded a SCC ofapproximately 0.45–0.50 across various cancer types. Furthermore, the identified mRNAs signatures offer explainable insights into the regulatory associations between miRNAs and their corresponding mRNA targets. These results underscore miSAM’s potential as a robust, interpretable, and scalable tool for miRNA inference in spatial transcriptomics, single-cell sequencing analyses, and precision oncology.
Yann-Lin Ho, Yann-Jen Ho, Shinn-Ying Ho, Tzong-Yi Lee
CIBCB3
2025 EL-DRP: An Evolutionary Learning-Based Multi-Omics Framework for Drug Response Prediction with Interpretable Biomarker Selection
abstract
Cancer is a complex disease driven by diverse genetic, epigenetic, and microenvironmental alterations that result in dysregulated cell proliferation and therapeutic resistance. Both inter- and intratumoral heterogeneity further complicate treatment outcomes, highlighting the critical need for effective and efficient drug response prediction in precision oncology. While deep learning models have yielded robust predictive performance, their limited interpretability remains a critical barrier to clinical translation. To overcome this, we propose EL-DRP, an evolutionary learning framework for drug response prediction that integrates multi-omics data, including mRNA expression, copy number variation, and single nucleotide polymorphisms. Central to EL-DRP is the use of an Inheritable Bi-objective Combinatorial Genetic Algorithm (IBCGA) for optimized feature selection within each omics modality. IBCGA identifies minimal yet informative subsets of features, which are subsequently integrated to construct a unified multi-omics predictive model. EL-DRP achieved an average AUROC of 0.791 ± 0.072 for 21 drugs using an average of 106 biomarkers per drug, most of which align with known drug mechanisms of action. For instance, among the camptothecin response–associated biomarkers, KDM4A and SCML2 are known to participate in camptothecin-induced DNA damage response, whereas SALL2, RBM10, BAHCC1, and PBRM1 regulate chromatin structure and genome maintenance, corroborating their mechanistic relevance. Importantly, this minimal biomarker panel alone retains robust predictive power with explainable biomarkers. These results provide a new perspective for discovering potential biomarkers associated with drug response prediction and lay the groundwork for future studies integrating clinical data with translational potential.
Yann-Jen Ho, Paik You Sheng, Yu-Ruo Chen, Shinn-Ying Ho, Tzong-Yi Lee
CIBCB5
2025 Establishing the Asia & Pacific Bioinformatics Joint Congress: a historic milestone in regional bioinformatics collaboration
abstract
In response to the need for greater cohesion among regional conferences, the Asia Pacific Bioinformatics Network (APBioNET) set out in 2015 to realize a long-held aspiration-a single, unifying bioinformatics "super conference" for the Asia & Pacific community. Nearly a decade of persistence, coordination, and coalition-building led to the inaugural Asia & Pacific Bioinformatics Joint Congress (APBJC2024) in Okinawa, Japan. Now established as a triennial event, APBJC stands as a testament to the power of collective vision and shared purpose, offering a unifying platform for regional collaboration and scientific exchange. Tagline: Bringing a Region Together: The Making of APBJC.
Asif M. Khan, Susumu Goto, Kenta Nakai, Limsoon Wong, Diane E. Kovats, Shinya Ikematsu, Yoshihiro Yamanishi, Nurul Salwanie Che Wahid, Pradeep Eranti, Yi-Ping Phoebe Chen, Tae-Min Kim, Shinn-Ying Ho, Jessica Cara Mar, Wataru Iwasaki 0001, Jayaraman Valadi, Prashanth Suravajhala, Christian Schönbach, Tin Wee Tan, Shoba Ranganathan, Kiyoko F. Aoki-Kinoshita
Briefings Bioinform.14
2020 GREMA: modelling of emulated gene regulatory networks with confidence levels based on evolutionary intelligence to cope with the underdetermined problem
abstract
MOTIVATION: Non-linear ordinary differential equation (ODE) models that contain numerous parameters are suitable for inferring an emulated gene regulatory network (eGRN). However, the number of experimental measurements is usually far smaller than the number of parameters of the eGRN model that leads to an underdetermined problem. There is no unique solution to the inference problem for an eGRN using insufficient measurements. RESULTS: This work proposes an evolutionary modelling algorithm (EMA) that is based on evolutionary intelligence to cope with the underdetermined problem. EMA uses an intelligent genetic algorithm to solve the large-scale parameter optimization problem. An EMA-based method, GREMA, infers a novel type of gene regulatory network with confidence levels for every inferred regulation. The higher the confidence level is, the more accurate the inferred regulation is. GREMA gradually determines the regulations of an eGRN with confidence levels in descending order using either an S-system or a Hill function-based ODE model. The experimental results showed that the regulations with high-confidence levels are more accurate and robust than regulations with low-confidence levels. Evolutionary intelligence enhanced the mean accuracy of GREMA by 19.2% when using the S-system model with benchmark datasets. An increase in the number of experimental measurements may increase the mean confidence level of the inferred regulations. GREMA performed well compared with existing methods that have been previously applied to the same S-system, DREAM4 challenge and SOS DNA repair benchmark datasets. AVAILABILITY AND IMPLEMENTATION: All of the datasets that were used and the GREMA-based tool are freely available at https://nctuiclab.github.io/GREMA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ming-Ju Tsai, Jyun-Rong Wang, Shinn-Jang Ho, Li-Sun Shu, Wen-Lin Huang, Shinn-Ying Ho
Bioinform.6
2018 Recognition-based character segmentation for multi-level writing style
Papangkorn Inkeaw, Jakramate Bootkrajang, Phasit Charoenkwan, Sanparith Marukatat, Shinn-Ying Ho, Jeerayut Chaijaruwanich
Int. J. Document Anal. Recognit.5
2017 ESA-UbiSite: accurate prediction of human ubiquitination sites by identifying a set of effective negatives
abstract
Motivation: Numerous ubiquitination sites remain undiscovered because of the limitations of mass spectrometry-based methods. Existing prediction methods use randomly selected non-validated sites as non-ubiquitination sites to train ubiquitination site prediction models. Results: We propose an evolutionary screening algorithm (ESA) to select effective negatives among non-validated sites and an ESA-based prediction method, ESA-UbiSite, to identify human ubiquitination sites. The ESA selects non-validated sites least likely to be ubiquitination sites as training negatives. Moreover, the ESA and ESA-UbiSite use a set of well-selected physicochemical properties together with a support vector machine for accurate prediction. Experimental results show that ESA-UbiSite with effective negatives achieved 0.92 test accuracy and a Matthews's correlation coefficient of 0.48, better than existing prediction methods. The ESA increased ESA-UbiSite's test accuracy from 0.75 to 0.92 and can improve other post-translational modification site prediction methods. Availability and Implementation: An ESA-UbiSite-based web server has been established at http://iclab.life.nctu.edu.tw/iclab_webtools/ESAUbiSite/ . Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Jyun-Rong Wang, Wen-Lin Huang, Ming-Ju Tsai, Kai-Ti Hsu, Hui-Ling Huang, Shinn-Ying Ho
Bioinform.6
2017 Recognition of handwritten Lanna Dhamma characters using a set of optimally designed moment features
Papangkorn Inkeaw, Phasit Charoenkwan, Hui-Ling Huang, Sanparith Marukatat, Shinn-Ying Ho, Jeerayut Chaijaruwanich
Int. J. Document Anal. Recognit.5
2016 A hydrophobic spine stabilizes a surface-exposed α-helix according to analysis of the solvent-accessible surface area
abstract
BACKGROUND: Most of hydrophilic and hydrophobic residues are thought to be exposed and buried in proteins, respectively. In contrast to the majority of the existing studies on protein folding characteristics using protein structures, in this study, our aim was to design predictors for estimating relative solvent accessibility (RSA) of amino acid residues to discover protein folding characteristics from sequences. METHODS: The proposed 20 real-value RSA predictors were designed on the basis of the support vector regression method with a set of informative physicochemical properties (PCPs) obtained by means of an optimal feature selection algorithm. Then, molecular dynamics simulations were performed for validating the knowledge discovered by analysis of the selected PCPs. RESULTS: The RSA predictors had the mean absolute error of 14.11% and a correlation coefficient of 0.69, better than the existing predictors. The hydrophilic-residue predictors preferred PCPs of buried amino acid residues to PCPs of exposed ones as prediction features. A hydrophobic spine composed of exposed hydrophobic residues of an α-helix was discovered by analyzing the PCPs of RSA predictors corresponding to hydrophobic residues. For example, the results of a molecular dynamics simulation of wild-type sequences and their mutants showed that proteins 1MOF and 2WRP_H16I (Protein Data Bank IDs), which have a perfectly hydrophobic spine, have more stable structures than 1MOF_I54D and 2WRP do (which do not have a perfectly hydrophobic spine). CONCLUSIONS: We identified informative PCPs to design high-performance RSA predictors and to analyze these PCPs for identification of novel protein folding characteristics. A hydrophobic spine in a protein can help to stabilize exposed α-helices.
Yi-Fan Liou, Hui-Ling Huang, Shinn-Ying Ho
BMC Bioinform.3
2016 SCMBYK: prediction and characterization of bacterial tyrosine-kinases based on propensity scores of dipeptides
abstract
BACKGROUND: Bacterial tyrosine-kinases (BY-kinases), which play an important role in numerous cellular processes, are characterized as a separate class of enzymes and share no structural similarity with their eukaryotic counterparts. However, in silico methods for predicting BY-kinases have not been developed yet. Since these enzymes are involved in key regulatory processes, and are promising targets for anti-bacterial drug design, it is desirable to develop a simple and easily interpretable predictor to gain new insights into bacterial tyrosine phosphorylation. This study proposes a novel SCMBYK method for predicting and characterizing BY-kinases. RESULTS: A dataset consisting of 797 BY-kinases and 783 non-BY-kinases was established to design the SCMBYK predictor, which achieved training and test accuracies of 97.55 and 96.73%, respectively. Furthermore, the leave-one-phylum-out method was used to predict specific bacterial phyla hosts of target sequences, gaining 97.39% average test accuracy. After analyzing SCMBYK-derived propensity scores, four characteristics of BY-kinases were determined: 1) BY-kinases tend to be composed of α-helices; 2) the amino-acid content of extracellular regions of BY-kinases is expected to be dominated by residues such as Val, Ile, Phe and Tyr; 3) BY-kinases structurally resemble nuclear proteins; 4) different domains play different roles in triggering BY-kinase activity. CONCLUSIONS: The SCMBYK predictor is an effective method for identification of possible BY-kinases. Furthermore, it can be used as a part of a novel drug repurposing method, which recognizes putative BY-kinases and matches them to approved drugs. Among other results, our analysis revealed that azathioprine could suppress the virulence of M. tuberculosis, and thus be considered as a potential antibiotic for tuberculosis treatment.
Tamara Vasylenko, Yi-Fan Liou, Po-Chin Chiou, Hsiao-Wei Chu, Yung-Sung Lai, Yu-Ling Chou, Hui-Ling Huang, Shinn-Ying Ho
BMC Bioinform.8
2015 GeNOSA: inferring and experimentally supporting quantitative gene regulatory networks in prokaryotes
abstract
MOTIVATION: The establishment of quantitative gene regulatory networks (qGRNs) through existing network component analysis (NCA) approaches suffers from shortcomings such as usage limitations of problem constraints and the instability of inferred qGRNs. The proposed GeNOSA framework uses a global optimization algorithm (OptNCA) to cope with the stringent limitations of NCA approaches in large-scale qGRNs. RESULTS: OptNCA performs well against existing NCA-derived algorithms in terms of utilization of connectivity information and reconstruction accuracy of inferred GRNs using synthetic and real Escherichia coli datasets. For comparisons with other non-NCA-derived algorithms, OptNCA without using known qualitative regulations is also evaluated in terms of qualitative assessments using a synthetic Saccharomyces cerevisiae dataset of the DREAM3 challenges. We successfully demonstrate GeNOSA in several applications including deducing condition-dependent regulations, establishing high-consensus qGRNs and validating a sub-network experimentally for dose-response and time-course microarray data, and discovering and experimentally confirming a novel regulation of CRP on AscG. AVAILABILITY AND IMPLEMENTATION: All datasets and the GeNOSA framework are freely available from http://e045.life.nctu.edu.tw/GeNOSA. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yi-Hsiung Chen, Chi-Dung Yang, Ching-Ping Tseng, Hsien-Da Huang, Shinn-Ying Ho
Bioinform.5
2015 Discovery of prognostic biomarkers for predicting lung cancer metastasis using microarray and survival data
abstract
BACKGROUND: Few studies have investigated prognostic biomarkers of distant metastases of lung cancer. One of the central difficulties in identifying biomarkers from microarray data is the availability of only a small number of samples, which results overtraining. Recently obtained evidence reveals that epithelial-mesenchymal transition (EMT) of tumor cells causes metastasis, which is detrimental to patients' survival. RESULTS: This work proposes a novel optimization approach to discovering EMT-related prognostic biomarkers to predict the distant metastasis of lung cancer using both microarray and survival data. This weighted objective function maximizes both the accuracy of prediction of distant metastasis and the area between the disease-free survival curves of the non-distant and distant metastases. Seventy-eight patients with lung cancer and a follow-up time of 120 months are used to identify a set of gene markers and an independent cohort of 26 patients is used to evaluate the identified biomarkers. The medical records of the 78 patients show a significant difference between the disease-free survival times of the 37 non-distant- and the 41 distant-metastasis patients. The experimental results thus obtained are as follows. 1) The use of disease-free survival curves can compensate for the shortcoming of insufficient samples and greatly increase the test accuracy by 11.10%; and 2) the support vector machine with a set of 17 transcripts, such as CCL16 and CDKN2AIP, can yield a leave-one-out cross-validation accuracy of 93.59%, a test accuracy of 76.92%, a large disease-free survival area of 74.81%, and a mean survival prediction error of 3.99 months. The identified putative biomarkers are examined using related studies and signaling pathways to reveal the potential effectiveness of the biomarkers in prospective confirmatory studies. CONCLUSIONS: The proposed new optimization approach to identifying prognostic biomarkers by combining multiple sources of data (microarray and survival) can facilitate the accurate selection of biomarkers that are most relevant to the disease while solving the problem of insufficient samples.
Hui-Ling Huang, Yu-Chung Wu, Li-Jen Su, Yun-Ju Huang, Phasit Charoenkwan, Hua-Chin Lee, William C. Chu, Shinn-Ying Ho
BMC Bioinform.9
2015 Characterizing informative sequence descriptors and predicting binding affinities of heterodimeric protein complexes
abstract
BACKGROUND: Protein-protein interactions (PPIs) are involved in various biological processes, and underlying mechanism of the interactions plays a crucial role in therapeutics and protein engineering. Most machine learning approaches have been developed for predicting the binding affinity of protein-protein complexes based on structure and functional information. This work aims to predict the binding affinity of heterodimeric protein complexes from sequences only. RESULTS: This work proposes a support vector machine (SVM) based binding affinity classifier, called SVM-BAC, to classify heterodimeric protein complexes based on the prediction of their binding affinity. SVM-BAC identified 14 of 580 sequence descriptors (physicochemical, energetic and conformational properties of the 20 amino acids) to classify 216 heterodimeric protein complexes into low and high binding affinity. SVM-BAC yielded the training accuracy, sensitivity, specificity, AUC and test accuracy of 85.80%, 0.89, 0.83, 0.86 and 83.33%, respectively, better than existing machine learning algorithms. The 14 features and support vector regression were further used to estimate the binding affinities (Pkd) of 200 heterodimeric protein complexes. Prediction performance of a Jackknife test was the correlation coefficient of 0.34 and mean absolute error of 1.4. We further analyze three informative physicochemical properties according to their contribution to prediction performance. Results reveal that the following properties are effective in predicting the binding affinity of heterodimeric protein complexes: apparent partition energy based on buried molar fractions, relations between chemical structure and biological activity in principal component analysis IV, and normalized frequency of beta turn. CONCLUSIONS: The proposed sequence-based prediction method SVM-BAC uses an optimal feature selection method to identify 14 informative features to classify and predict binding affinity of heterodimeric protein complexes. The characterization analysis revealed that the average numbers of beta turns and hydrogen bonds at protein-protein interfaces in high binding affinity complexes are more than those in low binding affinity complexes.
Yerukala Sathipati Srinivasulu, Jyun-Rong Wang, Kai-Ti Hsu, Ming-Ju Tsai, Phasit Charoenkwan, Wen-Lin Huang, Hui-Ling Huang, Shinn-Ying Ho
BMC Bioinform.8
2015 SCMPSP: Prediction and characterization of photosynthetic proteins based on a scoring card method
abstract
BACKGROUND: Photosynthetic proteins (PSPs) greatly differ in their structure and function as they are involved in numerous subprocesses that take place inside an organelle called a chloroplast. Few studies predict PSPs from sequences due to their high variety of sequences and structues. This work aims to predict and characterize PSPs by establishing the datasets of PSP and non-PSP sequences and developing prediction methods. RESULTS: A novel bioinformatics method of predicting and characterizing PSPs based on scoring card method (SCMPSP) was used. First, a dataset consisting of 649 PSPs was established by using a Gene Ontology term GO:0015979 and 649 non-PSPs from the SwissProt database with sequence identity <= 25%.- Several prediction methods are presented based on support vector machine (SVM), decision tree J48, Bayes, BLAST, and SCM. The SVM method using dipeptide features-performed well and yielded - a test accuracy of 72.31%. The SCMPSP method uses the estimated propensity scores of 400 dipeptides - as PSPs and has a test accuracy of 71.54%, which is comparable to that of the SVM method. The derived propensity scores of 20 amino acids were further used to identify informative physicochemical properties for characterizing PSPs. The analytical results reveal the following four characteristics of PSPs: 1) PSPs favour hydrophobic side chain amino acids; 2) PSPs are composed of the amino acids prone to form helices in membrane environments; 3) PSPs have low interaction with water; and 4) PSPs prefer to be composed of the amino acids of electron-reactive side chains. CONCLUSIONS: The SCMPSP method not only estimates the propensity of a sequence to be PSPs, it also discovers characteristics that further improve understanding of PSPs. The SCMPSP source code and the datasets used in this study are available at http://iclab.life.nctu.edu.tw/SCMPSP/.
Tamara Vasylenko, Yi-Fan Liou, Hong-An Chen, Phasit Charoenkwan, Hui-Ling Huang, Shinn-Ying Ho
BMC Bioinform.6
2014 SCMHBP: prediction and analysis of heme binding proteins using propensity scores of dipeptides
abstract
BACKGROUND: Heme binding proteins (HBPs) are metalloproteins that contain a heme ligand (an iron-porphyrin complex) as the prosthetic group. Several computational methods have been proposed to predict heme binding residues and thereby to understand the interactions between heme and its host proteins. However, few in silico methods for identifying HBPs have been proposed. RESULTS: This work proposes a scoring card method (SCM) based method (named SCMHBP) for predicting and analyzing HBPs from sequences. A balanced dataset of 747 HBPs (selected using a Gene Ontology term GO:0020037) and 747 non-HBPs (selected from 91,414 putative non-HBPs) with an identity of 25% was firstly established. Consequently, a set of scores that quantified the propensity of amino acids and dipeptides to be HBPs is estimated using SCM to maximize the predictive accuracy of SCMHBP. Finally, the informative physicochemical properties of 20 amino acids are identified by utilizing the estimated propensity scores to be used to categorize HBPs. The training and mean test accuracies of SCMHBP applied to three independent test datasets are 85.90% and 71.57%, respectively. SCMHBP performs well relative to comparison with such methods as support vector machine (SVM), decision tree J48, and Bayes classifiers. The putative non-HBPs with high sequence propensity scores are potential HBPs, which can be further validated by experimental confirmation. The propensity scores of individual amino acids and dipeptides are examined to elucidate the interactions between heme and its host proteins. The following characteristics of HBPs are derived from the propensity scores: 1) aromatic side chains are important to the effectiveness of specific HBP functions; 2) a hydrophobic environment is important in the interaction between heme and binding sites; and 3) the whole HBP has low flexibility whereas the heme binding residues are relatively flexible. CONCLUSIONS: SCMHBP yields knowledge that improves our understanding of HBPs rather than merely improves the prediction accuracy in predicting HBPs.
Yi-Fan Liou, Phasit Charoenkwan, Yerukala Sathipati Srinivasulu, Tamara Vasylenko, Shih-Chung Lai, Hua-Chin Lee, Yi-Hsiung Chen, Hui-Ling Huang, Shinn-Ying Ho
BMC Bioinform.9
2013 Prediction of Mouse Senescence from HE-Stain Liver Images Using an Ensemble SVM Classifier
Hui-Ling Huang, Ming-Hsin Hsu, Hua-Chin Lee, Phasit Charoenkwan, Shinn-Jang Ho, Shinn-Ying Ho
ACIIDS (2)6
2013 Designing predictors of halophilic and non-halophilic proteins using support vector machines
abstract
Finding the molecular features causes the halophilicity in the halostable organisms is helpful to understand the halophilic adaption. In this study, we proposed a prediction method for halophilic proteins by using a machine learning method. The stages of this study are six-fold. First, we establish a non-redundant dataset of the halophilic proteins, collected from NCBI, Uniprotkb and EMBL-EBI databases. The dataset consists of 245 positive and negative proteins with sequence identity <;25%. Second, the protein sequences are represented by three types of feature vector sets which include amino acid composition, dipeptide composition, and physicochemical properties. Third, we propose three classifiers based on support vector machine (SVM) to classify the halophilic proteins and non-halophilic proteins. Fourth, the independent test accuracies of the three efficient classifiers are larger than 83%. Fifth, an inheritable biobjective combinatory genetic algorithm is utilized to select a set of 11 physicochemical properties (PCPs). Sixth, these abundant amino acids, high different dipeptides (amino acid pair) and 11 informative PCPs can support to analyze the halophilic and non-halophilic proteins.
Hui-Ling Huang, Yerukala Sathipati Srinivasulu, Phasit Charoenkwan, Hua-Chin Lee, Shinn-Ying Ho
CIBCB5
2013 Predicting protein crystallization using a simple scoring card method
abstract
Many computational methods have been developed to predict protein crystallization. Most methods use amino acid and dipeptide compositions as part of the informative features. To advance the prediction accuracy, the support vector machine (SVM) based classifiers and ensemble approaches were effective and commonly-used techniques. However, these techniques suffer from the low interpretation ability of insight into crystallization. In this study, we utilize a newly-developed scoring card method (SCM) with a dipeptide composition feature to predict protein crystallization. This SCM classifier obtains prediction results 74%, 0.55 and 0.83 for accuracy, sensitivity and specificity, respectively, which is comparable to the SVM classifier using the same benchmarks. The experimental results show that the SCM classifier has advantages of simplicity, high interpretability, and high accuracy in predicting protein crystallization, compared with existing SVM-basedensemble classifiers.
Watshara Shoombuatong, Hui-Ling Huang, Jeerayut Chaijaruwanich, Phasit Charoenkwan, Hua-Chin Lee, Shinn-Ying Ho
CIBCB6
2013 HCS-Neurons: identifying phenotypic changes in multi-neuron images upon drug treatments of high-content screening
abstract
BACKGROUND: High-content screening (HCS) has become a powerful tool for drug discovery. However, the discovery of drugs targeting neurons is still hampered by the inability to accurately identify and quantify the phenotypic changes of multiple neurons in a single image (named multi-neuron image) of a high-content screen. Therefore, it is desirable to develop an automated image analysis method for analyzing multi-neuron images. RESULTS: We propose an automated analysis method with novel descriptors of neuromorphology features for analyzing HCS-based multi-neuron images, called HCS-neurons. To observe multiple phenotypic changes of neurons, we propose two kinds of descriptors which are neuron feature descriptor (NFD) of 13 neuromorphology features, e.g., neurite length, and generic feature descriptors (GFDs), e.g., Haralick texture. HCS-neurons can 1) automatically extract all quantitative phenotype features in both NFD and GFDs, 2) identify statistically significant phenotypic changes upon drug treatments using ANOVA and regression analysis, and 3) generate an accurate classifier to group neurons treated by different drug concentrations using support vector machine and an intelligent feature selection method. To evaluate HCS-neurons, we treated P19 neurons with nocodazole (a microtubule depolymerizing drug which has been shown to impair neurite development) at six concentrations ranging from 0 to 1000 ng/mL. The experimental results show that all the 13 features of NFD have statistically significant difference with respect to changes in various levels of nocodazole drug concentrations (NDC) and the phenotypic changes of neurites were consistent to the known effect of nocodazole in promoting neurite retraction. Three identified features, total neurite length, average neurite length, and average neurite area were able to achieve an independent test accuracy of 90.28% for the six-dosage classification problem. This NFD module and neuron image datasets are provided as a freely downloadable MatLab project at http://iclab.life.nctu.edu.tw/HCS-Neurons. CONCLUSIONS: Few automatic methods focus on analyzing multi-neuron images collected from HCS used in drug discovery. We provided an automatic HCS-based method for generating accurate classifiers to classify neurons based on their phenotypic changes upon drug treatments. The proposed HCS-neurons method is helpful in identifying and classifying chemical or biological molecules that alter the morphology of a group of neurons in HCS.
Phasit Charoenkwan, Eric Hwang, Robert W. Cutler, Hua-Chin Lee, Li-Wei Ko, Hui-Ling Huang, Shinn-Ying Ho
BMC Bioinform.7
2012 Motion sickness estimation system
abstract
Motion sickness occurs when the brain receives conflicting sensory information from body, inner ear and eyes [1]. In some cases, a decreased ability to actively control the body's postural motion also causes motion sickness [2][3]. Many previous studies have indicated that motion sickness had negative effect on driving performance, and sometimes lead to serious traffic accidents due to self-control ability decline. Therefore motion sickness becomes a very important issue in our daily life especially considering driving safety. There are many attempts made by researchers to realize motion sickness, and detect motion sickness in the early stage. Although many motion-sickness-related biomarkers have been identified, estimating human motion sickness level (MSL) remains a challenge in operational environment. In our past studies, we found that features in the occipital area were highly correlated with the driver's driving performance. In this study, we designed a virtual-reality (VR) based driving environment with instinct-MSL-reporting mechanism. When a subject performed a driving task, his/her brain EEG was recorded simultaneously. From those EEG data, features associated with left motor brain area, parietal brain area and occipital midline brain area which predicted MSL were extracted by an optimal classifier implemented by an inheritable bi-objective combinatorial genetic algorithm (IBCGA) with support vector machine. Unlike traditional correlation-based method, IBCGA aims to select a small set of EEG features and maximize the prediction accuracy simultaneously in BCI applications. Once the optimal feature set predicting MSL is successfully found, a driver's cognitive state can be monitored.
Chin-Teng Lin, Shu-Fang Tsai, Hua-Chin Lee, Hui-Ling Huang, Shinn-Ying Ho, Li-Wei Ko
IJCNN5
2012 Prediction and analysis of protein solubility using a novel scoring card method with dipeptide composition
abstract
BACKGROUND: Existing methods for predicting protein solubility on overexpression in Escherichia coli advance performance by using ensemble classifiers such as two-stage support vector machine (SVM) based classifiers and a number of feature types such as physicochemical properties, amino acid and dipeptide composition, accompanied with feature selection. It is desirable to develop a simple and easily interpretable method for predicting protein solubility, compared to existing complex SVM-based methods. RESULTS: This study proposes a novel scoring card method (SCM) by using dipeptide composition only to estimate solubility scores of sequences for predicting protein solubility. SCM calculates the propensities of 400 individual dipeptides to be soluble using statistic discrimination between soluble and insoluble proteins of a training data set. Consequently, the propensity scores of all dipeptides are further optimized using an intelligent genetic algorithm. The solubility score of a sequence is determined by the weighted sum of all propensity scores and dipeptide composition. To evaluate SCM by performance comparisons, four data sets with different sizes and variation degrees of experimental conditions were used. The results show that the simple method SCM with interpretable propensities of dipeptides has promising performance, compared with existing SVM-based ensemble methods with a number of feature types. Furthermore, the propensities of dipeptides and solubility scores of sequences can provide insights to protein solubility. For example, the analysis of dipeptide scores shows high propensity of α-helix structure and thermophilic proteins to be soluble. CONCLUSIONS: The propensities of individual dipeptides to be soluble are varied for proteins under altered experimental conditions. For accurately predicting protein solubility using SCM, it is better to customize the score card of dipeptide propensities by using a training data set under the same specified experimental conditions. The proposed method SCM with solubility scores and dipeptide propensities can be easily applied to the protein function prediction problems that dipeptide composition features play an important role. AVAILABILITY: The used datasets, source codes of SCM, and supplementary files are available at http://iclab.life.nctu.edu.tw/SCM/.
Hui-Ling Huang, Phasit Charoenkwan, Te-Fen Kao, Hua-Chin Lee, Fang-Lin Chang, Wen-Lin Huang, Shinn-Jang Ho, Li-Sun Shu, Shinn-Ying Ho
BMC Bioinform.10
2011 NeurphologyJ: an automatic neuronal morphology quantification method and its application in pharmacological discovery
abstract
BACKGROUND: Automatic quantification of neuronal morphology from images of fluorescence microscopy plays an increasingly important role in high-content screenings. However, there exist very few freeware tools and methods which provide automatic neuronal morphology quantification for pharmacological discovery. RESULTS: This study proposes an effective quantification method, called NeurphologyJ, capable of automatically quantifying neuronal morphologies such as soma number and size, neurite length, and neurite branching complexity (which is highly related to the numbers of attachment points and ending points). NeurphologyJ is implemented as a plugin to ImageJ, an open-source Java-based image processing and analysis platform. The high performance of NeurphologyJ arises mainly from an elegant image enhancement method. Consequently, some morphology operations of image processing can be efficiently applied. We evaluated NeurphologyJ by comparing it with both the computer-aided manual tracing method NeuronJ and an existing ImageJ-based plugin method NeuriteTracer. Our results reveal that NeurphologyJ is comparable to NeuronJ, that the coefficient correlation between the estimated neurite lengths is as high as 0.992. NeurphologyJ can accurately measure neurite length, soma number, neurite attachment points, and neurite ending points from a single image. Furthermore, the quantification result of nocodazole perturbation is consistent with its known inhibitory effect on neurite outgrowth. We were also able to calculate the IC50 of nocodazole using NeurphologyJ. This reveals that NeurphologyJ is effective enough to be utilized in applications of pharmacological discoveries. CONCLUSIONS: This study proposes an automatic and fast neuronal quantification method NeurphologyJ. The ImageJ plugin with supports of batch processing is easily customized for dealing with high-content screening applications. The source codes of NeurphologyJ (interactive and high-throughput versions) and the images used for testing are freely available (see Availability).
Shinn-Ying Ho, Chih-Yuan Chao, Hui-Ling Huang, Tzai-Wen Chiu, Phasit Charoenkwan, Eric Hwang
BMC Bioinform.1
2011 Predicting and analyzing DNA-binding domains using a systematic approach to identifying a set of informative physicochemical and biochemical properties
abstract
BACKGROUND: Existing methods of predicting DNA-binding proteins used valuable features of physicochemical properties to design support vector machine (SVM) based classifiers. Generally, selection of physicochemical properties and determination of their corresponding feature vectors rely mainly on known properties of binding mechanism and experience of designers. However, there exists a troublesome problem for designers that some different physicochemical properties have similar vectors of representing 20 amino acids and some closely related physicochemical properties have dissimilar vectors. RESULTS: This study proposes a systematic approach (named Auto-IDPCPs) to automatically identify a set of physicochemical and biochemical properties in the AAindex database to design SVM-based classifiers for predicting and analyzing DNA-binding domains/proteins. Auto-IDPCPs consists of 1) clustering 531 amino acid indices in AAindex into 20 clusters using a fuzzy c-means algorithm, 2) utilizing an efficient genetic algorithm based optimization method IBCGA to select an informative feature set of size m to represent sequences, and 3) analyzing the selected features to identify related physicochemical properties which may affect the binding mechanism of DNA-binding domains/proteins. The proposed Auto-IDPCPs identified m = 22 features of properties belonging to five clusters for predicting DNA-binding domains with a five-fold cross-validation accuracy of 87.12%, which is promising compared with the accuracy of 86.62% of the existing method PSSM-400. For predicting DNA-binding sequences, the accuracy of 75.50% was obtained using m = 28 features, where PSSM-400 has an accuracy of 74.22%. Auto-IDPCPs and PSSM-400 have accuracies of 80.73% and 82.81%, respectively, applied to an independent test data set of DNA-binding domains. Some typical physicochemical properties discovered are hydrophobicity, secondary structure, charge, solvent accessibility, polarity, flexibility, normalized Van Der Waals volume, pK (pK-C, pK-N, pK-COOH and pK-a(RCOOH)), etc. CONCLUSIONS: The proposed approach Auto-IDPCPs would help designers to investigate informative physicochemical and biochemical properties by considering both prediction accuracy and analysis of binding mechanism simultaneously. The approach Auto-IDPCPs can be also applicable to predict and analyze other protein functions from sequences.
Hui-Ling Huang, I-Che Lin, Yi-Fan Liou, Chia-Ta Tsai, Kai-Ti Hsu, Wen-Lin Huang, Shinn-Jang Ho, Shinn-Ying Ho
BMC Bioinform.8
2011 POPISK: T-cell reactivity prediction using support vector machines and string kernels
abstract
BACKGROUND: Accurate prediction of peptide immunogenicity and characterization of relation between peptide sequences and peptide immunogenicity will be greatly helpful for vaccine designs and understanding of the immune system. In contrast to the prediction of antigen processing and presentation pathway, the prediction of subsequent T-cell reactivity is a much harder topic. Previous studies of identifying T-cell receptor (TCR) recognition positions were based on small-scale analyses using only a few peptides and concluded different recognition positions such as positions 4, 6 and 8 of peptides with length 9. Large-scale analyses are necessary to better characterize the effect of peptide sequence variations on T-cell reactivity and design predictors of a peptide's T-cell reactivity (and thus immunogenicity). The identification and characterization of important positions influencing T-cell reactivity will provide insights into the underlying mechanism of immunogenicity. RESULTS: This work establishes a large dataset by collecting immunogenicity data from three major immunology databases. In order to consider the effect of MHC restriction, peptides are classified by their associated MHC alleles. Subsequently, a computational method (named POPISK) using support vector machine with a weighted degree string kernel is proposed to predict T-cell reactivity and identify important recognition positions. POPISK yields a mean 10-fold cross-validation accuracy of 68% in predicting T-cell reactivity of HLA-A2-binding peptides. POPISK is capable of predicting immunogenicity with scores that can also correctly predict the change in T-cell reactivity related to point mutations in epitopes reported in previous studies using crystal structures. Thorough analyses of the prediction results identify the important positions 4, 6, 8 and 9, and yield insights into the molecular basis for TCR recognition. Finally, we relate this finding to physicochemical properties and structural features of the MHC-peptide-TCR interaction. CONCLUSIONS: A computational method POPISK is proposed to predict immunogenicity with scores which are useful for predicting immunogenicity changes made by single-residue modifications. The web server of POPISK is freely available at http://iclab.life.nctu.edu.tw/POPISK.
Chun-Wei Tung, Matthias Ziehm, Andreas Kämper, Oliver Kohlbacher, Shinn-Ying Ho
BMC Bioinform.5
2009 Protein subcellular localization prediction of eukaryotes using a knowledge-based approach
abstract
BACKGROUND: The study of protein subcellular localization (PSL) is important for elucidating protein functions involved in various cellular processes. However, determining the localization sites of a protein through wet-lab experiments can be time-consuming and labor-intensive. Thus, computational approaches become highly desirable. Most of the PSL prediction systems are established for single-localized proteins. However, a significant number of eukaryotic proteins are known to be localized into multiple subcellular organelles. Many studies have shown that proteins may simultaneously locate or move between different cellular compartments and be involved in different biological processes with different roles. RESULTS: In this study, we propose a knowledge based method, called KnowPredsite, to predict the localization site(s) of both single-localized and multi-localized proteins. Based on the local similarity, we can identify the "related sequences" for prediction. We construct a knowledge base to record the possible sequence variations for protein sequences. When predicting the localization annotation of a query protein, we search against the knowledge base and used a scoring mechanism to determine the predicted sites. We downloaded the dataset from ngLOC, which consisted of ten distinct subcellular organelles from 1923 species, and performed ten-fold cross validation experiments to evaluate KnowPred site's performance. The experiment results show that KnowPred site achieves higher prediction accuracy than ngLOC and Blast-hit method. For single-localized proteins, the overall accuracy of KnowPred site is 91.7%. For multi-localized proteins, the overall accuracy of KnowPred site is 72.1%, which is significantly higher than that of ngLOC by 12.4%. Notably, half of the proteins in the dataset that cannot find any Blast hit sequence above a specified threshold can still be correctly predicted by KnowPred site. CONCLUSION: KnowPred site demonstrates the power of identifying related sequences in the knowledge base. The experiment results show that even though the sequence similarity is low, the local similarity is effective for prediction. Experiment results show that KnowPred site is a highly accurate prediction method for both single- and multi-localized proteins. It is worth-mentioning the prediction process of KnowPred site is transparent and biologically interpretable and it shows a set of template sequences to generate the prediction result. The KnowPred site prediction server is available at http://bio-cluster.iis.sinica.edu.tw/kbloc/.
Hsin-Nan Lin, Ching-Tai Chen, Ting-Yi Sung, Shinn-Ying Ho, Wen-Lian Hsu
BMC Bioinform.4
2008 Inferring S-system models of genetic networks from a time-series real data set of gene expression profiles
abstract
It is desirable to infer cellular dynamic regulation networks from gene expression profiles to discover more delicate and substantial functions in molecular biology, biochemistry, bioengineering, and pharmaceutics. The S-system model is suitable to characterize biochemical network systems and capable of analyzing the regulatory system dynamics. To cope with the problem ldquomultiplicity of solutionsrdquo, a sufficient amount of data sets of time-series gene expression profiles were often used. An efficient newly-developed method iTEA was proposed to effectively obtain S-system models from a large number (e.g., 15) of simulated data sets with/without noise. In this study, we propose an extended optimization method (named iTEAP) based on iTEA to infer the S-system models of genetic networks from a time-series real data set of gene expression profiles (using SOS DNA microarray data inE.colias an example). The algorithm iTEAP generated additionally multiple data sets of gene expression profiles by perturbing the given data set. The results reveal that 1) iTEAP can obtain S-system models with high-quality profiles to best fit the observed profiles; 2) the performance of using multiple data sets is better than that of using a single data set in terms of solution quality, and 3) the effectiveness of iTEAP using a single data set is close to that of iTEA using two real data sets.
Hui-Ling Huang, Kuan-Wei Chen, Shinn-Jang Ho, Shinn-Ying Ho
IEEE Congress on Evolutionary Computation4
2008 ProLoc-rGO: Using rule-based knowledge with Gene Ontology terms for prediction of protein subnuclear localization
abstract
Gene ontology (GO) annotation is a controlled vocabulary of terms and phrases describing the function of genes and gene products, which has been succeeded in predicting subcellular and subnuclear localization. Generally, each gene product is annotated by very few GO terms from more than 25,000 annotations available at present. How to represent a protein sequence using GO terms as features plays an important role in designing prediction systems for protein subnuclear localization. Our previous work ProLoc-GO can select a small number m out of a large number n GO terms, where m Lt n. However, its off-line time for training is large up to several days even though running on high speedily PC clusters. Therefore, this study proposes an efficient system (ProLoc-rGO) by using the decision tree method to speedily mine m informative GO terms and acquire interpretable rule-based knowledge for predicting subnuclear localization. The ProLoc-rGO performing on SNL9_80 (714 proteins in nine compartments with <80 identity) can mine m=17 informative GO terms, 17 interpretable rules and yield training and test accuracies of 84.9% and 78.2%. For comparison, an accuracy 82.6% (Matthews correlation coefficient (MCC) = 0.711) for ProLoc-rGO performed on SNL9_80 (714 proteins in nine compartments with <80 identity) is obtained, which is better than 67.4% (MCC = 0.50) for Nuc-PLoc that fuses the pseudo-amino acid composition of a protein and its position-specific scoring matrix.
Wen-Lin Huang, Chun-Wei Tung, Shih-Wen Ho, Shinn-Ying Ho
CIBCB4
2008 ProLoc-GO: Utilizing informative Gene Ontology terms for sequence-based prediction of protein subcellular localization
abstract
BACKGROUND: Gene Ontology (GO) annotation, which describes the function of genes and gene products across species, has recently been used to predict protein subcellular and subnuclear localization. Existing GO-based prediction methods for protein subcellular localization use the known accession numbers of query proteins to obtain their annotated GO terms. An accurate prediction method for predicting subcellular localization of novel proteins without known accession numbers, using only the input sequence, is worth developing. RESULTS: This study proposes an efficient sequence-based method (named ProLoc-GO) by mining informative GO terms for predicting protein subcellular localization. For each protein, BLAST is used to obtain a homology with a known accession number to the protein for retrieving the GO annotation. A large number n of all annotated GO terms that have ever appeared are then obtained from a large set of training proteins. A novel genetic algorithm based method (named GOmining) combined with a classifier of support vector machine (SVM) is proposed to simultaneously identify a small number m out of the n GO terms as input features to SVM, where m <<n. The m informative GO terms contain the essential GO terms annotating subcellular compartments such as GO:0005634 (Nucleus), GO:0005737 (Cytoplasm) and GO:0005856 (Cytoskeleton). Two existing data sets SCL12 (human protein with 12 locations) and SCL16 (Eukaryotic proteins with 16 locations) with <25% sequence identity are used to evaluate ProLoc-GO which has been implemented by using a single SVM classifier with the m = 44 and m = 60 informative GO terms, respectively. ProLoc-GO using input sequences yields test accuracies of 88.1% and 83.3% for SCL12 and SCL16, respectively, which are significantly better than the SVM-based methods, which achieve < 35% test accuracies using amino acid composition (AAC) with acid pairs and AAC with dipedtide composition. For comparison, ProLoc-GO using known accession numbers of query proteins yields test accuracies of 90.6% and 85.7%, which is also better than Hum-PLoc (85.0%) and Euk-OET-PLoc (83.7%) using ensemble classifiers with hybridization of GO terms and amphiphilic pseudo amino acid composition for SCL12 and SCL16, respectively. CONCLUSION: The growth of Gene Ontology in size and popularity has increased the effectiveness of GO-based features. GOmining can serve as a tool for selecting informative GO terms in solving sequence-based prediction problems. The prediction system using ProLoc-GO with input sequences of query proteins for protein subcellular localization has been implemented (see Availability).
Wen-Lin Huang, Chun-Wei Tung, Shih-Wen Ho, Shiow-Fen Hwang, Shinn-Ying Ho
BMC Bioinform.5
2008 Computational identification of ubiquitylation sites from protein sequences
abstract
BACKGROUND: Ubiquitylation plays an important role in regulating protein functions. Recently, experimental methods were developed toward effective identification of ubiquitylation sites. To efficiently explore more undiscovered ubiquitylation sites, this study aims to develop an accurate sequence-based prediction method to identify promising ubiquitylation sites. RESULTS: We established an ubiquitylation dataset consisting of 157 ubiquitylation sites and 3676 putative non-ubiquitylation sites extracted from 105 proteins in the UbiProt database. This study first evaluates promising sequence-based features and classifiers for the prediction of ubiquitylation sites by assessing three kinds of features (amino acid identity, evolutionary information, and physicochemical property) and three classifiers (support vector machine, k-nearest neighbor, and NaïveBayes). Results show that the set of used 531 physicochemical properties and support vector machine (SVM) are the best kind of features and classifier respectively that their combination has a prediction accuracy of 72.19% using leave-one-out cross-validation.Consequently, an informative physicochemical property mining algorithm (IPMA) is proposed to select an informative subset of 531 physicochemical properties. A prediction system UbiPred was implemented by using an SVM with the feature set of 31 informative physicochemical properties selected by IPMA, which can improve the accuracy from 72.19% to 84.44%. To further analyze the informative physicochemical properties, a decision tree method C5.0 was used to acquire if-then rule-based knowledge of predicting ubiquitylation sites. UbiPred can screen promising ubiquitylation sites from putative non-ubiquitylation sites using prediction scores. By applying UbiPred, 23 promising ubiquitylation sites were identified from an independent dataset of 3424 putative non-ubiquitylation sites, which were also validated by using the obtained prediction rules. CONCLUSION: We have proposed an algorithm IPMA for mining informative physicochemical properties from protein sequences to build an SVM-based prediction system UbiPred. UbiPred can predict ubiquitylation sites accompanied with a prediction score each to help biologists in identifying promising sites for experimental verification. UbiPred has been implemented as a web server and is available at http://iclab.life.nctu.edu.tw/ubipred.
Chun-Wei Tung, Shinn-Ying Ho
BMC Bioinform.2
2008 OPSO: Orthogonal Particle Swarm Optimization and Its Application to Task Assignment Problems
abstract
This paper proposes a novel variant of particle swarm optimization (PSO), named orthogonal PSO (OPSO), for solving intractable large parameter optimization problems. The standard version of PSO is associated with the lack of a mechanism responsible for the process of high-dimensional vector spaces. The high performance of OPSO arises mainly from a novel move behavior using an intelligent move mechanism (IMM) which applies orthogonal experimental design to adjust a velocity for each particle by using a systematic reasoning method instead of the conventional generate-and-go method. The IMM uses a divide-and- conquer approach to cope with the curse of dimensionality in determining the next move of particles. It is shown empirically that the OPSO performs well in solving parametric benchmark functions and a task assignment problem which is NP-complete compared with the standard PSO with the conventional move behavior. The OPSO with IMM is more specialized than the PSO and performs well on large-scale parameter optimization problems with few interactions between variables.
Shinn-Ying Ho, Hung-Sui Lin, Weei-Hurng Liauh, Shinn-Jang Ho
IEEE Trans. Syst. Man Cybern. Part A1
2008 A Novel Intelligent Multiobjective Simulated Annealing Algorithm for Designing Robust PID Controllers
abstract
This paper proposes an intelligent multiobjective simulated annealing algorithm (IMOSA) and its application to an optimal proportional integral derivative (PID) controller design problem. A well-designed PID-type controller should satisfy the following objectives: 1) disturbance attenuation; 2) robust stability; and 3) accurate setpoint tracking. The optimal PID controller design problem is a large-scale multiobjective optimization problem characterized by the following: 1) nonlinear multimodal search space; 2) large-scale search space; 3) three tight constraints; 4) multiple objectives; and 5) expensive objective function evaluations. In contrast to existing multiobjective algorithms of simulated annealing, the high performance in IMOSA arises mainly from a novel multiobjective generation mechanism using a Pareto-based scoring function without using heuristics. The multiobjective generation mechanism operates on a high-score nondominated solution using a systematic reasoning method based on an orthogonal experimental design, which exploits its neighborhood to economically generate a set of well-distributed nondominated solutions by considering individual and overall objectives. IMOSA is evaluated by using a practical design example of a super-maneuverable fighter aircraft system. An efficient existing multiobjective algorithm, the improved strength Pareto evolutionary algorithm, is also applied to the same example for comparison. Simulation results demonstrate high performance of the IMOSA-based method in designing robust PID controllers.
Ming-Hao Hung, Li-Sun Shu, Shinn-Jang Ho, Shiow-Fen Hwang, Shinn-Ying Ho
IEEE Trans. Syst. Man Cybern. Part A5
2007 Boosting Evolutionary Support Vector Machine for Designing Tumor Classifiers from Microarray Data
abstract
Since there are multiple sets of relevant genes having the same high accuracy in fitting training data called model uncertainty, to identify a small set of informative genes from microarray data for designing an accurate tumor classifier for unknown samples is intractable. Support vector machine (SVM), a supervised machine learning technique, is one of the methods successfully applied to cancer diagnosis problems. This study proposes an SVM-based classifier with automatic feature selection associated with a boosting strategy. The proposed boosting evolutionary support vector machine (named BESVM) hybridizes the advantages of SVM, boosting using a majority-voting ensemble and an intelligent genetic algorithm for gene selection. The merits of the BESVM-based classifier are threefold: 1) a small set of used genes, 2) accurate test classification using leave-one-out cross-validation, and 3) robust performance by avoiding overfitting training data. Five benchmark datasets were used to evaluate the BESVM-based classifier. Simulation results reveal that BESVM performs well having a mean accuracy 94.26% using only 10.1 genes averagely, compared with the existing SVM and non-SVM based classifiers
Hui-Ling Huang, Yi-Hsiung Chen, Dwight D. Koeberl, Shinn-Ying Ho
CIBCB4
2007 iPTREE-STAB: interpretable decision tree based method for predicting protein stability changes upon mutations
abstract
UNLABELLED: We have developed a web server, iPTREE-STAB for discriminating the stability of proteins (stabilizing or destabilizing) and predicting their stability changes (delta deltaG) upon single amino acid substitutions from amino acid sequence. The discrimination and prediction are mainly based on decision tree coupled with adaptive boosting algorithm, and classification and regression tree, respectively, using three neighboring residues of the mutant site along N- and C-terminals. Our method showed an accuracy of 82% for discriminating the stabilizing and destabilizing mutants, and a correlation of 0.70 for predicting protein stability changes upon mutations. AVAILABILITY: http://bioinformatics.myweb.hinet.net/iptree.htm. SUPPLEMENTARY INFORMATION: Dataset and other details are given.
Liang-Tsung Huang, M. Michael Gromiha, Shinn-Ying Ho
Bioinform.3
2007 POPI: predicting immunogenicity of MHC class I binding peptides by mining informative physicochemical properties
abstract
MOTIVATION: Both modeling of antigen-processing pathway including major histocompatibility complex (MHC) binding and immunogenicity prediction of those MHC-binding peptides are essential to develop a computer-aided system of peptide-based vaccine design that is one goal of immunoinformatics. Numerous studies have dealt with modeling the immunogenic pathway but not the intractable problem of immunogenicity prediction due to complex effects of many intrinsic and extrinsic factors. Moderate affinity of the MHC-peptide complex is essential to induce immune responses, but the relationship between the affinity and peptide immunogenicity is too weak to use for predicting immunogenicity. This study focuses on mining informative physicochemical properties from known experimental immunogenicity data to understand immune responses and predict immunogenicity of MHC-binding peptides accurately. RESULTS: This study proposes a computational method to mine a feature set of informative physicochemical properties from MHC class I binding peptides to design a support vector machine (SVM) based system (named POPI) for the prediction of peptide immunogenicity. High performance of POPI arises mainly from an inheritable bi-objective genetic algorithm, which aims to automatically determine the best number m out of 531 physicochemical properties, identify these m properties and tune SVM parameters simultaneously. The dataset consisting of 428 human MHC class I binding peptides belonging to four classes of immunogenicity was established from MHCPEP, a database of MHC-binding peptides (Brusic et al., 1998). POPI, utilizing the m = 23 selected properties, performs well with the accuracy of 64.72% using leave-one-out cross-validation, compared with two sequence alignment-based prediction methods ALIGN (54.91%) and PSI-BLAST (53.23%). POPI is the first computational system for prediction of peptide immunogenicity based on physicochemical properties. AVAILABILITY: A web server for prediction of peptide immunogenicity (POPI) and the used dataset of MHC class I binding peptides (PEPMHCI) are available at http://iclab.life.nctu.edu.tw/POPI
Chun-Wei Tung, Shinn-Ying Ho
Bioinform.2
2007 An Intelligent Two-Stage Evolutionary Algorithm for Dynamic Pathway Identification From Gene Expression Profiles
abstract
From gene expression profiles, it is desirable to rebuild cellular dynamic regulation networks to discover more delicate and substantial functions in molecular biology, biochemistry, bioengineering and pharmaceutics. S-system model is suitable to characterize biochemical network systems and capable to analyze the regulatory system dynamics. However, inference of an S-system model of N-gene genetic networks has 2N(N+1) parameters in a set of non-linear differential equations to be optimized. This paper proposes an intelligent two-stage evolutionary algorithm (iTEA) to efficiently infer the S-system models of genetic networks from time-series data of gene expression. To cope with curse of dimensionality, the proposed algorithm consists of two stages where each uses a divide-and-conquer strategy. The optimization problem is first decomposed into N subproblems having 2(N+1) parameters each. At the first stage, each subproblem is solved using a novel intelligent genetic algorithm (IGA) with intelligent crossover based on orthogonal experimental design (OED). At the second stage, the obtained N solutions to the N subproblems are combined and refined using an OED-based simulated annealing algorithm for handling noisy gene expression profiles. The effectiveness of iTEA is evaluated using simulated expression patterns with and without noise running on a single-processor PC. It is shown that 1) IGA is efficient enough to solve subproblems; 2) IGA is significantly superior to the existing method SPXGA; and 3) iTEA performs well in inferring S-system models for dynamic pathway identification.
Shinn-Ying Ho, Chih-Hung Hsieh, Fu-Chieh Yu, Hui-Ling Huang
IEEE ACM Trans. Comput. Biol. Bioinform.1
2006 Scoring Method for Tumor Prediction from Microarray Data Using an Evolutionary Fuzzy Classifier
Shinn-Ying Ho, Chih-Hung Hsieh, Kuan-Wei Chen, Hui-Ling Huang, Hung-Ming Chen, Shinn-Jang Ho
PAKDD1
2006 Intelligent Particle Swarm Optimization in Multi-objective Problems
Shinn-Jang Ho, Wen-Yuan Ku, Jun-Wun Jou, Ming-Hao Hung, Shinn-Ying Ho
PAKDD5
2006 Optimizing fuzzy neural networks for tuning PID controllers using an orthogonal simulated annealing algorithm OSA
abstract
In this paper, we formulate an optimization problem of establishing a fuzzy neural network model (FNNM) for efficiently tuning proportional-integral-derivative (PID) controllers of various test plants with under-damped responses using a large number P of training plants such that the mean tracking error J of the obtained P control systems is minimized. The FNNM consists of four fuzzy neural networks (FNNs) where each FNN models one of controller parameters (K, T/sub i/, T/sub d/, and b) of PID controllers. An existing indirect, two-stage approach used a dominant pole assignment method with P=198 to find the corresponding PID controllers. Consequently, an adaptive neuro-fuzzy inference system (ANFIS) is used to independently train the four individual FNNs using input the selected 176 of the 198 PID controllers that 22 controllers with parameters having large variation are abandoned. The innovation of the proposed approach is to directly and simultaneously optimize the four FNNs by using a novel orthogonal simulated annealing algorithm (OSA). High performance of the OSA-based approach arises from that OSA can effectively optimize lots of parameters of the FNNM to minimize J. It is shown that the OSA-based FNNM with P=176 can improve the ANFIS-based FNNM in averagely decreasing 13.08% error J and 88.07% tracking error of the 22 test plants by refining the solution of the ANFIS-based method. Furthermore, the OSA-based FNNMs using P=198 and 396 from an extensive tuning domain have similar good performance with that using P=176 in terms of J.
Shinn-Jang Ho, Li-Sun Shu, Shinn-Ying Ho
IEEE Trans. Fuzzy Syst.3
2005 Evolutionary divide-and-conquer approach to inferring S-system models of genetic networks
abstract
This paper proposes an efficient evolutionary divide-and-conquer approach (EDACA) to inferring S-system models of genetic networks from time-series data of gene expression. Inference of an S-system model has 2N(N+1) parameters to be optimized, where N is the number of genes in a genetic network. To cope with higher dimensionality, the proposed approach consists of two stages where each uses a divide-and-conquer strategy. The optimization problem is first decomposed into N subproblems having 2(N+1) parameters each. At the first stage, each subproblem is solved using a novel intelligent genetic algorithm (IGA) with intelligent crossover based on orthogonal experimental design (OED). The intelligent crossover divides two parents into n pairs of parameter groups, economically identifies the potentially better one of two groups of each pair, and systematically obtains a potentially good approximation to the best one of all 2/sup n/ combinations using at most 2n function evaluations. At the second stage, the obtained N solutions to the N subproblems are combined and refined using an OED-based simulated annealing algorithm (OSA) for handling noisy gene expression data. The effectiveness of EDACA is evaluated using simulated expression patterns with/without noise running on a single-CPU PC. It is shown that: 1) IGA is efficient enough to solve subproblems; 2) IGA is significantly superior to the existing method of using GA with simplex crossover; and 3) EDACA performs well in inferring S-system models of genetic networks from small-noise gene expression data.
Shinn-Ying Ho, Chih-Hung Hsieh, Fu-Chieh Yu, Hui-Ling Huang
Congress on Evolutionary Computation1
2005 Efficient gene selection for classification of microarray data
abstract
Microarray is a useful technique for measuring expression data of thousands of genes simultaneously. One of challenges in classification of microarray data is to select a minimal number of relevant genes which can maximize classification accuracy. Many gene selection methods as well as their corresponding classifiers have been proposed. One of existing analysis methods is the hybrid approach based on genetic algorithm and maximum likelihood classification (GA/MLHD). In this paper, an intelligent genetic algorithm (IGA) using control genes and an improved fitness function is proposed to determine the minimal number of relevant genes and identify these genes, while maximizing classification accuracy simultaneously. The experimental results show that our approach is superior to the existing method GA/MLHD in terms of the number of selected genes, classification accuracy, and robustness of selected genes and accuracy, especially for the datasets which have numerous categories and a large number of testing genes inside.
Shinn-Ying Ho, Chong-Cheng Lee, Hung-Ming Chen, Hui-Ling Huang
Congress on Evolutionary Computation1
2005 Flexible protein-ligand docking using particle swarm optimization
abstract
Many protein-ligand docking problems attempt to predict the bound conformations of two interacting molecules. Consequently, the docking problem requires a powerful search technique to explore the translations, orientations, and each torsion until an ideal site has been found. Therefore, protein-ligand docking can be formulated as a parameter optimization problem. However, highly flexible ligands have a lot of torsions. Therefore, the optimization problem of highly flexible docking would become more difficult due to the increment of parameter number and interactions among these parameters. We proposed a novel method SODOCK based on particle swarm optimization (PSO) for solving flexible protein-ligand docking problems. PSO has significant effect on the optimization of parameters with strong interactions. A commonly used efficient local search is incorporated into SODOCK to improve the efficiency and robustness of PSO. SODOCK is efficient for both types of ligands with small and large numbers of torsions. It is shown by computer simulation that SODOCK performs well in obtaining accurate conformations, compared with some of state-of-the-art methods. Moreover, it is also shown that SODOCK is superior to AutoDock using the same energy function in AutoDock 3.05 in terms of convergence speed, robustness, and docking energy, especially for highly flexible docking problems
Bo-Fu Liu, Hung-Ming Chen, Hui-Ling Huang, Shiow-Fen Hwang, Shinn-Ying Ho
Congress on Evolutionary Computation5
2005 Interpretable Prediction of Protein Stability Changes upon Mutation by Using Decision Tree
Larry Huang, Wen-Lin Huang, Shinn-Ying Ho, Shiow-Fen Hwang
CIBCB3
2005 Quality-time analysis of multi-objective evolutionary algorithms
abstract
A quality-time analysis of multi-objective evolutionary algorithms (MOEAs) based on schema theorem and building blocks hypothesis is developed. A bicriteria OneMax problem, a hypothesis of niche and species, and a definition of dissimilar schemata are introduced for the analysis. In this paper, the convergence time, the first and last hitting time models are constructed for analyzing the performance of MOEAs. Population sizing model is constructed for determining appropriate population sizes. The models are verified using the bicriteria OneMax problem. The theoretical results indicate how the convergence time and population size of a MOEA scale up with the problem size, the dissimilarity of Pareto-optimal solutions, and the number of Pareto-optimal solutions of a multi-objective optimization problem.
Jian-Hung Chen, Shinn-Ying Ho, David E. Goldberg
GECCO2
2005 MeSwarm: memetic particle swarm optimization
abstract
In this paper, a novel variant of particle swarm optimization (PSO), named memetic particle swarm optimization algorithm (MeSwarm), is proposed for tackling the overshooting problem in the motion behavior of PSO. The overshooting problem is a phenomenon in PSO due to the velocity update mechanism of PSO. While the overshooting problem occurs, particles may be led to wrong or opposite directions against the direction to the global optimum. As a result, MeSwarm integrates the standard PSO with the Solis and Wets local search strategy to avoid the overshooting problem and that is based on the recent probability of success to efficiently generate a new candidate solution around the current particle. Thus, six test functions and a real-world optimization problem, the flexible protein-ligand docking problem are used to validate the performance of MeSwarm. The experimental results indicate that MeSwarm outperforms the standard PSO and several evolutionary algorithms in terms of solution quality.
Bo-Fu Liu, Hung-Ming Chen, Jian-Hung Chen, Shiow-Fen Hwang, Shinn-Ying Ho
GECCO5
2005 Design of nearest neighbor classifiers: multi-objective approach
Jian-Hung Chen, Hung-Ming Chen, Shinn-Ying Ho
Int. J. Approx. Reason.3
2004 A Novel Multi-objective Orthogonal Simulated Annealing Algorithm for Solving Multi-objective Optimization Problems with a Large Number of Parameters
Li-Sun Shu, Shinn-Jang Ho, Shinn-Ying Ho, Jian-Hung Chen, Ming-Hao Hung
GECCO (1)3
2004 Design of Nearest Neighbor Classifiers Using an Intelligent Multi-objective Evolutionary Algorithm
Jian-Hung Chen, Hung-Ming Chen, Shinn-Ying Ho
PRICAI3
2004 Intelligent evolutionary algorithms for large parameter optimization problems
abstract
This work proposes two intelligent evolutionary algorithms IEA and IMOEA using a novel intelligent gene collector (IGC) to solve single and multiobjective large parameter optimization problems, respectively. IGC is the main phase in an intelligent recombination operator of IEA and IMOEA. Based on orthogonal experimental design, IGC uses a divide-and-conquer approach, which consists of adaptively dividing two individuals of parents into N pairs of gene segments, economically identifying the potentially better one of two gene segments of each pair, and systematically obtaining a potentially good approximation to the best one of all combinations using at most 2N fitness evaluations. IMOEA utilizes a novel generalized Pareto-based scale-independent fitness function for efficiently finding a set of Pareto-optimal solutions to a multiobjective optimization problem. The advantages of IEA and IMOEA are their simplicity, efficiency, and flexibility. It is shown empirically that IEA and IMOEA have high performance in solving benchmark functions comprising many parameters, as compared with some existing EAs.
Shinn-Ying Ho, Li-Sun Shu, Jian-Hung Chen
IEEE Trans. Evol. Comput.1
2004 Inheritable genetic algorithm for biobjective 0/1 combinatorial optimization problems and its applications
abstract
In this paper, we formulate a special type of multiobjective optimization problems, named biobjective 0/1 combinatorial optimization problem BOCOP, and propose an inheritable genetic algorithm IGA with orthogonal array crossover (OAX) to efficiently find a complete set of nondominated solutions to BOCOP. BOCOP with n binary variables has two incommensurable and often competing objectives: minimizing the sum r of values of all binary variables and optimizing the system performance. BOCOP is NP-hard having a finite number C(n, r) of feasible solutions for a limited number r. The merits of IGA are threefold as follows: 1) OAX with the systematic reasoning ability based on orthogonal experimental design can efficiently explore the search space of C(n, r); 2) IGA can efficiently search the space of C(n, r+/-1) by inheriting a good solution in the space of C(n, r); and 3) The single-objective IGA can economically obtain a complete set of high-quality nondominated solutions in a single run. Two applications of BOCOP are used to illustrate the effectiveness of the proposed algorithm: polygonal approximation problem (PAP) and the problem of editing a minimum reference set for nearest neighbor classification (MRSP). It is shown empirically that IGA is efficient in finding complete sets of nondominated solutions to PAP and MRSP, compared with some existing methods.
Shinn-Ying Ho, Jian-Hung Chen, Meng-Hsun Huang
IEEE Trans. Syst. Man Cybern. Part B1
2004 Design of accurate classifiers with a compact fuzzy-rule base using an evolutionary scatter partition of feature space
abstract
An evolutionary approach to designing accurate classifiers with a compact fuzzy-rule base using a scatter partition of feature space is proposed, in which all the elements of the fuzzy classifier design problem have been moved in parameters of a complex optimization problem. An intelligent genetic algorithm (IGA) is used to effectively solve the design problem of fuzzy classifiers with many tuning parameters. The merits of the proposed method are threefold: 1) the proposed method has high search ability to efficiently find fuzzy rule-based systems with high fitness values, 2) obtained fuzzy rules have high interpretability, and 3) obtained compact classifiers have high classification accuracy on unseen test patterns. The sensitivity of control parameters of the proposed method is empirically analyzed to show the robustness of the IGA-based method. The performance comparison and statistical analysis of experimental results using ten-fold cross validation show that the IGA-based method without heuristics is efficient in designing accurate and compact fuzzy classifiers using 11 well-known data sets with numerical attribute values.
Shinn-Ying Ho, Hung-Ming Chen, Shinn-Jang Ho, Tai-Kang Chen
IEEE Trans. Syst. Man Cybern. Part B1
2004 OSA: orthogonal simulated annealing algorithm and its application to designing mixed H2/H∞ optimal controllers
abstract
This paper proposes a novel orthogonal simulated annealing (OSA) algorithm for solving intractable large-scale engineering problems and its application to designing mixed H/sub 2//H/sub /spl infin// optimal structure-specified controllers with robust stability and disturbance attenuation. High performance of OSA arises mainly from an intelligent generation mechanism (IGM), which applies orthogonal experimental design to speed up the search. IGM can efficiently generate a good candidate solution for next move of OSA by using a systematic reasoning method. It is difficult for existing H/sub /spl infin//- and genetic algorithm (GA)-based methods to economically obtain an accurate solution to the design problem of multiple-input, multiple-output (MIMO) optimal control systems. The high performance and validity of OSA are demonstrated by parametric optimization functions and a MIMO super maneuverable F18/HARV fighter aircraft system with a proportional-integral-derivative (PID)-type controller. It is shown empirically that OSA performs well for parametric optimization functions and the performance of the OSA-based method without prior domain knowledge is superior to those of existing H/sub /spl infin//- and GA-based methods for designing MIMO optimal controllers.
Shin-Hang Ho, Shinn-Ying Ho, Li-Sun Shu
IEEE Trans. Syst. Man Cybern. Part A2
2004 An orthogonal simulated annealing algorithm for large floorplanning problems
abstract
The conventional simulated annealing with some random generation mechanism using the sequence-pair topological representation in block placement and floorplanning is effective for a very small number of modules (40-50). This paper proposes an orthogonal simulated annealing algorithm (OSA) with an efficient generation mechanism (EGM) for solving large floorplanning problems. EGM samples a small number of representative floorplans and then efficiently derives a high-performance floorplan by using a systematic reasoning method for the next move of OSA based on orthogonal experimental design. Furthermore, an improved swap operation is proposed which cooperates with EGM to make OSA efficient. Excellent experimental results using the Microelectronics Center of North Carolina and the Gigascale Systems Research Center benchmarks show that OSA performs better than existing methods for large floorplanning problems.
Shinn-Ying Ho, Shinn-Jang Ho, Yi-Kuang Lin, W. C.-C. Chu
IEEE Trans. Very Large Scale Integr. Syst.1
2003 A Non-parametric Image Segmentation Algorithm Using an Orthogonal Experimental Design Based Hill-Climbing
Kual-Zheng Lee, Wei-Che Chuang, Shinn-Ying Ho
IDEAL3
2003 Model-Based Pose Estimation of Human Motion Using Orthogonal Simulated Annealing
Kual-Zheng Lee, Ting-Wei Liu, Shinn-Ying Ho
IDEAL3
2003 A Novel Orthogonal Simulated Annealing Algorithm for Optimization of Electromagnetic Problems
Li-Sun Shu, Shinn-Jang Ho, Shinn-Ying Ho
IDEAL3
2003 Mesh optimization for surface approximation using an efficient coarse-to-fine evolutionary algorithm
Hui-Ling Huang, Shinn-Ying Ho
Pattern Recognit.2
2002 Design of high performance fuzzy controllers using flexible parameterized membership functions and intelligent genetic algorithms
abstract
This paper proposes a method for designing high performance fuzzy controllers with a compact rule system. The method is mainly derived from flexible parameterized membership functions (FPMFs) and a novel intelligent genetic algorithm (IGA). Each FPMF consists of flexible trapezoidal fuzzy sets and the fuzzy set is encoded by five parameters. Furthermore, the membership functions and fuzzy rules are simultaneously determined by effectively incorporating all the system parameters into chromosomes. Therefore, the optimal design of fuzzy controllers is formulated as a large parameter optimization problem, which can be effectively solved by IGA. The proposed method is demonstrated by two well-known problems, truck backing and cart centering problems. It is shown empirically that the performance of the proposed method is superior to those of existing methods in terms of the numbers of time steps and fuzzy rules.
Shinn-Ying Ho, Tai-Kang Chen, Shinn-Jang Ho
IEEE Congress on Evolutionary Computation1
2002 An evolutionary approach for pose determination and interpretation of occluded articulated objects
abstract
This paper proposes a novel evolutionary approach to a parameter solving problem for handling occluded articulated objects with any number of internal parameters representing articulation. The parameter solving problem is formulated as a parameter optimization problem and an objective function is also given based on the line segment features in the Hough space. The proposed approach uses a novel intelligent genetic algorithm (IGA) superior to conventional GAs in solving large parameter optimization problems to simultaneously solve pose determination and interpretation problems and consequently has the capabilities of accurate partial matching and robust pose determination. Effectiveness of the proposed IGA-based method is demonstrated by applying it to fitting a simplified artificial articulated model of a human body to monocular clutter images.
Shinn-Ying Ho, Zhen-Bang Huang, Shinn-Jang Ho
IEEE Congress on Evolutionary Computation1
2002 Design of an optimal nearest neighbor classifier using an intelligent genetic algorithm
abstract
The goal of designing an optimal nearest-neighbor classifier is to maximize the classification accuracy while minimizing the sizes of both the reference and feature sets. A novel intelligent genetic algorithm (IGA), which is superior to conventional genetic algorithms (GAs) in solving large parameter optimization problems, is used to effectively achieve this goal. It is shown empirically that the IGA-designed classifier outperforms existing GA-based and non-GA-based classifiers in terms of classification accuracy and the total number of parameters of the reduced sets.
Shinn-Ying Ho, Chia-Cheng Liu, Soundy Liu, Jun-Wen Jou
IEEE Congress on Evolutionary Computation1
2002 Fitness Inheritance In Multi-objective Optimization
Jian-Hung Chen, David E. Goldberg, Shinn-Ying Ho, Kumara Sastry
GECCO3
2002 Design of an optimal nearest neighbor classifier using an intelligent genetic algorithm
Shinn-Ying Ho, Chia-Cheng Liu, Soundy Liu
Pattern Recognit. Lett.1
2001 An efficient evolutionary image segmentation algorithm
abstract
In this paper, an efficient evolutionary image segmentation algorithm (EISA) is proposed. The existing evolutionary approach of image segmentation has the advantages over the other approaches such as continuous contour, non-oversegmentation, and non-thresholds, but suffers from long computation time. EISA uses a K-means algorithm to split an image into many homogeneous regions and then merges the split regions automatically using an evolutionary algorithm. The image segmentation problem is formulated as an optimization problem and the objective function is also given. EISA using a novel chromosome encoding method and a novel intelligent genetic algorithm makes the segmentation results robust and the computation time much shorter than the existing evolutionary image segmentation algorithms. Design and analysis of EISA are also presented. Experimental results of natural images with various degrees of noise demonstrate the effectiveness of EISA.
Shinn-Ying Ho, Kual-Zheng Lee
CEC1
2001 Mesh optimization for surface approximation using an efficient coarse-to-fine evolutionary algorithm
abstract
This paper investigates surface approximation using a mesh optimization approach. The mesh optimization problem is how to locate a limited number n of grid points such that the established mesh of n grid points approximates the digital surface of N sample points as closely as possible. The resulting combinatorial problem has an NP-hard search space of C(N, n) instances, i.e., the number of ways of choosing n grid points out of N sample points. A genetic algorithm-based method has been proposed for establishing optimal approximating mesh surfaces. It was shown that the GA-based method is effective in searching the combinatorial space which is intractable when n and N are in the order of thousands. This paper proposes an efficient coarse-to-fine evolutionary algorithm with a novel 2D orthogonal crossover for obtaining an optimal solution to the mesh optimization problem. It is shown empirically that the proposed coarse-to-fine evolutionary algorithm outperforms the existing GA-based method in solving the mesh optimization problem in terms of both approximation quality and convergence speed, especially in solving large mesh optimization problems.
Hui-Ling Huang, Shinn-Ying Ho
CEC2
2001 An efficient evolutionary algorithm for accurate polygonal approximation
Shinn-Ying Ho, Yeong-Chinq Chen
Pattern Recognit.1
2001 Facial modeling from an uncalibrated face image using a coarse-to-fine genetic algorithm
Shinn-Ying Ho, Hui-Ling Huang
Pattern Recognit.1
2001 Facial modeling from an uncalibrated face image using flexible generic parameterized facial models
abstract
The paper presents an optimization approach for facial modeling from an uncalibrated face image using flexible generic parameterized facial models (FGPFMs). An FGPFM consists of a topological structure and geometric knowledge of human faces. The topological description consists of a set of well-designed triangular polygons with a multilayered elastic structure in which the microstructural information can be expressed without complicated facial features. All the geometric values are obtained from a set of training facial models using statistical approaches and genetic algorithms. FGPFM can be easily modified using facial features as FGPFMs parameters to create an accurate specific three-dimensional (3D) facial model from only a photograph of an individual with a yawed face. In addition, the facial modeling problem is formulated as a parameter optimization problem. A hybrid optimization approach based on the Taguchi method and a best-first search algorithm is used to accelerate the search for a near optimal solution. Furthermore, sensitivity analysis and experimental results with texture mapping demonstrate the effectiveness of the proposed approach.
Shinn-Ying Ho, Hui-Ling Huang
IEEE Trans. Syst. Man Cybern. Part B1
2000 Designing an Efficient Fuzzy classifier Using an Intelligent Genetic Algorithm
abstract
The authors propose a method for designing an efficient fuzzy classifier that consists of a small number of fuzzy rules with only a few antecedent fuzzy sets using a novel intelligent genetic algorithm (IGA). It is known that the number of fuzzy rules will be exploded as the number of features increases. So the fuzzy classifier with many input variables has an extremely large number of fuzzy rules for high-dimensional pattern classification problems. To cope with this large rule base problem, our proposed method has the following three merits: (1) a flexible genetic parameterized fuzzy region is proposed to efficiently partition the feature space. (2) The parametric genes for representing the membership functions and fuzzy rules and the control genes used for useful pattern feature selection and dummy fuzzy rule deletion are incorporated into a single chromosome. This means that the participated features, the membership function of each antecedent fuzzy set, and the fuzzy rules are simultaneously determined. (3) The efficient fuzzy classifier design is formulated as a large parameter optimization problem (LPOP). We solve LPOPs using a novel IGA which is superior to the conventional genetic algorithms in solving LPOPs. The high performance of the proposed method is illustrated by computer simulations on the iris and wine classification problems and the simulation results are superior to those of existing methods.
Shinn-Ying Ho, Tai-Kang Chen, Shinn-Jang Ho
COMPSAC1
2000 An Efficient Quadratic Curve Approximation Using an Intelligent Genetic Algorithm
Shinn-Ying Ho, Meng-Hsun Huang
GECCO1
2000 A Simple and Fast GA-SA hybrid Image Segmentation Algorithm
Shinn-Ying Ho, Kual-Zheng Lee
GECCO1
1998 An analytic solution for the pose determination of human faces from a monocular image
Shinn-Ying Ho, Hui-Ling Huang
Pattern Recognit. Lett.1
1994 Measuring 3-D location and shape parameters of cylinders by a spatial encoding technique
abstract
We are concerned with the problem of estimation of true cylinder radius, height, location, and orientation. We present mathematical models for measuring these location and shape parameters using a spatial encoding technique. A crucial step in the proposed method is to convert the estimation problem with complex curved stripe patterns to an equivalent, but simpler, estimation problem with line stripe patterns. The notions of silhouette edges and virtual plane are introduced. Various projective geometry techniques are applied to derive the cylinder location and shape parameters. The actual experiment apparatus is set up to employ the developed mathematical models to measure the cylinder geometric parameters. Description of the image-processing tasks for extracting perceived curved stripes and stripe endpoints is given. Sensitivity analysis and proper measures are taken to consider the effects of the uncertainties of stripe endpoints and silhouette edge locations on the measurement. Experiments have confirmed that our measurement method yields quite good results for different cylinders under various measurement environment conditions.>
Zen Chen, Tsorng-Lin Chia, Shinn-Ying Ho
IEEE Trans. Robotics Autom.3
1993 Incremental model building of polyhedral objects using structured light
Zen Chen, Shinn-Ying Ho
Pattern Recognit.2
1993 An effective search approach to camera parameter estimation using an arbitrary planar calibration object
Zen Chen, Chao-Ming Wang, Shinn-Ying Ho
Pattern Recognit.3
1993 Polyhedral face reconstruction and modeling from a single image with structured light
abstract
The determination of the 3-D geometric model of visible polyhedral faces from a single view is addressed. It applies a grid coding technique to derive the normal vector of each visible polyhedral face and the depth parameter of the face equation based on the given dimensions of the grid on the code plate. It is shown that the method is correspondenceless. Furthermore, it also determines the 3-D face model for all visible polyhedral faces including the occlusion relations between faces. Two edge reconstruction methods are given and their goodness are compared. For the determination of the final object model, three different integration methods are given. Some experiments are reported to illustrate the method.>
Zen Chen, Shinn-Ying Ho, Din-Chang Tseng
IEEE Trans. Syst. Man Cybern.2
1991 Computer vision for robust 3D aircraft recognition with fast library search
Zen Chen, Shinn-Ying Ho
Pattern Recognit.2