Hui-Ling Huang

dblp:34/6358 · DBLP profile ↗
← Back
38ranked-venue papers
11as first author
0since 2021 · last 2017
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 5 first-authorArtificial intelligence and machine learning · 12 · 4 first-authorSystems, architecture and hardware · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-authorTheory of computation · 2Computer networks · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Interconnection networks and networks-on-chip · 33% Distributed systems · 33% Hardware reliability and fault tolerance · 33%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
negative sampling
0.312017
ESA-UbiSite: accurate prediction of human ubiquitination sites by identifying a set of effective negatives · Bioinform. 2017
Bioinformatics and computational biology › proteomics
post-translational modification prediction
0.312017
ESA-UbiSite: accurate prediction of human ubiquitination sites by identifying a set of effective negatives · Bioinform. 2017
Bioinformatics and computational biology › protein sequence analysis › protein sequence annotation
ubiquitination site prediction
0.312017
ESA-UbiSite: accurate prediction of human ubiquitination sites by identifying a set of effective negatives · Bioinform. 2017
Hardware reliability and fault tolerance › network fault tolerance
fault diameter
0.011999
Combinatorial Properties of Two-Level Hypernet Networks · IEEE Trans. Parallel Distributed Syst. 1999
Distributed systems
fault tolerance
0.011999
Combinatorial Properties of Two-Level Hypernet Networks · IEEE Trans. Parallel Distributed Syst. 1999
Interconnection networks and networks-on-chip
network topology
0.011999
Combinatorial Properties of Two-Level Hypernet Networks · IEEE Trans. Parallel Distributed Syst. 1999
Graph algorithms and graph theory
graph theory
0.011999
Combinatorial Properties of Two-Level Hypernet Networks · IEEE Trans. Parallel Distributed Syst. 1999

Methods — techniques the papers use, named apart from their topics

support vector machine · 0.3physicochemical properties · 0.3evolutionary screening algorithm · 0.3wide diameter · 0.0container analysis · 0.0
YearPublicationVenuePosition
2017 ESA-UbiSite: accurate prediction of human ubiquitination sites by identifying a set of effective negatives
abstract
Motivation: Numerous ubiquitination sites remain undiscovered because of the limitations of mass spectrometry-based methods. Existing prediction methods use randomly selected non-validated sites as non-ubiquitination sites to train ubiquitination site prediction models. Results: We propose an evolutionary screening algorithm (ESA) to select effective negatives among non-validated sites and an ESA-based prediction method, ESA-UbiSite, to identify human ubiquitination sites. The ESA selects non-validated sites least likely to be ubiquitination sites as training negatives. Moreover, the ESA and ESA-UbiSite use a set of well-selected physicochemical properties together with a support vector machine for accurate prediction. Experimental results show that ESA-UbiSite with effective negatives achieved 0.92 test accuracy and a Matthews's correlation coefficient of 0.48, better than existing prediction methods. The ESA increased ESA-UbiSite's test accuracy from 0.75 to 0.92 and can improve other post-translational modification site prediction methods. Availability and Implementation: An ESA-UbiSite-based web server has been established at http://iclab.life.nctu.edu.tw/iclab_webtools/ESAUbiSite/ . Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Jyun-Rong Wang, Wen-Lin Huang, Ming-Ju Tsai, Kai-Ti Hsu, Hui-Ling Huang, Shinn-Ying Ho
Bioinform.5
2017 Recognition of handwritten Lanna Dhamma characters using a set of optimally designed moment features
Papangkorn Inkeaw, Phasit Charoenkwan, Hui-Ling Huang, Sanparith Marukatat, Shinn-Ying Ho, Jeerayut Chaijaruwanich
Int. J. Document Anal. Recognit.3
2016 A hydrophobic spine stabilizes a surface-exposed α-helix according to analysis of the solvent-accessible surface area
abstract
BACKGROUND: Most of hydrophilic and hydrophobic residues are thought to be exposed and buried in proteins, respectively. In contrast to the majority of the existing studies on protein folding characteristics using protein structures, in this study, our aim was to design predictors for estimating relative solvent accessibility (RSA) of amino acid residues to discover protein folding characteristics from sequences. METHODS: The proposed 20 real-value RSA predictors were designed on the basis of the support vector regression method with a set of informative physicochemical properties (PCPs) obtained by means of an optimal feature selection algorithm. Then, molecular dynamics simulations were performed for validating the knowledge discovered by analysis of the selected PCPs. RESULTS: The RSA predictors had the mean absolute error of 14.11% and a correlation coefficient of 0.69, better than the existing predictors. The hydrophilic-residue predictors preferred PCPs of buried amino acid residues to PCPs of exposed ones as prediction features. A hydrophobic spine composed of exposed hydrophobic residues of an α-helix was discovered by analyzing the PCPs of RSA predictors corresponding to hydrophobic residues. For example, the results of a molecular dynamics simulation of wild-type sequences and their mutants showed that proteins 1MOF and 2WRP_H16I (Protein Data Bank IDs), which have a perfectly hydrophobic spine, have more stable structures than 1MOF_I54D and 2WRP do (which do not have a perfectly hydrophobic spine). CONCLUSIONS: We identified informative PCPs to design high-performance RSA predictors and to analyze these PCPs for identification of novel protein folding characteristics. A hydrophobic spine in a protein can help to stabilize exposed α-helices.
Yi-Fan Liou, Hui-Ling Huang, Shinn-Ying Ho
BMC Bioinform.2
2016 SCMBYK: prediction and characterization of bacterial tyrosine-kinases based on propensity scores of dipeptides
abstract
BACKGROUND: Bacterial tyrosine-kinases (BY-kinases), which play an important role in numerous cellular processes, are characterized as a separate class of enzymes and share no structural similarity with their eukaryotic counterparts. However, in silico methods for predicting BY-kinases have not been developed yet. Since these enzymes are involved in key regulatory processes, and are promising targets for anti-bacterial drug design, it is desirable to develop a simple and easily interpretable predictor to gain new insights into bacterial tyrosine phosphorylation. This study proposes a novel SCMBYK method for predicting and characterizing BY-kinases. RESULTS: A dataset consisting of 797 BY-kinases and 783 non-BY-kinases was established to design the SCMBYK predictor, which achieved training and test accuracies of 97.55 and 96.73%, respectively. Furthermore, the leave-one-phylum-out method was used to predict specific bacterial phyla hosts of target sequences, gaining 97.39% average test accuracy. After analyzing SCMBYK-derived propensity scores, four characteristics of BY-kinases were determined: 1) BY-kinases tend to be composed of α-helices; 2) the amino-acid content of extracellular regions of BY-kinases is expected to be dominated by residues such as Val, Ile, Phe and Tyr; 3) BY-kinases structurally resemble nuclear proteins; 4) different domains play different roles in triggering BY-kinase activity. CONCLUSIONS: The SCMBYK predictor is an effective method for identification of possible BY-kinases. Furthermore, it can be used as a part of a novel drug repurposing method, which recognizes putative BY-kinases and matches them to approved drugs. Among other results, our analysis revealed that azathioprine could suppress the virulence of M. tuberculosis, and thus be considered as a potential antibiotic for tuberculosis treatment.
Tamara Vasylenko, Yi-Fan Liou, Po-Chin Chiou, Hsiao-Wei Chu, Yung-Sung Lai, Yu-Ling Chou, Hui-Ling Huang, Shinn-Ying Ho
BMC Bioinform.7
2015 Discovery of prognostic biomarkers for predicting lung cancer metastasis using microarray and survival data
abstract
BACKGROUND: Few studies have investigated prognostic biomarkers of distant metastases of lung cancer. One of the central difficulties in identifying biomarkers from microarray data is the availability of only a small number of samples, which results overtraining. Recently obtained evidence reveals that epithelial-mesenchymal transition (EMT) of tumor cells causes metastasis, which is detrimental to patients' survival. RESULTS: This work proposes a novel optimization approach to discovering EMT-related prognostic biomarkers to predict the distant metastasis of lung cancer using both microarray and survival data. This weighted objective function maximizes both the accuracy of prediction of distant metastasis and the area between the disease-free survival curves of the non-distant and distant metastases. Seventy-eight patients with lung cancer and a follow-up time of 120 months are used to identify a set of gene markers and an independent cohort of 26 patients is used to evaluate the identified biomarkers. The medical records of the 78 patients show a significant difference between the disease-free survival times of the 37 non-distant- and the 41 distant-metastasis patients. The experimental results thus obtained are as follows. 1) The use of disease-free survival curves can compensate for the shortcoming of insufficient samples and greatly increase the test accuracy by 11.10%; and 2) the support vector machine with a set of 17 transcripts, such as CCL16 and CDKN2AIP, can yield a leave-one-out cross-validation accuracy of 93.59%, a test accuracy of 76.92%, a large disease-free survival area of 74.81%, and a mean survival prediction error of 3.99 months. The identified putative biomarkers are examined using related studies and signaling pathways to reveal the potential effectiveness of the biomarkers in prospective confirmatory studies. CONCLUSIONS: The proposed new optimization approach to identifying prognostic biomarkers by combining multiple sources of data (microarray and survival) can facilitate the accurate selection of biomarkers that are most relevant to the disease while solving the problem of insufficient samples.
Hui-Ling Huang, Yu-Chung Wu, Li-Jen Su, Yun-Ju Huang, Phasit Charoenkwan, Hua-Chin Lee, William C. Chu, Shinn-Ying Ho
BMC Bioinform.1
2015 Characterizing informative sequence descriptors and predicting binding affinities of heterodimeric protein complexes
abstract
BACKGROUND: Protein-protein interactions (PPIs) are involved in various biological processes, and underlying mechanism of the interactions plays a crucial role in therapeutics and protein engineering. Most machine learning approaches have been developed for predicting the binding affinity of protein-protein complexes based on structure and functional information. This work aims to predict the binding affinity of heterodimeric protein complexes from sequences only. RESULTS: This work proposes a support vector machine (SVM) based binding affinity classifier, called SVM-BAC, to classify heterodimeric protein complexes based on the prediction of their binding affinity. SVM-BAC identified 14 of 580 sequence descriptors (physicochemical, energetic and conformational properties of the 20 amino acids) to classify 216 heterodimeric protein complexes into low and high binding affinity. SVM-BAC yielded the training accuracy, sensitivity, specificity, AUC and test accuracy of 85.80%, 0.89, 0.83, 0.86 and 83.33%, respectively, better than existing machine learning algorithms. The 14 features and support vector regression were further used to estimate the binding affinities (Pkd) of 200 heterodimeric protein complexes. Prediction performance of a Jackknife test was the correlation coefficient of 0.34 and mean absolute error of 1.4. We further analyze three informative physicochemical properties according to their contribution to prediction performance. Results reveal that the following properties are effective in predicting the binding affinity of heterodimeric protein complexes: apparent partition energy based on buried molar fractions, relations between chemical structure and biological activity in principal component analysis IV, and normalized frequency of beta turn. CONCLUSIONS: The proposed sequence-based prediction method SVM-BAC uses an optimal feature selection method to identify 14 informative features to classify and predict binding affinity of heterodimeric protein complexes. The characterization analysis revealed that the average numbers of beta turns and hydrogen bonds at protein-protein interfaces in high binding affinity complexes are more than those in low binding affinity complexes.
Yerukala Sathipati Srinivasulu, Jyun-Rong Wang, Kai-Ti Hsu, Ming-Ju Tsai, Phasit Charoenkwan, Wen-Lin Huang, Hui-Ling Huang, Shinn-Ying Ho
BMC Bioinform.7
2015 SCMPSP: Prediction and characterization of photosynthetic proteins based on a scoring card method
abstract
BACKGROUND: Photosynthetic proteins (PSPs) greatly differ in their structure and function as they are involved in numerous subprocesses that take place inside an organelle called a chloroplast. Few studies predict PSPs from sequences due to their high variety of sequences and structues. This work aims to predict and characterize PSPs by establishing the datasets of PSP and non-PSP sequences and developing prediction methods. RESULTS: A novel bioinformatics method of predicting and characterizing PSPs based on scoring card method (SCMPSP) was used. First, a dataset consisting of 649 PSPs was established by using a Gene Ontology term GO:0015979 and 649 non-PSPs from the SwissProt database with sequence identity <= 25%.- Several prediction methods are presented based on support vector machine (SVM), decision tree J48, Bayes, BLAST, and SCM. The SVM method using dipeptide features-performed well and yielded - a test accuracy of 72.31%. The SCMPSP method uses the estimated propensity scores of 400 dipeptides - as PSPs and has a test accuracy of 71.54%, which is comparable to that of the SVM method. The derived propensity scores of 20 amino acids were further used to identify informative physicochemical properties for characterizing PSPs. The analytical results reveal the following four characteristics of PSPs: 1) PSPs favour hydrophobic side chain amino acids; 2) PSPs are composed of the amino acids prone to form helices in membrane environments; 3) PSPs have low interaction with water; and 4) PSPs prefer to be composed of the amino acids of electron-reactive side chains. CONCLUSIONS: The SCMPSP method not only estimates the propensity of a sequence to be PSPs, it also discovers characteristics that further improve understanding of PSPs. The SCMPSP source code and the datasets used in this study are available at http://iclab.life.nctu.edu.tw/SCMPSP/.
Tamara Vasylenko, Yi-Fan Liou, Hong-An Chen, Phasit Charoenkwan, Hui-Ling Huang, Shinn-Ying Ho
BMC Bioinform.5
2014 SCMHBP: prediction and analysis of heme binding proteins using propensity scores of dipeptides
abstract
BACKGROUND: Heme binding proteins (HBPs) are metalloproteins that contain a heme ligand (an iron-porphyrin complex) as the prosthetic group. Several computational methods have been proposed to predict heme binding residues and thereby to understand the interactions between heme and its host proteins. However, few in silico methods for identifying HBPs have been proposed. RESULTS: This work proposes a scoring card method (SCM) based method (named SCMHBP) for predicting and analyzing HBPs from sequences. A balanced dataset of 747 HBPs (selected using a Gene Ontology term GO:0020037) and 747 non-HBPs (selected from 91,414 putative non-HBPs) with an identity of 25% was firstly established. Consequently, a set of scores that quantified the propensity of amino acids and dipeptides to be HBPs is estimated using SCM to maximize the predictive accuracy of SCMHBP. Finally, the informative physicochemical properties of 20 amino acids are identified by utilizing the estimated propensity scores to be used to categorize HBPs. The training and mean test accuracies of SCMHBP applied to three independent test datasets are 85.90% and 71.57%, respectively. SCMHBP performs well relative to comparison with such methods as support vector machine (SVM), decision tree J48, and Bayes classifiers. The putative non-HBPs with high sequence propensity scores are potential HBPs, which can be further validated by experimental confirmation. The propensity scores of individual amino acids and dipeptides are examined to elucidate the interactions between heme and its host proteins. The following characteristics of HBPs are derived from the propensity scores: 1) aromatic side chains are important to the effectiveness of specific HBP functions; 2) a hydrophobic environment is important in the interaction between heme and binding sites; and 3) the whole HBP has low flexibility whereas the heme binding residues are relatively flexible. CONCLUSIONS: SCMHBP yields knowledge that improves our understanding of HBPs rather than merely improves the prediction accuracy in predicting HBPs.
Yi-Fan Liou, Phasit Charoenkwan, Yerukala Sathipati Srinivasulu, Tamara Vasylenko, Shih-Chung Lai, Hua-Chin Lee, Yi-Hsiung Chen, Hui-Ling Huang, Shinn-Ying Ho
BMC Bioinform.8
2013 Prediction of Mouse Senescence from HE-Stain Liver Images Using an Ensemble SVM Classifier
Hui-Ling Huang, Ming-Hsin Hsu, Hua-Chin Lee, Phasit Charoenkwan, Shinn-Jang Ho, Shinn-Ying Ho
ACIIDS (2)1
2013 Designing predictors of halophilic and non-halophilic proteins using support vector machines
abstract
Finding the molecular features causes the halophilicity in the halostable organisms is helpful to understand the halophilic adaption. In this study, we proposed a prediction method for halophilic proteins by using a machine learning method. The stages of this study are six-fold. First, we establish a non-redundant dataset of the halophilic proteins, collected from NCBI, Uniprotkb and EMBL-EBI databases. The dataset consists of 245 positive and negative proteins with sequence identity <;25%. Second, the protein sequences are represented by three types of feature vector sets which include amino acid composition, dipeptide composition, and physicochemical properties. Third, we propose three classifiers based on support vector machine (SVM) to classify the halophilic proteins and non-halophilic proteins. Fourth, the independent test accuracies of the three efficient classifiers are larger than 83%. Fifth, an inheritable biobjective combinatory genetic algorithm is utilized to select a set of 11 physicochemical properties (PCPs). Sixth, these abundant amino acids, high different dipeptides (amino acid pair) and 11 informative PCPs can support to analyze the halophilic and non-halophilic proteins.
Hui-Ling Huang, Yerukala Sathipati Srinivasulu, Phasit Charoenkwan, Hua-Chin Lee, Shinn-Ying Ho
CIBCB1
2013 Predicting protein crystallization using a simple scoring card method
abstract
Many computational methods have been developed to predict protein crystallization. Most methods use amino acid and dipeptide compositions as part of the informative features. To advance the prediction accuracy, the support vector machine (SVM) based classifiers and ensemble approaches were effective and commonly-used techniques. However, these techniques suffer from the low interpretation ability of insight into crystallization. In this study, we utilize a newly-developed scoring card method (SCM) with a dipeptide composition feature to predict protein crystallization. This SCM classifier obtains prediction results 74%, 0.55 and 0.83 for accuracy, sensitivity and specificity, respectively, which is comparable to the SVM classifier using the same benchmarks. The experimental results show that the SCM classifier has advantages of simplicity, high interpretability, and high accuracy in predicting protein crystallization, compared with existing SVM-basedensemble classifiers.
Watshara Shoombuatong, Hui-Ling Huang, Jeerayut Chaijaruwanich, Phasit Charoenkwan, Hua-Chin Lee, Shinn-Ying Ho
CIBCB2
2013 HCS-Neurons: identifying phenotypic changes in multi-neuron images upon drug treatments of high-content screening
abstract
BACKGROUND: High-content screening (HCS) has become a powerful tool for drug discovery. However, the discovery of drugs targeting neurons is still hampered by the inability to accurately identify and quantify the phenotypic changes of multiple neurons in a single image (named multi-neuron image) of a high-content screen. Therefore, it is desirable to develop an automated image analysis method for analyzing multi-neuron images. RESULTS: We propose an automated analysis method with novel descriptors of neuromorphology features for analyzing HCS-based multi-neuron images, called HCS-neurons. To observe multiple phenotypic changes of neurons, we propose two kinds of descriptors which are neuron feature descriptor (NFD) of 13 neuromorphology features, e.g., neurite length, and generic feature descriptors (GFDs), e.g., Haralick texture. HCS-neurons can 1) automatically extract all quantitative phenotype features in both NFD and GFDs, 2) identify statistically significant phenotypic changes upon drug treatments using ANOVA and regression analysis, and 3) generate an accurate classifier to group neurons treated by different drug concentrations using support vector machine and an intelligent feature selection method. To evaluate HCS-neurons, we treated P19 neurons with nocodazole (a microtubule depolymerizing drug which has been shown to impair neurite development) at six concentrations ranging from 0 to 1000 ng/mL. The experimental results show that all the 13 features of NFD have statistically significant difference with respect to changes in various levels of nocodazole drug concentrations (NDC) and the phenotypic changes of neurites were consistent to the known effect of nocodazole in promoting neurite retraction. Three identified features, total neurite length, average neurite length, and average neurite area were able to achieve an independent test accuracy of 90.28% for the six-dosage classification problem. This NFD module and neuron image datasets are provided as a freely downloadable MatLab project at http://iclab.life.nctu.edu.tw/HCS-Neurons. CONCLUSIONS: Few automatic methods focus on analyzing multi-neuron images collected from HCS used in drug discovery. We provided an automatic HCS-based method for generating accurate classifiers to classify neurons based on their phenotypic changes upon drug treatments. The proposed HCS-neurons method is helpful in identifying and classifying chemical or biological molecules that alter the morphology of a group of neurons in HCS.
Phasit Charoenkwan, Eric Hwang, Robert W. Cutler, Hua-Chin Lee, Li-Wei Ko, Hui-Ling Huang, Shinn-Ying Ho
BMC Bioinform.6
2012 Motion sickness estimation system
abstract
Motion sickness occurs when the brain receives conflicting sensory information from body, inner ear and eyes [1]. In some cases, a decreased ability to actively control the body's postural motion also causes motion sickness [2][3]. Many previous studies have indicated that motion sickness had negative effect on driving performance, and sometimes lead to serious traffic accidents due to self-control ability decline. Therefore motion sickness becomes a very important issue in our daily life especially considering driving safety. There are many attempts made by researchers to realize motion sickness, and detect motion sickness in the early stage. Although many motion-sickness-related biomarkers have been identified, estimating human motion sickness level (MSL) remains a challenge in operational environment. In our past studies, we found that features in the occipital area were highly correlated with the driver's driving performance. In this study, we designed a virtual-reality (VR) based driving environment with instinct-MSL-reporting mechanism. When a subject performed a driving task, his/her brain EEG was recorded simultaneously. From those EEG data, features associated with left motor brain area, parietal brain area and occipital midline brain area which predicted MSL were extracted by an optimal classifier implemented by an inheritable bi-objective combinatorial genetic algorithm (IBCGA) with support vector machine. Unlike traditional correlation-based method, IBCGA aims to select a small set of EEG features and maximize the prediction accuracy simultaneously in BCI applications. Once the optimal feature set predicting MSL is successfully found, a driver's cognitive state can be monitored.
Chin-Teng Lin, Shu-Fang Tsai, Hua-Chin Lee, Hui-Ling Huang, Shinn-Ying Ho, Li-Wei Ko
IJCNN4
2012 Prediction and analysis of protein solubility using a novel scoring card method with dipeptide composition
abstract
BACKGROUND: Existing methods for predicting protein solubility on overexpression in Escherichia coli advance performance by using ensemble classifiers such as two-stage support vector machine (SVM) based classifiers and a number of feature types such as physicochemical properties, amino acid and dipeptide composition, accompanied with feature selection. It is desirable to develop a simple and easily interpretable method for predicting protein solubility, compared to existing complex SVM-based methods. RESULTS: This study proposes a novel scoring card method (SCM) by using dipeptide composition only to estimate solubility scores of sequences for predicting protein solubility. SCM calculates the propensities of 400 individual dipeptides to be soluble using statistic discrimination between soluble and insoluble proteins of a training data set. Consequently, the propensity scores of all dipeptides are further optimized using an intelligent genetic algorithm. The solubility score of a sequence is determined by the weighted sum of all propensity scores and dipeptide composition. To evaluate SCM by performance comparisons, four data sets with different sizes and variation degrees of experimental conditions were used. The results show that the simple method SCM with interpretable propensities of dipeptides has promising performance, compared with existing SVM-based ensemble methods with a number of feature types. Furthermore, the propensities of dipeptides and solubility scores of sequences can provide insights to protein solubility. For example, the analysis of dipeptide scores shows high propensity of α-helix structure and thermophilic proteins to be soluble. CONCLUSIONS: The propensities of individual dipeptides to be soluble are varied for proteins under altered experimental conditions. For accurately predicting protein solubility using SCM, it is better to customize the score card of dipeptide propensities by using a training data set under the same specified experimental conditions. The proposed method SCM with solubility scores and dipeptide propensities can be easily applied to the protein function prediction problems that dipeptide composition features play an important role. AVAILABILITY: The used datasets, source codes of SCM, and supplementary files are available at http://iclab.life.nctu.edu.tw/SCM/.
Hui-Ling Huang, Phasit Charoenkwan, Te-Fen Kao, Hua-Chin Lee, Fang-Lin Chang, Wen-Lin Huang, Shinn-Jang Ho, Li-Sun Shu, Shinn-Ying Ho
BMC Bioinform.1
2012 Variable patrol Planning of Multi-Robot Systems by a Cooperative Auction System
abstract
A cooperative auction system (CAS) is proposed to solve the large-scale multi-robot patrol planning problem. Each robot picks its own patrol points via the cooperative auction system and the system continuously re-auctions, based on the team work performance. The proposed method not only works in static environments but also considers variable path planning when the number of mobile robots increases or decreases during patrol. From the results of the simulation, the proposed approach demonstrates decreased time complexity, a lower routing path cost, improved balance of workload among robots, and the potential to scale to a large number of robots and is adaptive to environmental perturbations when the number of robots changes during patrol.
Jin-Ling Lin, Kao-Shing Hwang, Hui-Ling Huang
Cybern. Syst.3
2012 Knowledge management fit and its implications for business performance: A profile deviation analysis
Yue-Yang Chen, Hui-Ling Huang
Knowl. Based Syst.2
2011 NeurphologyJ: an automatic neuronal morphology quantification method and its application in pharmacological discovery
abstract
BACKGROUND: Automatic quantification of neuronal morphology from images of fluorescence microscopy plays an increasingly important role in high-content screenings. However, there exist very few freeware tools and methods which provide automatic neuronal morphology quantification for pharmacological discovery. RESULTS: This study proposes an effective quantification method, called NeurphologyJ, capable of automatically quantifying neuronal morphologies such as soma number and size, neurite length, and neurite branching complexity (which is highly related to the numbers of attachment points and ending points). NeurphologyJ is implemented as a plugin to ImageJ, an open-source Java-based image processing and analysis platform. The high performance of NeurphologyJ arises mainly from an elegant image enhancement method. Consequently, some morphology operations of image processing can be efficiently applied. We evaluated NeurphologyJ by comparing it with both the computer-aided manual tracing method NeuronJ and an existing ImageJ-based plugin method NeuriteTracer. Our results reveal that NeurphologyJ is comparable to NeuronJ, that the coefficient correlation between the estimated neurite lengths is as high as 0.992. NeurphologyJ can accurately measure neurite length, soma number, neurite attachment points, and neurite ending points from a single image. Furthermore, the quantification result of nocodazole perturbation is consistent with its known inhibitory effect on neurite outgrowth. We were also able to calculate the IC50 of nocodazole using NeurphologyJ. This reveals that NeurphologyJ is effective enough to be utilized in applications of pharmacological discoveries. CONCLUSIONS: This study proposes an automatic and fast neuronal quantification method NeurphologyJ. The ImageJ plugin with supports of batch processing is easily customized for dealing with high-content screening applications. The source codes of NeurphologyJ (interactive and high-throughput versions) and the images used for testing are freely available (see Availability).
Shinn-Ying Ho, Chih-Yuan Chao, Hui-Ling Huang, Tzai-Wen Chiu, Phasit Charoenkwan, Eric Hwang
BMC Bioinform.3
2011 Predicting and analyzing DNA-binding domains using a systematic approach to identifying a set of informative physicochemical and biochemical properties
abstract
BACKGROUND: Existing methods of predicting DNA-binding proteins used valuable features of physicochemical properties to design support vector machine (SVM) based classifiers. Generally, selection of physicochemical properties and determination of their corresponding feature vectors rely mainly on known properties of binding mechanism and experience of designers. However, there exists a troublesome problem for designers that some different physicochemical properties have similar vectors of representing 20 amino acids and some closely related physicochemical properties have dissimilar vectors. RESULTS: This study proposes a systematic approach (named Auto-IDPCPs) to automatically identify a set of physicochemical and biochemical properties in the AAindex database to design SVM-based classifiers for predicting and analyzing DNA-binding domains/proteins. Auto-IDPCPs consists of 1) clustering 531 amino acid indices in AAindex into 20 clusters using a fuzzy c-means algorithm, 2) utilizing an efficient genetic algorithm based optimization method IBCGA to select an informative feature set of size m to represent sequences, and 3) analyzing the selected features to identify related physicochemical properties which may affect the binding mechanism of DNA-binding domains/proteins. The proposed Auto-IDPCPs identified m = 22 features of properties belonging to five clusters for predicting DNA-binding domains with a five-fold cross-validation accuracy of 87.12%, which is promising compared with the accuracy of 86.62% of the existing method PSSM-400. For predicting DNA-binding sequences, the accuracy of 75.50% was obtained using m = 28 features, where PSSM-400 has an accuracy of 74.22%. Auto-IDPCPs and PSSM-400 have accuracies of 80.73% and 82.81%, respectively, applied to an independent test data set of DNA-binding domains. Some typical physicochemical properties discovered are hydrophobicity, secondary structure, charge, solvent accessibility, polarity, flexibility, normalized Van Der Waals volume, pK (pK-C, pK-N, pK-COOH and pK-a(RCOOH)), etc. CONCLUSIONS: The proposed approach Auto-IDPCPs would help designers to investigate informative physicochemical and biochemical properties by considering both prediction accuracy and analysis of binding mechanism simultaneously. The approach Auto-IDPCPs can be also applicable to predict and analyze other protein functions from sequences.
Hui-Ling Huang, I-Che Lin, Yi-Fan Liou, Chia-Ta Tsai, Kai-Ti Hsu, Wen-Lin Huang, Shinn-Jang Ho, Shinn-Ying Ho
BMC Bioinform.1
2011 Edge-bipancyclicity of star graphs with faulty elements
Chao-Wen Huang, Hui-Ling Huang, Sun-Yuan Hsieh
Theor. Comput. Sci.2
2009 1-vertex-fault-tolerant cycles embedding on folded hypercubes
Sun-Yuan Hsieh, Che-Nan Kuo, Hui-Ling Huang
Discret. Appl. Math.3
2008 Inferring S-system models of genetic networks from a time-series real data set of gene expression profiles
abstract
It is desirable to infer cellular dynamic regulation networks from gene expression profiles to discover more delicate and substantial functions in molecular biology, biochemistry, bioengineering, and pharmaceutics. The S-system model is suitable to characterize biochemical network systems and capable of analyzing the regulatory system dynamics. To cope with the problem ldquomultiplicity of solutionsrdquo, a sufficient amount of data sets of time-series gene expression profiles were often used. An efficient newly-developed method iTEA was proposed to effectively obtain S-system models from a large number (e.g., 15) of simulated data sets with/without noise. In this study, we propose an extended optimization method (named iTEAP) based on iTEA to infer the S-system models of genetic networks from a time-series real data set of gene expression profiles (using SOS DNA microarray data inE.colias an example). The algorithm iTEAP generated additionally multiple data sets of gene expression profiles by perturbing the given data set. The results reveal that 1) iTEAP can obtain S-system models with high-quality profiles to best fit the observed profiles; 2) the performance of using multiple data sets is better than that of using a single data set in terms of solution quality, and 3) the effectiveness of iTEAP using a single data set is close to that of iTEA using two real data sets.
Hui-Ling Huang, Kuan-Wei Chen, Shinn-Jang Ho, Shinn-Ying Ho
IEEE Congress on Evolutionary Computation1
2007 Boosting Evolutionary Support Vector Machine for Designing Tumor Classifiers from Microarray Data
abstract
Since there are multiple sets of relevant genes having the same high accuracy in fitting training data called model uncertainty, to identify a small set of informative genes from microarray data for designing an accurate tumor classifier for unknown samples is intractable. Support vector machine (SVM), a supervised machine learning technique, is one of the methods successfully applied to cancer diagnosis problems. This study proposes an SVM-based classifier with automatic feature selection associated with a boosting strategy. The proposed boosting evolutionary support vector machine (named BESVM) hybridizes the advantages of SVM, boosting using a majority-voting ensemble and an intelligent genetic algorithm for gene selection. The merits of the BESVM-based classifier are threefold: 1) a small set of used genes, 2) accurate test classification using leave-one-out cross-validation, and 3) robust performance by avoiding overfitting training data. Five benchmark datasets were used to evaluate the BESVM-based classifier. Simulation results reveal that BESVM performs well having a mean accuracy 94.26% using only 10.1 genes averagely, compared with the existing SVM and non-SVM based classifiers
Hui-Ling Huang, Yi-Hsiung Chen, Dwight D. Koeberl, Shinn-Ying Ho
CIBCB1
2007 The m-pancycle-connectivity of a WK-Recursive network
Jywe-Fei Fang, Yuh-Rau Wang, Hui-Ling Huang
Inf. Sci.3
2007 An Intelligent Two-Stage Evolutionary Algorithm for Dynamic Pathway Identification From Gene Expression Profiles
abstract
From gene expression profiles, it is desirable to rebuild cellular dynamic regulation networks to discover more delicate and substantial functions in molecular biology, biochemistry, bioengineering and pharmaceutics. S-system model is suitable to characterize biochemical network systems and capable to analyze the regulatory system dynamics. However, inference of an S-system model of N-gene genetic networks has 2N(N+1) parameters in a set of non-linear differential equations to be optimized. This paper proposes an intelligent two-stage evolutionary algorithm (iTEA) to efficiently infer the S-system models of genetic networks from time-series data of gene expression. To cope with curse of dimensionality, the proposed algorithm consists of two stages where each uses a divide-and-conquer strategy. The optimization problem is first decomposed into N subproblems having 2(N+1) parameters each. At the first stage, each subproblem is solved using a novel intelligent genetic algorithm (IGA) with intelligent crossover based on orthogonal experimental design (OED). At the second stage, the obtained N solutions to the N subproblems are combined and refined using an OED-based simulated annealing algorithm for handling noisy gene expression profiles. The effectiveness of iTEA is evaluated using simulated expression patterns with and without noise running on a single-processor PC. It is shown that 1) IGA is efficient enough to solve subproblems; 2) IGA is significantly superior to the existing method SPXGA; and 3) iTEA performs well in inferring S-system models for dynamic pathway identification.
Shinn-Ying Ho, Chih-Hung Hsieh, Fu-Chieh Yu, Hui-Ling Huang
IEEE ACM Trans. Comput. Biol. Bioinform.4
2007 Panconnectivity and edge-pancyclicity of 3-ary N -cubes
Sun-Yuan Hsieh, Tsong-Jie Lin, Hui-Ling Huang
J. Supercomput.3
2006 Scoring Method for Tumor Prediction from Microarray Data Using an Evolutionary Fuzzy Classifier
Shinn-Ying Ho, Chih-Hung Hsieh, Kuan-Wei Chen, Hui-Ling Huang, Hung-Ming Chen, Shinn-Jang Ho
PAKDD4
2005 Evolutionary divide-and-conquer approach to inferring S-system models of genetic networks
abstract
This paper proposes an efficient evolutionary divide-and-conquer approach (EDACA) to inferring S-system models of genetic networks from time-series data of gene expression. Inference of an S-system model has 2N(N+1) parameters to be optimized, where N is the number of genes in a genetic network. To cope with higher dimensionality, the proposed approach consists of two stages where each uses a divide-and-conquer strategy. The optimization problem is first decomposed into N subproblems having 2(N+1) parameters each. At the first stage, each subproblem is solved using a novel intelligent genetic algorithm (IGA) with intelligent crossover based on orthogonal experimental design (OED). The intelligent crossover divides two parents into n pairs of parameter groups, economically identifies the potentially better one of two groups of each pair, and systematically obtains a potentially good approximation to the best one of all 2/sup n/ combinations using at most 2n function evaluations. At the second stage, the obtained N solutions to the N subproblems are combined and refined using an OED-based simulated annealing algorithm (OSA) for handling noisy gene expression data. The effectiveness of EDACA is evaluated using simulated expression patterns with/without noise running on a single-CPU PC. It is shown that: 1) IGA is efficient enough to solve subproblems; 2) IGA is significantly superior to the existing method of using GA with simplex crossover; and 3) EDACA performs well in inferring S-system models of genetic networks from small-noise gene expression data.
Shinn-Ying Ho, Chih-Hung Hsieh, Fu-Chieh Yu, Hui-Ling Huang
Congress on Evolutionary Computation4
2005 Efficient gene selection for classification of microarray data
abstract
Microarray is a useful technique for measuring expression data of thousands of genes simultaneously. One of challenges in classification of microarray data is to select a minimal number of relevant genes which can maximize classification accuracy. Many gene selection methods as well as their corresponding classifiers have been proposed. One of existing analysis methods is the hybrid approach based on genetic algorithm and maximum likelihood classification (GA/MLHD). In this paper, an intelligent genetic algorithm (IGA) using control genes and an improved fitness function is proposed to determine the minimal number of relevant genes and identify these genes, while maximizing classification accuracy simultaneously. The experimental results show that our approach is superior to the existing method GA/MLHD in terms of the number of selected genes, classification accuracy, and robustness of selected genes and accuracy, especially for the datasets which have numerous categories and a large number of testing genes inside.
Shinn-Ying Ho, Chong-Cheng Lee, Hung-Ming Chen, Hui-Ling Huang
Congress on Evolutionary Computation4
2005 Flexible protein-ligand docking using particle swarm optimization
abstract
Many protein-ligand docking problems attempt to predict the bound conformations of two interacting molecules. Consequently, the docking problem requires a powerful search technique to explore the translations, orientations, and each torsion until an ideal site has been found. Therefore, protein-ligand docking can be formulated as a parameter optimization problem. However, highly flexible ligands have a lot of torsions. Therefore, the optimization problem of highly flexible docking would become more difficult due to the increment of parameter number and interactions among these parameters. We proposed a novel method SODOCK based on particle swarm optimization (PSO) for solving flexible protein-ligand docking problems. PSO has significant effect on the optimization of parameters with strong interactions. A commonly used efficient local search is incorporated into SODOCK to improve the efficiency and robustness of PSO. SODOCK is efficient for both types of ligands with small and large numbers of torsions. It is shown by computer simulation that SODOCK performs well in obtaining accurate conformations, compared with some of state-of-the-art methods. Moreover, it is also shown that SODOCK is superior to AutoDock using the same energy function in AutoDock 3.05 in terms of convergence speed, robustness, and docking energy, especially for highly flexible docking problems
Bo-Fu Liu, Hung-Ming Chen, Hui-Ling Huang, Shiow-Fen Hwang, Shinn-Ying Ho
Congress on Evolutionary Computation3
2003 Mesh optimization for surface approximation using an efficient coarse-to-fine evolutionary algorithm
Hui-Ling Huang, Shinn-Ying Ho
Pattern Recognit.1
2001 Mesh optimization for surface approximation using an efficient coarse-to-fine evolutionary algorithm
abstract
This paper investigates surface approximation using a mesh optimization approach. The mesh optimization problem is how to locate a limited number n of grid points such that the established mesh of n grid points approximates the digital surface of N sample points as closely as possible. The resulting combinatorial problem has an NP-hard search space of C(N, n) instances, i.e., the number of ways of choosing n grid points out of N sample points. A genetic algorithm-based method has been proposed for establishing optimal approximating mesh surfaces. It was shown that the GA-based method is effective in searching the combinatorial space which is intractable when n and N are in the order of thousands. This paper proposes an efficient coarse-to-fine evolutionary algorithm with a novel 2D orthogonal crossover for obtaining an optimal solution to the mesh optimization problem. It is shown empirically that the proposed coarse-to-fine evolutionary algorithm outperforms the existing GA-based method in solving the mesh optimization problem in terms of both approximation quality and convergence speed, especially in solving large mesh optimization problems.
Hui-Ling Huang, Shinn-Ying Ho
CEC1
2001 A general broadcasting scheme for recursive networks with complete connection
Gen-Huey Chen, Shien-Ching Hwang, Hui-Ling Huang, Ming-Yang Su, Dyi-Rong Duh
Parallel Comput.3
2001 Facial modeling from an uncalibrated face image using a coarse-to-fine genetic algorithm
Shinn-Ying Ho, Hui-Ling Huang
Pattern Recognit.2
2001 Facial modeling from an uncalibrated face image using flexible generic parameterized facial models
abstract
The paper presents an optimization approach for facial modeling from an uncalibrated face image using flexible generic parameterized facial models (FGPFMs). An FGPFM consists of a topological structure and geometric knowledge of human faces. The topological description consists of a set of well-designed triangular polygons with a multilayered elastic structure in which the microstructural information can be expressed without complicated facial features. All the geometric values are obtained from a set of training facial models using statistical approaches and genetic algorithms. FGPFM can be easily modified using facial features as FGPFMs parameters to create an accurate specific three-dimensional (3D) facial model from only a photograph of an individual with a yawed face. In addition, the facial modeling problem is formulated as a parameter optimization problem. A hybrid optimization approach based on the Taguchi method and a best-first search algorithm is used to accelerate the search for a near optimal solution. Furthermore, sensitivity analysis and experimental results with texture mapping demonstrate the effectiveness of the proposed approach.
Shinn-Ying Ho, Hui-Ling Huang
IEEE Trans. Syst. Man Cybern. Part B2
2000 Node-disjoint paths in incomplete WK-recursive networks
Ming-Yang Su, Hui-Ling Huang, Gen-Huey Chen, Dyi-Rong Duh
Parallel Comput.2
1999 Combinatorial Properties of Two-Level Hypernet Networks
abstract
The purpose of the paper is to investigate combinatorial properties of the hypernet network. The hypernet network owns two structural advantages: expansibility and equal degree. In addition, it was shown to be efficient in both communication and computation. Since the number of nodes contained in the hypernet network increases very rapidly with expansion level, we emphasize the hypernet network of two levels (denoted by HN(d, 2)) with a practical view. Recently, combinatorial properties such as container (i.e., node-disjoint paths), wide diameter, and fault diameter have received much attention due to their increasing importance and applications in networks. The following results are obtained for HN(d, 2): (1) best containers with width d-1, (2) containers with (maximum) width d, (3) the (d-1)-wide diameter, (4) the d-wide diameter, (5) the (d-2)-fault diameter, and (6) the (d-1)-fault diameter. More specifically, between every two nodes of HN(d, 2), d (or d-1) packets can be transmitted simultaneously with at most D+2 (or D+1) parallel steps, where D=2d+1 is the diameter of HN(d, 2). Besides, the diameter of HN(d, 2) will increase by at most two (or one), if there are at most d-1 (or d-2) node faults. Our results reveal that HN(d, 2) is not only efficient in parallel transmission, but robust in fault tolerance.
Hui-Ling Huang, Gen-Huey Chen
IEEE Trans. Parallel Distributed Syst.1
1998 Topological properties and algorithms for two-level hypernet networks
abstract
Although many networks have been proposed as the topology of a large-scale parallel and distributed system, most of them are neither expansible nor of equal degree. A network with these two properties will gain the advantages of easy implementation and low cost when it is manufactured. The hypernet, which was proposed by Hwang and Ghosh, represents a family of recursively scalable networks that are both expansible and of equal degree. In addition to the two merits, the hypernet has proven efficient for communication and computation. But, unfortunately, most topological properties and the problem of shortest-path routing for the hypernet are still unsolved. The reason is that the structure of the hypernet is complex and asymmetric, and, especially, no mathematical description was given before. In this paper, considering current hardware restrictions, we concentrate our effort on the hypernet of moderate size. We first give a concise mathematical definition for the hypernet and then solve the following problems for the hypernet of two levels: (1) shortest-path routing, (2) diameter, (3) connectivity, (4) minimum-height spanning trees, and (5) embedding of rings, tori, and hypercubes. © 1998 John Wiley & Sons, Inc. Networks 31:105–118, 1998
Hui-Ling Huang, Gen-Huey Chen
Networks1
1998 An analytic solution for the pose determination of human faces from a monocular image
Shinn-Ying Ho, Hui-Ling Huang
Pattern Recognit. Lett.2