VLDB 2026 Research / reviewers in the wild / expert
Cheng-Hong Yang 0001
dblp:18/6168-1
· DBLP profile ↗
57ranked-venue papers
29as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 31 · 19 first-author · 4 since 2021Artificial intelligence and machine learning · 24 · 10 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AKBs & MKBs: Knowledge-based systems to predict breast cancer mortality
Cheng-Hong Yang 0001, Sin-Hua Moi, Ming-Feng Hou, Li-Yeh Chuang, Yu-Da Lin |
Knowl. Based Syst. | 1 |
| 2025 | Shoreline change prediction along the Cijin coastline of Taiwan using deep learning and satellite imagery
Li-Hung Tsai, Chih-Hsien Wu, Qin-Sen Zhang, Jen-Chung Shao, Chih-Min Hsieh, Cheng-Hong Yang 0001 |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | An autoencoder-based arithmetic optimization clustering algorithm to enhance principal component analysis to study the relations between industrial market stock indices in real estate
Cheng-Hong Yang 0001, Borcy Lee, Yi-In Lee, Yu-Fang Chung, Yu-Da Lin |
Expert Syst. Appl. | 1 |
| 2025 | An Information Fusion System-Driven Deep Neural Networks With Application to Cancer Mortality Risk EstimateabstractNext-generation sequencing (NGS) genomic data offer valuable high-throughput genomic information for computational applications in medicine. Using genomic data to identify disease-associated genes to estimate cancer mortality risk remains challenging regarding to computational efficiency and risk integration. For determining mortality-related genes, we propose an information fusion system based on a fuzzy system to fuse the numerous deep-learning-based risk scores, consider the significance of features related to time-varying effects and risk stratifications, and interpret the directional relationship and interaction between outcome and predictors. Fuzzy rules were implemented to integrate the considerations mentioned above by merging all the risk score models to achieve advanced risk estimation. The genomic data of head and neck squamous cell carcinoma (HNSCC) were used to evaluate the performance of the proposed computational approach. The results indicated that the proposed computational approach exhibited optimal ability to identify mortality risk-related genes in HNSCC patients. The results also suggest that HNSCC mortality is associated with cancer inflammatory response, the interleukin-17A signaling pathway, stellate cell activation, and the extracellular-regulated protein kinase five signaling pathway, which might offer new therapeutic targets HNSCC through immunologic or antiangiogenic mechanisms. The proposed information fusion system can promote the determination of high-risk genes related to cancer mortality. This study contributes a valid cancer mortality risk estimate that can identify mortality-related genes. Cheng-Hong Yang 0001, Sin-Hua Moi, Li-Yeh Chuang, Yu-Da Lin |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Dimensionality reduction approach for many-objective epistasis analysisabstractIn epistasis analysis, single-nucleotide polymorphism-single-nucleotide polymorphism interactions (SSIs) among genes may, alongside other environmental factors, influence the risk of multifactorial diseases. To identify SSI between cases and controls (i.e. binary traits), the score for model quality is affected by different objective functions (i.e. measurements) because of potential disease model preferences and disease complexities. Our previous study proposed a multiobjective approach-based multifactor dimensionality reduction (MOMDR), with the results indicating that two objective functions could enhance SSI identification with weak marginal effects. However, SSI identification using MOMDR remains a challenge because the optimal measure combination of objective functions has yet to be investigated. This study extended MOMDR to the many-objective version (i.e. many-objective MDR, MaODR) by integrating various disease probability measures based on a two-way contingency table to improve the identification of SSI between cases and controls. We introduced an objective function selection approach to determine the optimal measure combination in MaODR among 10 well-known measures. In total, 6 disease models with and 40 disease models without marginal effects were used to evaluate the general algorithms, namely those based on multifactor dimensionality reduction, MOMDR and MaODR. Our results revealed that the MaODR-based three objective function model, correct classification rate, likelihood ratio and normalized mutual information (MaODR-CLN) exhibited the higher 6.47% detection success rates (Accuracy) than MOMDR and higher 17.23% detection success rates than MDR through the application of an objective function selection approach. In a Wellcome Trust Case Control Consortium, MaODR-CLN successfully identified the significant SSIs (P < 0.001) associated with coronary artery disease. We performed a systematic analysis to identify the optimal measure combination in MaODR among 10 objective functions. Our combination detected SSIs-based binary traits with weak marginal effects and thus reduced spurious variables in the score model. MOAI is freely available at https://sites.google.com/view/maodr/home. Cheng-Hong Yang 0001, Ming-Feng Hou, Li-Yeh Chuang, Cheng-San Yang, Yu-Da Lin |
Briefings Bioinform. | 1 |
| 2023 | Export- and import-based economic models for predicting global trade using deep learning
Cheng-Hong Yang 0001, Cheng-Feng Lee, Po-Yin Chang |
Expert Syst. Appl. | 1 |
| 2023 | Fuzzy-Based Multiobjective Multifactor Dimensionality Reduction for Epistasis AnalysisabstractEpistasis detection is vital for understanding disease susceptibility in genetics. Multiobjective multifactor dimensionality reduction (MOMDR) was previously proposed to detect epistasis. MOMDR was performed using binary classification to distinguish the high-risk (H) and low-risk (L) groups to reduce multifactor dimensionality. However, the binary classification does not reflect the uncertainty of the H and L classification. In this study, we proposed an empirical fuzzy MOMDR (EFMOMDR) to address the limitations of binary classification using the degree of membership through an empirical fuzzy approach. The EFMOMDR can simultaneously consider two incorporated fuzzy-based measures, including correct classification rate and likelihood rate, and does not require parameter tuning. Simulation studies revealed that EFMOMDR has higher 7.14% detection success rates than MOMDR, indicating that the limitations of binary classification of MOMDR have been successfully improved by empirical fuzzy. Moreover, EFMOMDR was used to analyze coronary artery disease in the Wellcome Trust Case Control Consortium dataset. Cheng-Hong Yang 0001, Hsiu-Chen Huang, Ming-Feng Hou, Li-Yeh Chuang, Yu-Da Lin |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2022 | Multiobjective optimization-driven primer design mechanism: towards user-specified parameters of PCR primerabstractPrimers are critical for polymerase chain reaction (PCR) and influence PCR experimental outcomes. Designing numerous combinations of forward and reverse primers involves various primer constraints, posing a computational challenge. Most PCR primer design methods limit parameters because the available algorithms use general fitness functions. This study designed new fitness functions based on user-specified parameters and used the functions in a primer design approach based on the multiobjective particle swarm optimization (MOPSO) algorithm to address the challenge of primer design with user-specified parameters. Multicriteria evaluation was conducted simultaneously based on primer constraints. The fitness functions were evaluated using 7425 DNA sequences and compared with a predominant primer design approach based on optimization algorithms. Each DNA sequence was run 100 times to calculate the difference between the user-specified parameters and primer constraint values. The algorithms based on fitness functions with user-specified parameters outperformed the algorithms based on general fitness functions for 11 primer constraints. Moreover, MOPSO exhibited superior implementation in all experiments. Practical gel electrophoresis was conducted to verify the PCR experiments and established that MOPSO effectively designs primers based on user-specified parameters. Cheng-Hong Yang 0001, Yu-Huei Cheng, Li-Yeh Chuang, Yu-Da Lin |
Briefings Bioinform. | 1 |
| 2022 | DeepBarcoding: Deep Learning for Species Classification Using DNA BarcodingabstractDNA barcodes with short sequence fragments are used for species identification. Because of advances in sequencing technologies, DNA barcodes have gradually been emphasized. DNA sequences from different organisms are easily and rapidly acquired. Therefore, DNA sequence analysis tools play an increasingly crucial role in species identification. This study proposed deep barcoding, a deep learning framework for species classification by using DNA barcodes. Deep barcoding uses raw sequence data as the input to represent one-hot encoding as a one-dimensional image and uses a deep convolutional neural network with a fully connected deep neural network for sequence analysis. It can achieve an average accuracy of >90 percent for both simulation and real datasets. Although deep learning yields outstanding performance for species classification with DNA sequences, its application remains a challenge. The deep barcoding model can be a potential tool for species classification and can elucidate DNA barcode-based species identification. Cheng-Hong Yang 0001, Kuo-Chuan Wu, Li-Yeh Chuang, Hsueh-Wei Chang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2021 | Effective multinational trade forecasting using LSTM recurrent neural network
Mei-Li Shen, Cheng-Feng Lee, Hsiou-Hsiang Liu, Po-Yin Chang, Cheng-Hong Yang 0001 |
Expert Syst. Appl. | 5 |
| 2021 | Applications of Deep Learning and Fuzzy Systems to Detect Cancer Mortality in Next-Generation Genomic DataabstractIn the era of advanced precision medicine, next-generation genomic data are crucial to achieve breakthroughs in cancer medicine. Effective cancer mortality risk estimation for genomic data associated with cancer remains a vital challenge. The combination of machine learning algorithms and conventional survival analysis can advance the detection of high-risk missense mutation variants and candidate genes associated with cancer mortality in next-generation genomic data. In this article, a fuzzy logic system combined with machine learning algorithms and conventional survival analysis named FuzzyDeepCoxPH was proposed to identify high-risk missense mutation variants and candidate genes highly associated with cancer mortality. DL-derived abstracted weights and Cox proportional hazards (CoxPH) ratios were used to develop four model-based risk scores to consider the factor importance associated with risk stratification, time-varying effects, and individual and interaction effects among features. Fuzzy rules based on a fuzzy logic system were designed to integrate these considerations by merging four model-based risk scores to develop advanced risk estimation. The clinical features and next-generation sequencing of deoxyribonucleic acid and ribonucleic acid genomic data were used to evaluate FuzzyDeepCoxPH performance. The results indicated that FuzzyDeepCoxPH can effectively distinguish high-risk variants and candidate genes related to cancer mortality. In FuzzyDeepCoxPH, the fuzzy logic system was applied to combine DL-based and CoxPH-based models to provide a comprehensive cancer mortality risk estimation for cancer medicine. Cheng-Hong Yang 0001, Sin-Hua Moi, Ming-Feng Hou, Li-Yeh Chuang, Yu-Da Lin |
IEEE Trans. Fuzzy Syst. | 1 |
| 2020 | Identification of Kidney Clear Cell Carcinoma Mortality Risk-Associated Gene Mutation by Using a Random Survival Forest ApproachabstractKidney clear cell carcinoma is commonly characterized by poor prognosis, which is associated with the function and differential expression of specific genes. The combination of a typical statistical survival model and nonparametric random forest algorithm provides a more precise estimation for the association between gene mutations and all-cause mortality risk. This study identifies mortality risk-associated gene mutations in kidney clear cell carcinoma by using a random survival forest algorithm. Of 22 candidate genes, VHL (variable importance [VIMP] = 0.097), EDIL3 (VIMP = 0.037), PBRM1 (VIMP = 0.027), PTEN (VIMP = 0.012), BAP1 (VIMP = 0.010), and HMGN5 (VIMP = 0.002) were selected and used to develop a dichotomous risk model for all-cause mortality by using the estimated risk threshold. The high-risk group exhibited a relatively poor survival rate than did the low-risk group (95.5% vs. 93.0%). In conclusion, this study provides a simple dichotomous model for mutation risk, according to the gene mutation risk threshold, by using a random survival forest model. For the gene mutation risk model, VHL, EDIL3, PBRM1, PTEN, BAP1, and HMGN5 were selected to effectively determine the effects of gene mutation on all-cause mortality from kidney clear cell carcinoma. Cheng-Hong Yang 0001, Yin-Syuan Chen, Sin-Hua Moi, Li-Yeh Chuang, Yu-Da Lin |
BIBE | 1 |
| 2020 | New Evaluation Measures for Multifactor Dimensionality Reduction in SNP-SNP Interaction AnalysisabstractStudies have proven that single nucleotide polymorphism (SNP)-SNP interaction detection is helpful for understanding the susceptibility of an individual to genetic diseases. Although multifactor dimensionality reduction (MDR) is an effective SNP-SNP interaction detection algorithm, the mechanism of SNP-SNP interaction detection based on MDR contingency tables has not been widely studied. In this study, we propose a multi-objective MDR to detect SNP-SNP interactions. In the proposed multi-objective MDR, multiple measures can be simultaneously considered for detecting epistatic interactions. Then, set theory is used to select the best epistatic interactions in k-fold cross-validation to achieve high identification accuracy for SNP-SNP interactions. Two MDR parameters, namely the correct classification rate (CCR) and predictive summary index (PSI), were used for evaluating the algorithms. The results revealed that the detection success rates of multi-objective MDR were higher than those of other MDR-based algorithms in identifying epistatic interactions. Based on the CCR and PSI, our study demonstrated that the proposed multi-objective MDR can effectively detect SNP-SNP interactions. Cheng-Hong Yang 0001, Sin-Hua Moi, Li-Yeh Chuang, Yu-Da Lin |
BIBE | 1 |
| 2020 | An improved fuzzy set-based multifactor dimensionality reduction for detecting epistasis
Cheng-Hong Yang 0001, Li-Yeh Chuang, Yu-Da Lin |
Artif. Intell. Medicine | 1 |
| 2020 | Class Balanced Multifactor Dimensionality Reduction to Detect Gene-Gene InteractionsabstractDetecting gene-gene interactions in single-nucleotide polymorphism data is vital for understanding disease susceptibility. However, existing approaches may be limited by the sample size in case-control studies. Herein, we propose a balance approach for the multifactor dimensionality reduction (BMDR) method to increase the accuracy of estimates of the prediction error rate in small samples. BMDR explicitly selects the best model by evaluating the average of prediction error rates over k-fold cross-validation without cross-validation consistency selection. In this study, we used several epistatic models with and without marginal effects under different parameter settings (heritability and minor allele frequencies) to evaluate the performance of existing approaches. Using simulated data sets, BMDR successfully detected gene-gene interactions, particularly for data sets with small sample sizes. A large data set was obtained from the Wellcome Trust Case Control Consortium, and results indicated that BMDR could effectively detect significant gene-gene interactions. Cheng-Hong Yang 0001, Yu-Da Lin, Li-Yeh Chuang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2020 | Epistasis Analysis Using an Improved Fuzzy C-Means-Based Entropy ApproachabstractEpistasis detection is vital to determining disease susceptibility in the human genome. With rapid advances in technology, multifactor dimensionality reduction (MDR) has become an effective algorithm for epistasis detection. Classification of high-risk (H) and low-risk (L) groups in MDR operations is a key topic, but it has not been thoroughly investigated. In this paper, we propose an improved fuzzy c-means-based entropy (FCME) approach to address the limitations of binary classification. For this approach, the degree of membership in MDR, referred to as FCMEMDR, was used. The FCME approach and MDR measure were integrated to enable more precise differentiation between similar frequencies of multifactor genotypes in the cases of possible epistasis. We used the MDR measures of correct classification rate and likelihood ratio. Numerous simulated datasets were applied, and the experimental results revealed two measures of FCMEMDR with higher detection rates than those of other MDR-based algorithms. Our analysis of binary and fuzzy classifications in MDR operations may offer insights into the problem of uncertainty in H/L classification. Two measures of FCMEMDR detected significant instances of epistasis associated with coronary artery disease in the Wellcome Trust Case Control Consortium dataset. FCMEMDR is freely available at https://gitlab.com/yudalinemail/fcmemdr. Cheng-Hong Yang 0001, Li-Yeh Chuang, Yu-Da Lin |
IEEE Trans. Fuzzy Syst. | 1 |
| 2019 | Multiple-Criteria Decision Analysis-Based Multifactor Dimensionality Reduction for Detecting Gene-Gene InteractionsabstractGene-gene interactions (GGIs) are important markers for determining susceptibility to a disease. Multifactor dimensionality reduction (MDR) is a popular algorithm for detecting GGIs and primarily adopts the correct classification rate (CCR) to assess the quality of a GGI. However, CCR measurement alone may not successfully detect certain GGIs because of potential model preferences and disease complexities. In this study, multiple-criteria decision analysis (MCDA) based on MDR was named MCDA-MDR and proposed for detecting GGIs. MCDA facilitates MDR to simultaneously adopt multiple measures within the two-way contingency table of MDR to assess GGIs; the CCR and rule utility measure were employed. Cross-validation consistency was adopted to determine the most favorable GGIs among the Pareto sets. Simulation studies were conducted to compare the detection success rates of the MDR-only-based measure and MCDA-MDR, revealing that MCDA-MDR had superior detection success rates. The Wellcome Trust Case Control Consortium dataset was analyzed using MCDA-MDR to detect GGIs associated with coronary artery disease, and MCDA-MDR successfully detected numerous significant GGIs (p < 0.001). MCDA-MDR performance assessment revealed that the applied MCDA successfully enhanced the GGI detection success rate of the MDR-based method compared with MDR alone. Cheng-Hong Yang 0001, Yu-Da Lin, Li-Yeh Chuang |
IEEE J. Biomed. Health Informatics | 1 |
| 2018 | Improved Multifactor Dimensionality Reduction for Epistasis DetectionabstractEpistasis detection facilitates determining susceptibility to disease. Multifactor dimensionality reduction (MDR) and multiobjective MDR (MOMDR) were proposed for epistasis detection. However, more measures must be investigated for MOMDR. In this study, we incorporated the Youden index (YI) and correct classification rate (CCR) into MOMDR (MOMDR-YC) for epistasis detection. Simulations were conducted to compare MDR-based YI (MDR-Y), MDR-based CCR (MDR-C), and MOMDR-YC. Moreover, the detection success rates of the three approaches are presented. MOMDR-YC revealed that the YI and CCR measures can enhance the detection success rates of MDR. The simulation results revealed that epistasis could be successfully detected by incorporating YI and CCR into MOMDR. Li-Yeh Chuang, Cheng-Hong Yang 0001, Yu-Da Lin |
BIBE | 2 |
| 2018 | Decision Theory-Based DNA Barcoding Through Quick Response Code RepresentationabstractDNA barcoding is widely used in fields, such as taxonomy and species identification. Conventional DNA barcoding sequences employ uninformative or repeat nucleotides in known groups of taxa within a monophylum. Herein, we propose a decision theory-based DNA barcode that tests for the ribulose bisphosphate carboxylase gene (rbcL). The proposed method can generate shorter DNA barcodes called single nucleotide polymorphism (SNP) tags, which shorten rbcL sequences from their full length (400-654 bp) to 25-bp DNA tags. These DNA tags are then represented by quick response (QR) codes containing the species names, accession numbers, and DNA tag sequences. Our proposed method can efficiently reduce data storage and provide DNA barcoding for various plant species. Cheng-Hong Yang 0001, Kuo-Chuan Wu, Hsueh-Wei Chang, Li-Yeh Chuang |
BIBE | 1 |
| 2018 | Multiobjective multifactor dimensionality reduction to detect SNP-SNP interactionsabstractMotivation: Single-nucleotide polymorphism (SNP)-SNP interactions (SSIs) are popular markers for understanding disease susceptibility. Multifactor dimensionality reduction (MDR) can successfully detect considerable SSIs. Currently, MDR-based methods mainly adopt a single-objective function (a single measure based on contingency tables) to detect SSIs. However, generally, a single-measure function might not yield favorable results due to potential model preferences and disease complexities. Approach: This study proposes a multiobjective MDR (MOMDR) method that is based on a contingency table of MDR as an objective function. MOMDR considers the incorporated measures, including correct classification and likelihood rates, to detect SSIs and adopts set theory to predict the most favorable SSIs with cross-validation consistency. MOMDR enables simultaneously using multiple measures to determine potential SSIs. Results: Three simulation studies were conducted to compare the detection success rates of MOMDR and single-objective MDR (SOMDR), revealing that MOMDR had higher detection success rates than SOMDR. Furthermore, the Wellcome Trust Case Control Consortium dataset was analyzed by MOMDR to detect SSIs associated with coronary artery disease. Availability and implementation: MOMDR is freely available at https://goo.gl/M8dpDg. Supplementary information: Supplementary data are available at Bioinformatics online. Cheng-Hong Yang 0001, Li-Yeh Chuang, Yu-Da Lin |
Bioinform. | 1 |
| 2017 | CMDR based differential evolution identifies the epistatic interaction in genome-wide association studiesabstractMOTIVATION: Detecting epistatic interactions in genome-wide association studies (GWAS) is a computational challenge. Such huge numbers of single-nucleotide polymorphism (SNP) combinations limit the some of the powerful algorithms to be applied to detect the potential epistasis in large-scale SNP datasets. APPROACH: We propose a new algorithm which combines the differential evolution (DE) algorithm with a classification based multifactor-dimensionality reduction (CMDR), termed DECMDR. DECMDR uses the CMDR as a fitness measure to evaluate values of solutions in DE process for scanning the potential statistical epistasis in GWAS. RESULTS: The results indicated that DECMDR outperforms the existing algorithms in terms of detection success rate by the large simulation and real data obtained from the Wellcome Trust Case Control Consortium. For running time comparison, DECMDR can efficient to apply the CMDR to detect the significant association between cases and controls amongst all possible SNP combinations in GWAS. AVAILABILITY AND IMPLEMENTATION: DECMDR is freely available at https://goo.gl/p9sLuJ . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Cheng-Hong Yang 0001, Li-Yeh Chuang, Yu-Da Lin |
Bioinform. | 1 |
| 2016 | A comparative analysis of chaotic particle swarm optimizations for detecting single nucleotide polymorphism barcodes
Li-Yeh Chuang, Sin-Hua Moi, Yu-Da Lin, Cheng-Hong Yang 0001 |
Artif. Intell. Medicine | 4 |
| 2016 | Analysis of high-order SNP barcodes in mitochondrial D-loop for chronic dialysis susceptibility
Cheng-Hong Yang 0001, Yu-Da Lin, Li-Yeh Chuang, Hsueh-Wei Chang |
J. Biomed. Informatics | 1 |
| 2015 | An improved GA for identifying susceptibility genes in the presence of epistasisabstractIdentifying the epistasis models between single nucleotide polymorphisms (SNPs) in several genes can explain the susceptibility to diseases. The statistical methods have been used to identify the significant epistasis models according to the related statistical values, including odds ratio (OR), chi-square test (χ2), p-value, etc. However, the high calculations limit the statistic to identify the high-order epistasis. In this study, we proposed an lsGA algorithm, genetic algorithm based on local search algorithm, to identify the significant epistasis model amongst the large SNP combinations. Two disease models were used to simulate the large data sets considering the minor allele frequency (MAF), number of SNP, and number of sample. The 3-order epistasis models were identified by chi-square test (χ2) for evaluating the significance (P-value <; 0.05). lsGA was compared with GA to analyze the improvement in the search abilities, and results showed that lsGA provided higher chi-square test values than that of GA. Jyh-Ferng Yang, Yu-Da Lin, Li-Yeh Chuang, Cheng-Hong Yang 0001 |
CEC | 4 |
| 2013 | Drug-SNPing: an integrated drug-based, protein interaction-based tagSNP-based pharmacogenomics platform for SNP genotypingabstractMany drug or single nucleotide polymorphism (SNP)-related resources and tools have been developed, but connecting and integrating them is still a challenge. Here, we describe a user-friendly web-based software package, named Drug-SNPing, which provides a platform for the integration of drug information (DrugBank and PharmGKB), protein-protein interactions (STRING), tagSNP selection (HapMap) and genotyping information (dbSNP, REBASE and SNP500Cancer). DrugBank-based inputs include the following: (i) common name of the drug, (ii) synonym or drug brand name, (iii) gene name (HUGO) and (iv) keywords. PharmGKB-based inputs include the following: (i) gene name (HUGO), (ii) drug name and (iii) disease-related keywords. The output provides drug-related information, metabolizing enzymes and drug targets, as well as protein-protein interaction data. Importantly, tagSNPs of the selected genes are retrieved for genotyping analyses. All drug-based and protein-protein interaction-based SNP genotyping information are provided with PCR-RFLP (PCR-restriction enzyme length polymorphism) and TaqMan probes. Thus, users can enter any drug keywords/brand names to obtain immediate information that is highly relevant to genotyping for pharmacogenomics research. Cheng-Hong Yang 0001, Yu-Huei Cheng, Li-Yeh Chuang, Hsueh-Wei Chang |
Bioinform. | 1 |
| 2013 | Operon Prediction Using Chaos Embedded Particle Swarm OptimizationabstractOperons contain valuable information for drug design and determining protein functions. Genes within an operon are co-transcribed to a single-strand mRNA and must be coregulated. The identification of operons is, thus, critical for a detailed understanding of the gene regulations. However, currently used experimental methods for operon detection are generally difficult to implement and time consuming. In this paper, we propose a chaotic binary particle swarm optimization (CBPSO) to predict operons in bacterial genomes. The intergenic distance, participation in the same metabolic pathway and the cluster of orthologous groups (COG) properties of the Escherichia coli genome are used to design a fitness function. Furthermore, the Bacillus subtilis, Pseudomonas aeruginosa PA01, Staphylococcus aureus and Mycobacterium tuberculosis genomes are tested and evaluated for accuracy, sensitivity, and specificity. The computational results indicate that the proposed method works effectively in terms of enhancing the performance of the operon prediction. The proposed method also achieved a good balance between sensitivity and specificity when compared to methods from the literature. Li-Yeh Chuang, Cheng-Huei Yang, Jui-Hung Tsai, Cheng-Hong Yang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2013 | Evaluation of Breast Cancer Susceptibility Using Improved Genetic Algorithms to Generate Genotype SNP BarcodesabstractGenetic association is a challenging task for the identification and characterization of genes that increase the susceptibility to common complex multifactorial diseases. To fully execute genetic studies of complex diseases, modern geneticists face the challenge of detecting interactions between loci. A genetic algorithm (GA) is developed to detect the association of genotype frequencies of cancer cases and noncancer cases based on statistical analysis. An improved genetic algorithm (IGA) is proposed to improve the reliability of the GA method for high-dimensional SNP-SNP interactions. The strategy offers the top five results to the random population process, in which they guide the GA toward a significant search course. The IGA increases the likelihood of quickly detecting the maximum ratio difference between cancer cases and noncancer cases. The study systematically evaluates the joint effect of 23 SNP combinations of six steroid hormone metabolisms, and signaling-related genes involved in breast carcinogenesis pathways were systematically evaluated, with IGA successfully detecting significant ratio differences between breast cancer cases and noncancer cases. The possible breast cancer risks were subsequently analyzed by odds-ratio (OR) and risk-ratio analysis. The estimated OR of the best SNP barcode is significantly higher than 1 (between 1.15 and 7.01) for specific combinations of two to 13 SNPs. Analysis results support that the IGA provides higher ratio difference values than the GA between breast cancer cases and noncancer cases over 3-SNP to 13-SNP interactions. A more specific SNP-SNP interaction profile for the risk of breast cancer is also provided. Cheng-Hong Yang 0001, Yu-Da Lin, Li-Yeh Chuang, Hsueh-Wei Chang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2012 | A Quantum Genetic Algorithm for Operon PredictionabstractOperon is a fundamental unit of transcription which is usually used to understand gene regulations and functions in entire genomes. Detecting operon experimentally is difficult and time-consuming, thus many bioinformatics algorithms have been proposed to predict operon. In this paper, we use an improved discrete genetic algorithm based on quantum theory for operon prediction. It is simpler and more powerful than the algorithms available, and thus avoids local optima while searching for a better solution. We utilize intergenic distance, participation in the same metabolic pathway and cluster of orthologous groups (COG) gene functions to design fitness function base on reward and penalty (RP). The RP operation can improve fitness value of chromosome in proportion to the accuracy. Experimental results show that the detection accuracy of our method reached 0.872, 0.925, 0.943, 0.954 and 0.926 respectively for the E. coli, B. subtilis, P. aeruginosa PA01, S. aureus and M. tuberculosis genomes. Results demonstrate that our proposed method can predict operons with high accuracy. Li-Yeh Chuang, Cheng-Yi Chiang, Cheng-Hong Yang 0001 |
AINA | 3 |
| 2012 | Chaos Embedded Particle Swarm Optimization for Tag Single Nucleotide Polymorphism SelectionabstractSingle Nucleotide Polymorphisms (SNPs) are the most common variants in the human genome. Disease analysis costs can be reduced by selecting meaningful SNPs, i.e., tagging the SNP selection. We propose a method, called chaos particle swarm optimization (CPSO), to select tag SNPs, and use linkage disequilibrium (LD) and the K-nearest neighbor (K-NN) method to respectively reduce and evaluate the tag SNPs. To measure the quality of the correction rate and the tag SNPs number, the Hap Map database was used to test CPSO's ability and to compare the proposed method with other methods. The results indicate that the proposed method is effectively to enhance the tag SNP prediction in terms of the result achieves a good accuracy when compared to methods from the literature. Li-Yeh Chuang, Li-Wei Huang, Cheng-Hong Yang 0001 |
AINA | 3 |
| 2012 | Complementary distribution BPSO for feature selectionabstractFeature selection is a preprocessing technique in the field of data analysis, which is used to reduce the number of features by removing irrelevant, noisy, and redundant data, thus resulting in acceptable classification accuracy. This process constit Li-Yeh Chuang, Cheng-Hong Yang 0001, Sheng-Wei Tsai |
Intell. Data Anal. | 2 |
| 2012 | Mutagenic Primer Design for Mismatch PCR-RFLP SNP Genotyping Using a Genetic AlgorithmabstractPolymerase chain reaction-restriction fragment length polymorphism (PCR-RFLP) is useful in small-scale basic research studies of complex genetic diseases that are associated with single nucleotide polymorphism (SNP). Designing a feasible primer pair is an important work before performing PCR-RFLP for SNP genotyping. However, in many cases, restriction enzymes to discriminate the target SNP resulting in the primer design is not applicable. A mutagenic primer is introduced to solve this problem. GA-based Mismatch PCR-RFLP Primers Design (GAMPD) provides a method that uses a genetic algorithm to search for optimal mutagenic primers and available restriction enzymes from REBASE. In order to improve the efficiency of the proposed method, a mutagenic matrix is employed to judge whether a hypothetical mutagenic primer can discriminate the target SNP by digestion with available restriction enzymes. The available restriction enzymes for the target SNP are mined by the updated core of SNP-RFLPing. GAMPD has been used to simulate the SNPs in the human SLC6A4 gene under different parameter settings and compared with SNP Cutter for mismatch PCR-RFLP primer design. The in silico simulation of the proposed GAMPD program showed that it designs mismatch PCR-RFLP primers. The GAMPD program is implemented in JAVA and is freely available at http://bio.kuas.edu.tw/gampd/. Cheng-Hong Yang 0001, Yu-Huei Cheng, Cheng-Huei Yang, Li-Yeh Chuang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2011 | An Improved Natural PCR-RFLP Primer Design MethodabstractThe polymerase chain reaction restriction fragment length polymorphism (PCR-RFLP) technique is often used in laboratories and many basic research studies of complex genetic diseases associated with single nucleotide polymorphisms (SNP). When performing PCR-RFLP for SNP genotyping, feasible primer pairs are subject to numerous constraints and require a restriction enzyme for discriminating the target SNP. In this study, we develop a method for natural PCR-RFLP primer design for SNP genotyping using a particle swarm optimization algorithm. The in silico simulation with SNPs of the SLC6A4 gene demonstrates that this method reliably produces designs for natural PCR-RFLP primers which best fit the common primer constraints and also identifies available restriction enzymes. Li-Yeh Chuang, Yu-Da Lin, Hsueh-Wei Chang, Cheng-Hong Yang 0001 |
BIBE | 4 |
| 2011 | Analysis of SNP Interaction Combinations to Determine Breast Cancer Risk with PSOabstractMany association studies analyze the genotype frequencies of case and control data to predict susceptibility to diseases and cancers. An increasing number of studies has shown that the risk of getting diseases and cancers is associated with the co-occurrence of some contain single nucleotide polymorphisms (SNPs). Determining the disease-causing SNPs has become an important objective. In order to study the SNP-SNP interaction in breast cancer, we used a particle swarm optimization (PSO) algorithm to compute the difference between the control and case data and performed a feature selection from different SNP combinations with their corresponding genotypes. The best combination of SNP-SNP interactions is the maximal difference of co-occurrences between the control and case groups. In this study, we explored the SNP interaction of 19 SNPs in 372 controls and 398 cases of breast cancer association using simulated SNP data of breast cancers. The odds ratio (OR) were used to evaluate the breast cancer risk in terms of the best combination of SNP-SNP interactions. Compared to their corresponding non-SNP combinations, the estimated OR of the best predicted SNP combination with genotypes for breast cancer is significantly greater than 1 (about 1.771 and 2.417; confidence interval (CI): 1.223-4.371; p <; 0.05-0.001) for specific SNP combinations of two to five SNPs. The SNP interaction associated with a high risk of breast cancer could be successfully predicted using the proposed PSO method. Li-Yeh Chuang, Ming-Cheng Lin, Hsueh-Wei Chang, Cheng-Hong Yang 0001 |
BIBE | 4 |
| 2011 | Particle Swarm Optimization with Extremal Optimization for the Prediction of CpG Islands in the Mammal GenomeabstractRegions with abundant GC nucleotides in a genome, which are often referred to as CpG islands, have been used in methylation analysis and the prediction of promoter regions. In this study, we propose PSOEO (Particle Swarm Optimization with Extremal Optimization), a method for the prediction of CpG islands in the mammal genome. This method adopts the GGF criteria (GC content ≥ 50%, observed/expected (O/E) ratio ≥0.6 and length ≥200 bp) for the search of CpG islands. First, we used the PSO algorithm to predict CpG islands. In a second stage, we used EO to search for various output states (local search) in order to find a better result. Extremal optimization is a developed heuristic local search method. Finally, we used five evaluation criteria, namely the sensitivity (SN), specificity (SP), accuracy (ACC), correlation coefficient (CC) and performance coefficient (PC) to compare other methods in the literature. PSOEO method provided better SN and CC predictions for the locations of CpG islands than the other methods it was compared to. Li-Yeh Chuang, Ming-Cheng Lin, Cheng-Hong Yang 0001 |
BIBE | 3 |
| 2011 | Chaotic particle swarm optimization for data clustering
Li-Yeh Chuang, Chih-Jen Hsiao, Cheng-Hong Yang 0001 |
Expert Syst. Appl. | 3 |
| 2011 | Improved binary particle swarm optimization using catfish effect for feature selection
Li-Yeh Chuang, Sheng-Wei Tsai, Cheng-Hong Yang 0001 |
Expert Syst. Appl. | 3 |
| 2011 | Gene selection and classification using Taguchi chaotic binary particle swarm optimization
Li-Yeh Chuang, Cheng-San Yang, Kuo-Chuan Wu, Cheng-Hong Yang 0001 |
Expert Syst. Appl. | 4 |
| 2010 | Confronting Two-Pair Primer Design Using Particle Swarm Optimization
Cheng-Hong Yang 0001, Yu-Huei Cheng, Li-Yeh Chuang |
ICCCI (3) | 1 |
| 2010 | PPO: Predictor for Prokaryotic OperonsabstractSUMMARY: We present an operon predictor for prokaryotic operons (PPO), which can predict operons in the entire prokaryotic genome. The prediction algorithm used in PPO allows the user to select binary particle swarm optimization (BPSO), a genetic algorithm (GA) or some other methods introduced in the literature to predict operons. The operon predictor on our web server and the provided database are easy to access and use. The main features offered are: (i) selection of the prediction algorithm; (ii) adjustable parameter settings of the prediction algorithm; (iii) graphic visualization of results; (iv) integrated database queries; (v) listing of experimentally verified operons; and (vi) related tools. AVAILABILITY AND IMPLEMENTATION: PPO is freely available at http://bio.kuas.edu.tw/PPO/. Li-Yeh Chuang, Jui-Hung Tsai, Cheng-Hong Yang 0001 |
Bioinform. | 3 |
| 2010 | SNP-RFLPing 2: an updated and integrated PCR-RFLP tool for SNP genotypingabstractBACKGROUND: PCR-restriction fragment length polymorphism (RFLP) assay is a cost-effective method for SNP genotyping and mutation detection, but the manual mining for restriction enzyme sites is challenging and cumbersome. Three years after we constructed SNP-RFLPing, a freely accessible database and analysis tool for restriction enzyme mining of SNPs, significant improvements over the 2006 version have been made and incorporated into the latest version, SNP-RFLPing 2. RESULTS: The primary aim of SNP-RFLPing 2 is to provide comprehensive PCR-RFLP information with multiple functionality about SNPs, such as SNP retrieval to multiple species, different polymorphism types (bi-allelic, tri-allelic, tetra-allelic or indels), gene-centric searching, HapMap tagSNPs, gene ontology-based searching, miRNAs, and SNP500Cancer. The RFLP restriction enzymes and the corresponding PCR primers for the natural and mutagenic types of each SNP are simultaneously analyzed. All the RFLP restriction enzyme prices are also provided to aid selection. Furthermore, the previously encountered updating problems for most SNP related databases are resolved by an on-line retrieval system. CONCLUSIONS: The user interfaces for functional SNP analyses have been substantially improved and integrated. SNP-RFLPing 2 offers a new and user-friendly interface for RFLP genotyping that can be used in association studies and is freely available at http://bio.kuas.edu.tw/snp-rflping2. Hsueh-Wei Chang, Yu-Huei Cheng, Li-Yeh Chuang, Cheng-Hong Yang 0001 |
BMC Bioinform. | 4 |
| 2010 | Confronting two-pair primer design for enzyme-free SNP genotyping based on a genetic algorithmabstractBACKGROUND: Polymerase chain reaction with confronting two-pair primers (PCR-CTPP) method produces allele-specific DNA bands of different lengths by adding four designed primers and it achieves the single nucleotide polymorphism (SNP) genotyping by electrophoresis without further steps. It is a time- and cost-effective SNP genotyping method that has the advantage of simplicity. However, computation of feasible CTPP primers is still challenging. RESULTS: In this study, we propose a GA (genetic algorithm)-based method to design a feasible CTPP primer set to perform a reliable PCR experiment. The SLC6A4 gene was tested with 288 SNPs for dry dock experiments which indicated that the proposed algorithm provides CTPP primers satisfied most primer constraints. One SNP rs12449783 in the SLC6A4 gene was taken as an example for the genotyping experiments using electrophoresis which validated the GA-based design method as providing reliable CTPP primer sets for SNP genotyping. CONCLUSIONS: The GA-based CTPP primer design method provides all forms of estimation for the common primer constraints of PCR-CTPP. The GA-CTPP program is implemented in JAVA and a user-friendly input interface is freely available at http://bio.kuas.edu.tw/ga-ctpp/. Cheng-Hong Yang 0001, Yu-Huei Cheng, Li-Yeh Chuang, Hsueh-Wei Chang |
BMC Bioinform. | 1 |
| 2009 | A Hybrid Feature Selection Method Using Gene Expression DataabstractIn this paper, correlation-based feature selection (CFS) and the Taguchi-genetic algorithm (TGA) method were combined in a hybrid method, and the K-nearest neighbor (KNN) method with leave-one-out cross-validation (LOOCV) served as a classifier for eleven classification profiles. With the help of this classifier classification accuracy were calculated. Experimental results show that this method effectively simplifies features selection by reducing the total number of features needed. The proposed method obtained the highest classification accuracy in five out of the six gene expression data set test problems when compared to other classification methods from the literature. Li-Yeh Chuang, Kuo-Chuan Wu, Cheng-Hong Yang 0001 |
BIBE | 3 |
| 2009 | Genetic Algorithm for the Design of Confronting Two-Pair PrimersabstractMany single nucleotide polymorphisms (SNPs) genotyping techniques have been developed but most of them are expensive. Polymerase chain reaction with confronting two-pair primers (PCR-CTPP) is a restriction enzyme-free and economic genotyping but its primer design is still computationally challenged. Here, we introduced a genetic algorithm (GA)-based PCR-CTPP primer design method. Thirty SNPs of the Janus kinase 2 gene with their SNP flanking length for 500 bps were tested. These GA-based designing CTPP primers were characterized with close values for melting temperature (Tm) and specificity, and their corresponding PCR products were provided with the optimal length. In conclusion, this novel PCR-CTPP primer designing method provides the computation for the primer information of cost- and time-effective enzyme-free SNP genotyping. Cheng-Hong Yang 0001, Yu-Huei Cheng, Li-Yeh Chuang, Hsueh-Wei Chang |
BIBE | 1 |
| 2009 | Designing of a novel GA based on fuzzy system for prediction of CpG islands in the human genomeabstractIn this paper we proposed a novel genetic algorithm based on fuzzy system for identification CpG islands in human genome, called FGA-CGI (fuzzy GA-CpG Island). CpG islands play a fundamental role in genome analysis and annotation and contribute to increase the accuracy of promoter prediction. Recently, some approaches rely on large parameter space algorithms of predicting the CpG islands have been proposed in the literature. The goal of our proposed method was that using the evolutionary algorithms with fuzzy system and machine learning to identify CpG islands. A fuzzy expert system was implemented to dynamically adapt the crossover rate and mutation rate in GA for identify significant of CpG islands in human genome, and reinforcement learning serve as extend operation for combined the best subset of islands. In this study, three public tools for identification CpG islands were used to compare with FGA-CGI for the assessment of five prediction performance and statistically analysis. Experimental results reveal that our method can adjust the two variables to escape local optimal by fuzzy system and identify more number of CpG islands. In addition, FGA-CGI had capable of higher performance and precisely predicting statistically significant CpG islands in target sequences than these previous tools. Li-Yeh Chuang, Yu-Jung Chen, Cheng-Hong Yang 0001 |
FUZZ-IEEE | 3 |
| 2009 | Fuzzy guided BPSO method for haplotype tag SNP selectionabstractIn the current researches of disease-gene association, Single Nucleotide Polymorphism (SNP) is the most interested topic. However, genotyping all existing SNPs for a large number of samples is still challenging even though SNP arrays have been developed to facilitate the task. Therefore, it is essential to select only informative SNPs (tag SNP) representing the rest SNPs for genome-wide association studies. Accordingly, the cost of genotyping is expected to be largely reduced. In this study, the fuzzy guided binary particle swarm optimization (FBPSO) based approach make it possible to select tag SNPs with higher accuracy. The fuzzy logic is employed to tuning the inertia weight (w) of BPSO. Three publicly data sets from the literature have been used for testing the performance of FBPSO. The experimental results indicated that the fuzzy logic will reinforce the search capability of BPSO, which is more accurate than the state-of-the-art methods. On the average of testing results, it also outperforms SVM/STSA method about 3.7%. Li-Yeh Chuang, Yu-Jen Hou, Cheng-Hong Yang 0001 |
FUZZ-IEEE | 3 |
| 2009 | Improved catfish particle swarm optimization with fuzzy adaptationabstractCatfish particle swarm optimization (CatfishPSO) algorithm is a novel swarm intelligence optimization, which inspired by the behavior between sardines and catfish, i.e. the so-called catfish effect is applied to improve the performance of particle swarm optimization (PSO). In this paper, we propose an improved CatfishPSO with fuzzy adaptive (F-CatfishPSO), which a fuzzy system is implemented to dynamically adapt the inertia weight of the CatfishPSO. In the conducted experiments, we adapt the inertia weight to strengthen the solution quality of PSO and CatfishPSO via fuzzy system. Six benchmark functions with unimodal and multimodal different trait are selected as the test functions. The experimental results indicate that the performance of the F-CatfishPSO is better than methods from the literature by statistical analysis. Li-Yeh Chuang, Sheng-Wei Tsai, Cheng-Hong Yang 0001 |
FUZZ-IEEE | 3 |
| 2009 | Fuzzy adaptive particle swarm optimization for a specific primer design problemabstractA fuzzy system with dynamically adapt the inertia weight of the particle swarm optimization (FAPSO) had been implemented to select a specific feasible primer pair for PCR experiments. Overall, fifty accession nucleotide sequences between 1900 bps and 2100 bps were sampled for primer design with specific PCR product lengths of 150~300 bps and 500~800 bps. Total five hundred runs of the proposed primer design approaches were performed for each accession nucleotide sequence to calculate the optimum accuracy. The proposed approach is compared to standard PSO primer design method. The results generated in a dry dock experiment showed that the FAPSO primer design yielded approximately optimal primer sets and had a relatively short CPU-time than standard PSO. Related materials are available online at http://bio.kuas.edu.tw/fapso-pd/. Cheng-Hong Yang 0001, Yu-Huei Cheng, Hsueh-Wei Chang, Li-Yeh Chuang |
FUZZ-IEEE | 1 |
| 2009 | Improved Catfish Particle Swarm Optimization with Embedded Chaotic MapabstractChaotic catfish particle swarm optimization (C-CatfishPSO) is a novel optimization algorithm proposed in this paper. C-CatfishPSO introduces chaotic maps into catfish particle swarm optimization (CatfishPSO), which increase the search capability of CatfishPSO via the chaos approach. Simple CatfishPSO relies on the incorporation of catfish particles into particle swarm optimization (PSO). The introduced catfish particles improve the performance of PSO considerably. Unlike other ordinary particles, catfish particles initialize a new search from extreme points of the search space when the gbest fitness value (the best previously encountered value) has not changed for a certain number of consecutive iterations. This results in further opportunities of finding better solutions for the swarm by guiding the entire swarm to promising new regions of the search space, and by accelerating search efficiency. In this study, we adopted chaotic maps to strengthen the solution quality of PSO and CatfishPSO. After the introduction of chaotic maps into the process, the improved PSO and CatfishPSO are called chaotic PSO (C-PSO) and chaotic CatfishPSO (C-CatfishPSO), respectively. PSO, C-PSO, CatfishPSO and C-CatfishPSO were extensively compared on six benchmark functions. Statistical analysis of the experimental results indicates that the performance of C-CatfishPSO is better than the performance of PSO, C-PSO, and CatfishPSO. Li-Yeh Chuang, Sheng-Wei Tsai, Cheng-Hong Yang 0001 |
SMC | 3 |
| 2008 | Boolean binary particle swarm optimization for feature selectionabstractFeature selection is the process of choosing a subset of features from an original set. This subset should be necessary, reasonably represent the original data, and useful for identification classification. The task of feature selection is to search for an optimal solution in a - usually large - search space. However, if the search space too large, difficulties can occur during the search process, often resulting in a considerable increase in computational time. A particle swarm optimization algorithm (PSO) is a relatively new evolutionary computation technique, which has previously been used to implement feature selection. However, particle swarm optimization, like other evolutionary algorithms, tends to converge at a local optimum early. In this paper, we introduce a Boolean function which improves on the disadvantages of standard particle swarm optimization and use it to implement a feature selection for six microarray data sets. The experimental results show that the proposed method selects a smaller number of feature subsets and obtains better classification accuracy than standard PSO. Cheng-San Yang, Li-Yeh Chuang, Chao-Hsuan Ke, Cheng-Hong Yang 0001 |
IEEE Congress on Evolutionary Computation | 4 |
| 2008 | Improved tag SNP selection using binary particle swarm optimizationabstractSingle nucleotide polymorphisms (SNPs) hold much promise as a basis for disease-gene association. However, they are limited by the cost of genotyping the tremendous number of SNPs. It is therefore essential to select only informative subsets (tag SNPs) out of all SNPs. Several promising methods for tag SNP selection have been proposed, such as the haplotype block-based and block-free approaches. The block-free methods are preferred by some researchers because most of the block-based methods rely on strong assumptions, such as prior block-partitioning, bi-allelic SNPs, or a fixed number or locations for tagging SNPs. We employed the feature selection idea of binary particle swarm optimization (binary PSO) to find informative tag SNPs. This method is very efficient, as it does not rely on block partitioning of the genomic region. Using four public data sets, the method consistently identified tag SNPs with considerably better prediction ability than STAMPA. Moreover, this method retains its performance even when a very small number and 100% prediction accuracy are used for the tag SNPs. Cheng-Hong Yang 0001, Chang-Hsuan Ho, Li-Yeh Chuang |
IEEE Congress on Evolutionary Computation | 1 |
| 2008 | A Novel GA-Taguchi-Based Feature Selection Method
Cheng-Hong Yang 0001, Chi-Chun Huang, Kuo-Chuan Wu, Hsin-Yun Chang |
IDEAL | 1 |
| 2008 | A novel BPSO approach for gene selection and classification of microarray dataabstractSelecting relevant genes from microarray data poses a huge challenge due to the high-dimensionality of the features, multi-class categories and a relatively small sample size. The main task of the classification process is to decrease the microarray data dimensionality. In order to analyze microarray data, an optimal subset of features (genes) which adequately represents the original set of features has to be found. In this study, we used a novel binary particle swarm optimization (NBPSO) algorithm to perform microarray data selection and classification. The K-nearest neighbor (K-NN) method with leave-one-out cross-validation (LOOCV) served as a classifier. The experimental results showed that the proposed method not only effectively reduced the number of gene expression levels, but also achieved lower classification error rates. Cheng-San Yang, Li-Yeh Chuang, Jung-Chike Li, Cheng-Hong Yang 0001 |
IJCNN | 4 |
| 2008 | A hybrid filter/wrapper approach of feature selection for gene expression dataabstractIn recent years, many studies have shown that microarray gene expression data is useful for disease identification and cancer classification. However, since gene expression data may contain thousands of genes simultaneously, successful microarray classification can be rather difficult. Feature (gene) selection is a frequently used pre-processing technology for successful classification of microarray gene expression data. Selecting a useful gene subset as a classifier not only decreases the computational time and cost, but also increases the classification accuracy. It is therefore imperative to extract only a small number of genes, which are exclusively relevant for the classification of a particular cancer/disease type. In this paper, correlation-based binary particle swarm optimizations is proposed to select the relevant genes, and a K-nearest neighbor with the leave-one-out cross-validation method serves as a classifier to evaluate the classification performance on six published cancer classification data sets. The experimental results show that the proposed method selects fewer gene subsets, while still resulting in higher prediction accuracy than the other literature methods. Chao-Hsuan Ke, Cheng-Hong Yang 0001, Li-Yeh Chuang, Cheng-San Yang |
SMC | 2 |
| 2008 | Information gain with chaotic genetic algorithm for gene selection and classification problemabstractFor microarray data classification problem, selecting relevant genes from microarray data pose a formidable challenge to researchers due to the high-dimensionality of features, multi-class categories being involved and the usually small sample size. In order to correctly analyze microarray data, the goal of feature (gene) selection is to select those subsets of differentially expressed genes that are potentially relevant for distinguishing the sample classes. In this paper, information gain and chaotic genetic algorithm are proposed to select the relevant genes, and a K-nearest neighbor with the leave-one-out cross-validation method serves as a classifier. Chaotic genetic algorithm is modified by using the chaotic mutation operator to increase the population diversity. The experimental results show that the proposed method not only effectively reduced the number of gene expression levels, but also achieved lower classification error rates. Cheng-San Yang, Li-Yeh Chuang, Jung-Chike Li, Cheng-Hong Yang 0001 |
SMC | 4 |
| 2008 | Catfish particle swarm optimizationabstractCatfish particle swarm optimization (CatfishPSO) is a novel optimization algorithm proposed in this paper. The mechanism is dependent on the incorporation of a catfish particle into the linearly decreasing weight particle swarm optimization (LDWPSO). The introduced catfish particle improves the performance of LDWPSO. Unlike other ordinary particles, the catfish particles will initialize a new search from the extreme points of the search space when the gbest fitness value (global optimum at each iteration) has not been changed for a given time, which results in further opportunities to find better solutions for the swarm by guiding the whole swarm to promising new regions of the search space, and accelerating convergence. In our experiment, CatfishPSO, LDWPSO and other improved PSO procedures were extensively compared on three benchmark test functions with 10, 20 and 30 different dimensions. Experimental results indicate that CatfishPSO achieves better performance than LDWPSO procedure and other improved PSO algorithms from the literature. Li-Yeh Chuang, Sheng-Wei Tsai, Cheng-Hong Yang 0001 |
SIS | 3 |
| 2006 | V-MitoSNP: visualization of human mitochondrial SNPsabstractBACKGROUND: Mitochondrial single nucleotide polymorphisms (mtSNPs) constitute important data when trying to shed some light on human diseases and cancers. Unfortunately, providing relevant mtSNP genotyping information in mtDNA databases in a neatly organized and transparent visual manner still remains a challenge. Amongst the many methods reported for SNP genotyping, determining the restriction fragment length polymorphisms (RFLPs) is still one of the most convenient and cost-saving methods. In this study, we prepared the visualization of the mtDNA genome in a way, which integrates the RFLP genotyping information with mitochondria related cancers and diseases in a user-friendly, intuitive and interactive manner. The inherent problem associated with mtDNA sequences in BLAST of the NCBI database was also solved. DESCRIPTION: V-MitoSNP provides complete mtSNP information for four different kinds of inputs: (1) color-coded visual input by selecting genes of interest on the genome graph, (2) keyword search by locus, disease and mtSNP rs# ID, (3) visualized input of nucleotide range by clicking the selected region of the mtDNA sequence, and (4) sequences mtBLAST. The V-MitoSNP output provides 500 bp (base pairs) flanking sequences for each SNP coupled with the RFLP enzyme and the corresponding natural or mismatched primer sets. The output format enables users to see the SNP genotype pattern of the RFLP by virtual electrophoresis of each mtSNP. The rate of successful design of enzymes and primers for RFLPs in all mtSNPs was 99.1%. The RFLP information was validated by actual agarose electrophoresis and showed successful results for all mtSNPs tested. The mtBLAST function in V-MitoSNP provides the gene information within the input sequence rather than providing the complete mitochondrial chromosome as in the NCBI BLAST database. All mtSNPs with rs number entries in NCBI are integrated in the corresponding SNP in V-MitoSNP. CONCLUSION: V-MitoSNP is a web-based software platform that provides a user-friendly and interactive interface for mtSNP information, especially with regard to RFLP genotyping. Visual input and output coupled with integrated mtSNP information from MITOMAP and NCBI make V-MitoSNP an ideal and complete visualization interface for human mtSNPs association studies. Li-Yeh Chuang, Cheng-Hong Yang 0001, Yu-Huei Cheng, De-Leung Gu, Phei-Lang Chang, Ke-Hung Tsui, Hsueh-Wei Chang |
BMC Bioinform. | 2 |
| 2005 | A Wireless Emulation Management System for Learning Mandarin Phonetic and Chanjei Morse CodeabstractAssistive technology (AT) is becoming increasingly important in improving mobility, language and learning capabilities of persons who have disabilities enabling them to function independently and to improve their social opportunities. Morse code has been shown to be a valuable tool in assistive technology, augmentative and alternative communication, rehabilitation, and education, as well as adapted computer access methods via special software programs, hardware devices, and switches. In this study, we designed and implemented an interactive Mandarin phonetic/chanjei Morse code typing emulation system for persons with disabilities, with three adaptive recognition methods. Cheng-Hong Yang 0001, Li-Yeh Chuang, Shyang-Lung Lin, Chi-Min Wang |
ICALT | 1 |