EDBT 2026 Demo / reviewers in the wild / expert
Hsien-Da Huang
dblp:96/401
· DBLP profile ↗
34ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0003-2857-7023ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 26 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 8Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Integrative Framework for Functional analysis of multiple Traditional Chinese Medicines based on transcriptome-driven systems biology and supervised learning strategyabstractTraditional Chinese Medicines (TCMs) contain a wide variety of ingredients and are rich in bioactive chemical sources. However, their complex and unknown effects on the human body prove a challenge in TCM research. Multiple genes are involved in a biological system and disease progression; therefore, transcriptome-driven systems biology approach directs our attention to perturbed pathways or enriched gene sets from whole gene expression profile, allowing more explanatory power and less analytical complexity than solely considering differentially expressed genes (DEGs). This framework integrated differential expression analysis (DEA), pathway-based expression analysis and supervised learning approaches to explore TCM functions and extract signature genes of high importance. Supervised learning approaches were employed to extract signature genes that effectively discriminate between four TCMs (92% accuracy). These signature genes were further examined through database annotations and literature survey, where some were found to be TCM compound targets, such as fatty acid synthase (FASN) and thioredoxin reductase 1 (TXNRD1). Based on network pharmacology concept, the proposed TCM-compound-target-pathway network can link TCM compounds to signature genes, DEGs and perturbed pathways, revealing the rationale behind TCM’s mechanism of action in treatment. Integrating a supervised learning strategy in identifying TCM-associated genes can offer a new perspective for discovering potential drug targets. Chia-Ru Chung, Hsi-Yuan Huang, Yang-Chi-Dung Lin, Hua-Li Zuo, Siyao Hu, Hsiao-Chin Hong, Jinrui Bai, Hsien-Da Huang |
CIBCB | 12 |
| 2025 | DeepADR: multimodal prediction of adverse drug reaction frequency by integrating early-stage drug discovery information via Kolmogorov-Arnold networksabstractAdverse drug reactions (ADRs) are a major cause of clinical trial failure and postmarket withdrawal, posing significant risks to public health and impeding drug development. While computational methods offer an alternative to costly preclinical testing, existing models often fail with novel compounds by requiring pre-existing information such as drug-ADR associations or by inadequately integrating diverse data sources. Here, we introduce DeepADR, a multimodal deep learning framework for predicting both the occurrence and frequency of ADRs using early-stage, readily available data. DeepADR integrates chemical structures and biological target profiles with semantic representations of ADR terms derived from a large language model (LLMs). These heterogeneous parameters are fused using a Kolmogorov-Arnold Network (KAN), which enhances the modeling of complex, nonlinear relationships among modalities to improve predictive performance. Our model outperforms existing methods in predicting both ADR occurrence and frequency, demonstrating robust generalization to new chemical entities. DeepADR showed consistently better performance than other models across both classification and regression tasks. By effectively integrating chemical, biological, and semantic datasets, DeepADR provides a powerful, scalable tool for the early-stage safety assessment and candidate prioritization. This framework not only facilitates the prioritization of safer drug candidates but also offers a methodology for predicting the toxicity of other hazardous materials, holding significant promise for advancing public health. Jingting Wan, Chenyang Jia, Danhong Dong, Yigang Chen 0001, Yang-Chi-Dung Lin, Yisheng He, Hsi-Yuan Huang, Hsien-Da Huang |
Briefings Bioinform. | 8 |
| 2023 | Quantitative model for genome-wide cyclic AMP receptor protein binding site identification and characteristic analysisabstractCyclic AMP receptor proteins (CRPs) are important transcription regulators in many species. The prediction of CRP-binding sites was mainly based on position-weighted matrixes (PWMs). Traditional prediction methods only considered known binding motifs, and their ability to discover inflexible binding patterns was limited. Thus, a novel CRP-binding site prediction model called CRPBSFinder was developed in this research, which combined the hidden Markov model, knowledge-based PWMs and structure-based binding affinity matrixes. We trained this model using validated CRP-binding data from Escherichia coli and evaluated it with computational and experimental methods. The result shows that the model not only can provide higher prediction performance than a classic method but also quantitatively indicates the binding affinity of transcription factor binding sites by prediction scores. The prediction result included not only the most knowns regulated genes but also 1089 novel CRP-regulated genes. The major regulatory roles of CRPs were divided into four classes: carbohydrate metabolism, organic acid metabolism, nitrogen compound metabolism and cellular transport. Several novel functions were also discovered, including heterocycle metabolic and response to stimulus. Based on the functional similarity of homologous CRPs, we applied the model to 35 other species. The prediction tool and the prediction results are online and are available at: https://awi.cuhk.edu.cn/∼CRPBSFinder. Yigang Chen 0001, Yang-Chi-Dung Lin, Yijun Luo, Xiao-Xuan Cai, Peng Qiu, Shi-Dong Cui, Hsi-Yuan Huang, Hsien-Da Huang |
Briefings Bioinform. | 9 |
| 2023 | Extraction of microRNA-target interaction sentences from biomedical literature by deep learning approachabstractMicroRNA (miRNA)-target interaction (MTI) plays a substantial role in various cell activities, molecular regulations and physiological processes. Published biomedical literature is the carrier of high-confidence MTI knowledge. However, digging out this knowledge in an efficient manner from large-scale published articles remains challenging. To address this issue, we were motivated to construct a deep learning-based model. We applied the pre-trained language models to biomedical text to obtain the representation, and subsequently fed them into a deep neural network with gate mechanism layers and a fully connected layer for the extraction of MTI information sentences. Performances of the proposed models were evaluated using two datasets constructed on the basis of text data obtained from miRTarBase. The validation and test results revealed that incorporating both PubMedBERT and SciBERT for sentence level encoding with the long short-term memory (LSTM)-based deep neural network can yield an outstanding performance, with both F1 and accuracy being higher than 80% on validation data and test data. Additionally, the proposed deep learning method outperformed the following machine learning methods: random forest, support vector machine, logistic regression and bidirectional LSTM. This work would greatly facilitate studies on MTI analysis and regulations. It is anticipated that this work can assist in large-scale screening of miRNAs, thereby revealing their functional roles in various diseases, which is important for the development of highly specific drugs with fewer side effects. Source code and corpus are publicly available at https://github.com/qi29. Mengqi Luo, Shangfu Li, Yuxuan Pang, Lantian Yao, Renfei Ma, Hsi-Yuan Huang, Hsien-Da Huang, Tzong-Yi Lee |
Briefings Bioinform. | 7 |
| 2023 | Holistic similarity-based prediction of phosphorylation sites for understudied kinasesabstractPhosphorylation is an essential mechanism for regulating protein activities. Determining kinase-specific phosphorylation sites by experiments involves time-consuming and expensive analyzes. Although several studies proposed computational methods to model kinase-specific phosphorylation sites, they typically required abundant experimentally verified phosphorylation sites to yield reliable predictions. Nevertheless, the number of experimentally verified phosphorylation sites for most kinases is relatively small, and the targeting phosphorylation sites are still unidentified for some kinases. In fact, there is little research related to these understudied kinases in the literature. Thus, this study aims to create predictive models for these understudied kinases. A kinase-kinase similarity network was generated by merging the sequence-, functional-, protein-domain- and 'STRING'-related similarities. Thus, besides sequence data, protein-protein interactions and functional pathways were also considered to aid predictive modelling. This similarity network was then integrated with a classification of kinase groups to yield highly similar kinases to a specific understudied type of kinase. Their experimentally verified phosphorylation sites were leveraged as positive sites to train predictive models. The experimentally verified phosphorylation sites of the understudied kinase were used for validation. Results demonstrate that 82 out of 116 understudied kinases were predicted with adequate performance via the proposed modelling strategy, achieving a balanced accuracy of 0.81, 0.78, 0.84, 0.84, 0.85, 0.82, 0.90, 0.82 and 0.85, for the 'TK', 'Other', 'STE', 'CAMK', 'TKL', 'CMGC', 'AGC', 'CK1' and 'Atypical' groups, respectively. Therefore, this study demonstrates that web-like predictive networks can reliably capture the underlying patterns in such understudied kinases by harnessing relevant sources of similarities to predict their specific phosphorylation sites. Renfei Ma, Shangfu Li, Luca Parisi, Hsien-Da Huang, Tzong-Yi Lee |
Briefings Bioinform. | 5 |
| 2023 | Identification of species-specific RNA N6-methyladinosine modification sites from RNA sequencesabstractN6-methyladinosine (m6A) modification is the most abundant co-transcriptional modification in eukaryotic RNA and plays important roles in cellular regulation. Traditional high-throughput sequencing experiments used to explore functional mechanisms are time-consuming and labor-intensive, and most of the proposed methods focused on limited species types. To further understand the relevant biological mechanisms among different species with the same RNA modification, it is necessary to develop a computational scheme that can be applied to different species. To achieve this, we proposed an attention-based deep learning method, adaptive-m6A, which consists of convolutional neural network, bi-directional long short-term memory and an attention mechanism, to identify m6A sites in multiple species. In addition, three conventional machine learning (ML) methods, including support vector machine, random forest and logistic regression classifiers, were considered in this work. In addition to the performance of ML methods for multi-species prediction, the optimal performance of adaptive-m6A yielded an accuracy of 0.9832 and the area under the receiver operating characteristic curve of 0.98. Moreover, the motif analysis and cross-validation among different species were conducted to test the robustness of one model towards multiple species, which helped improve our understanding about the sequence characteristics and biological functions of RNA modifications in different species. Rulan Wang, Chia-Ru Chung, Hsien-Da Huang, Tzong-Yi Lee |
Briefings Bioinform. | 3 |
| 2020 | Multi-omics profiling reveals microRNA-mediated insulin signaling networksabstractBACKGROUND: MicroRNAs (miRNAs) play a key role in mediating the action of insulin on cell growth and the development of diabetes. However, few studies have been conducted to provide a comprehensive overview of the miRNA-mediated signaling network in response to glucose in pancreatic beta cells. In our study, we established a computational framework integrating multi-omics profiles analyses, including RNA sequencing (RNA-seq) and small RNA sequencing (sRNA-seq) data analysis, inverse expression pattern analysis, public data integration, and miRNA targets prediction to illustrate the miRNA-mediated regulatory network at different glucose concentrations in INS-1 pancreatic beta cells (INS-1), which display important characteristics of the pancreatic beta cells. RESULTS: We applied our computational framework to the expression profiles of miRNA/mRNA of INS-1, at different glucose concentrations. A total of 1437 differentially expressed genes (DEGs) and 153 differentially expressed miRNAs (DEmiRs) were identified from multi-omics profiles. In particular, 121 DEmiRs putatively regulated a total of 237 DEGs involved in glucose metabolism, fatty acid oxidation, ion channels, exocytosis, homeostasis, and insulin gene regulation. Moreover, Argonaute 2 immunoprecipitation sequencing, qRT-PCR, and luciferase assay identified Crem, Fn1, and Stc1 are direct targets of miR-146b and elucidated that miR-146b acted as a potential regulator and promising target to understand the insulin signaling network. CONCLUSIONS: In this study, the integration of experimentally verified data with system biology framework extracts the miRNA network for exploring potential insulin-associated miRNA and their target genes. The findings offer a potentially significant effect on the understanding of miRNA-mediated insulin signaling network in the development and progression of pancreatic diabetes. Yang-Chi-Dung Lin, Hsi-Yuan Huang, Sirjana Shrestha, Chih-Hung Chou, Yen-Hua Chen, Chi-Ru Chen, Hsiao-Chin Hong, Yi-An Chang, Men-Yee Chiew, Ya-Rong Huang, Siang-Jyun Tu, Ting-Hsuan Sun, Shun-Long Weng, Ching-Ping Tseng, Hsien-Da Huang |
BMC Bioinform. | 16 |
| 2019 | Biogenesis mechanisms of circular RNA can be categorized through feature extraction of a machine learning modelabstractMOTIVATION: In recent years, multiple circular RNAs (circRNA) biogenesis mechanisms have been discovered. Although each reported mechanism has been experimentally verified in different circRNAs, no single biogenesis mechanism has been proposed that can universally explain the biogenesis of all tens of thousands of discovered circRNAs. Under the hypothesis that human circRNAs can be categorized according to different biogenesis mechanisms, we designed a contextual regression model trained to predict the formation of circular RNA from a random genomic locus on human genome, with potential biogenesis factors of circular RNA as the features of the training data. RESULTS: After achieving high prediction accuracy, we found through the feature extraction technique that the examined human circRNAs can be categorized into seven subgroups, according to the presence of the following sequence features: RNA editing sites, simple repeat sequences, self-chains, RNA binding protein binding sites and CpG islands within the flanking regions of the circular RNA back-spliced junction sites. These results support all of the previously reported biogenesis mechanisms of circRNA and solidify the idea that multiple biogenesis mechanisms co-exist for different subset of human circRNAs. Furthermore, we uncover a potential new links between circRNA biogenesis and flanking CpG island. We have also identified RNA binding proteins putatively correlated with circRNA biogenesis. AVAILABILITY AND IMPLEMENTATION: Scripts and tutorial are available at http://wanglab.ucsd.edu/star/circRNA. This program is under GNU General Public License v3.0. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hsien-Da Huang, Wei Wang 0051 |
Bioinform. | 3 |
| 2016 | PARRoT- a homology-based strategy to quantify and compare RNA-sequencing from non-model organismsabstractBACKGROUND: Next-generation sequencing promises the de novo genomic and transcriptomic analysis of samples of interests. However, there are only a few organisms having reference genomic sequences and even fewer having well-defined or curated annotations. For transcriptome studies focusing on organisms lacking proper reference genomes, the common strategy is de novo assembly followed by functional annotation. However, things become even more complicated when multiple transcriptomes are compared. RESULTS: Here, we propose a new analysis strategy and quantification methods for quantifying expression level which not only generate a virtual reference from sequencing data, but also provide comparisons between transcriptomes. First, all reads from the transcriptome datasets are pooled together for de novo assembly. The assembled contigs are searched against NCBI NR databases to find potential homolog sequences. Based on the searched result, a set of virtual transcripts are generated and served as a reference transcriptome. By using the same reference, normalized quantification values including RC (read counts), eRPKM (estimated RPKM) and eTPM (estimated TPM) can be obtained that are comparable across transcriptome datasets. In order to demonstrate the feasibility of our strategy, we implement it in the web service PARRoT. PARRoT stands for Pipeline for Analyzing RNA Reads of Transcriptomes. It analyzes gene expression profiles for two transcriptome sequencing datasets. For better understanding of the biological meaning from the comparison among transcriptomes, PARRoT further provides linkage between these virtual transcripts and their potential function through showing best hits in SwissProt, NR database, assigning GO terms. Our demo datasets showed that PARRoT can analyze two paired-end transcriptomic datasets of approximately 100 million reads within just three hours. CONCLUSIONS: In this study, we proposed and implemented a strategy to analyze transcriptomes from non-reference organisms which offers the opportunity to quantify and compare transcriptome profiles through a homolog based virtual transcriptome reference. By using the homolog based reference, our strategy effectively avoids the problems that may cause from inconsistencies among transcriptomes. This strategy will shed lights on the field of comparative genomics for non-model organism. We have implemented PARRoT as a web service which is freely available at http://parrot.cgu.edu.tw . Richie Ruei-Chi Gan, Ting-Wen Chen, Timothy H. Wu, Po-Jung Huang, Chi-Ching Lee, Yuan-Ming Yeh, Cheng-Hsun Chiu, Hsien-Da Huang, Petrus Tang |
BMC Bioinform. | 8 |
| 2015 | GeNOSA: inferring and experimentally supporting quantitative gene regulatory networks in prokaryotesabstractMOTIVATION: The establishment of quantitative gene regulatory networks (qGRNs) through existing network component analysis (NCA) approaches suffers from shortcomings such as usage limitations of problem constraints and the instability of inferred qGRNs. The proposed GeNOSA framework uses a global optimization algorithm (OptNCA) to cope with the stringent limitations of NCA approaches in large-scale qGRNs. RESULTS: OptNCA performs well against existing NCA-derived algorithms in terms of utilization of connectivity information and reconstruction accuracy of inferred GRNs using synthetic and real Escherichia coli datasets. For comparisons with other non-NCA-derived algorithms, OptNCA without using known qualitative regulations is also evaluated in terms of qualitative assessments using a synthetic Saccharomyces cerevisiae dataset of the DREAM3 challenges. We successfully demonstrate GeNOSA in several applications including deducing condition-dependent regulations, establishing high-consensus qGRNs and validating a sub-network experimentally for dose-response and time-course microarray data, and discovering and experimentally confirming a novel regulation of CRP on AscG. AVAILABILITY AND IMPLEMENTATION: All datasets and the GeNOSA framework are freely available from http://e045.life.nctu.edu.tw/GeNOSA. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yi-Hsiung Chen, Chi-Dung Yang, Ching-Ping Tseng, Hsien-Da Huang, Shinn-Ying Ho |
Bioinform. | 4 |
| 2015 | Guest Editorial for the 13th Asia Pacific Bioinformatics ConferenceabstractPresents papers that were presented at the 13th Asia Pacific Bioinformatics Conference. Hsien-Da Huang, Yi-Ping Phoebe Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2013 | An enhanced computational platform for investigating the roles of regulatory RNA and for identifying functional RNA motifsabstractBACKGROUND: Functional RNA molecules participate in numerous biological processes, ranging from gene regulation to protein synthesis. Analysis of functional RNA motifs and elements in RNA sequences can obtain useful information for deciphering RNA regulatory mechanisms. Our previous work, RegRNA, is widely used in the identification of regulatory motifs, and this work extends it by incorporating more comprehensive and updated data sources and analytical approaches into a new platform. METHODS AND RESULTS: An integrated web-based system, RegRNA 2.0, has been developed for comprehensively identifying the functional RNA motifs and sites in an input RNA sequence. Numerous data sources and analytical approaches are integrated, and several types of functional RNA motifs and sites can be identified by RegRNA 2.0: (i) splicing donor/acceptor sites; (ii) splicing regulatory motifs; (iii) polyadenylation sites; (iv) ribosome binding sites; (v) rho-independent terminator; (vi) motifs in mRNA 5'-untranslated region (5'UTR) and 3'UTR; (vii) AU-rich elements; (viii) C-to-U editing sites; (ix) riboswitches; (x) RNA cis-regulatory elements; (xi) transcriptional regulatory motifs; (xii) user-defined motifs; (xiii) similar functional RNA sequences; (xiv) microRNA target sites; (xv) non-coding RNA hybridization sites; (xvi) long stems; (xvii) open reading frames; (xviii) related information of an RNA sequence. User can submit an RNA sequence and obtain the predictive results through RegRNA 2.0 web page. CONCLUSIONS: RegRNA 2.0 is an easy to use web server for identifying regulatory RNA motifs and functional sites. Through its integrated user-friendly interface, user is capable of using various analytical approaches and observing results with graphical visualization conveniently. RegRNA 2.0 is now available at http://regrna2.mbc.nctu.edu.tw. Tzu-Hao Chang, Hsi-Yuan Huang, Justin Bo-Kai Hsu, Shun-Long Weng, Jorng-Tzong Horng, Hsien-Da Huang |
BMC Bioinform. | 6 |
| 2012 | dbSNO: a database of cysteine S-nitrosylationabstractUNLABELLED: S-nitrosylation (SNO), a selective and reversible protein post-translational modification that involves the covalent attachment of nitric oxide (NO) to the sulfur atom of cysteine, critically regulates protein activity, localization and stability. Due to its importance in regulating protein functions and cell signaling, a mass spectrometry-based proteomics method rapidly evolved to increase the dataset of experimentally determined SNO sites. However, there is currently no database dedicated to the integration of all experimentally verified S-nitrosylation sites with their structural or functional information. Thus, the dbSNO database is created to integrate all available datasets and to provide their structural analysis. Up to April 15, 2012, the dbSNO has manually accumulated >3000 experimentally verified S-nitrosylated peptides from 219 research articles using a text mining approach. To solve the heterogeneity among the data collected from different sources, the sequence identity of these reported S-nitrosylated peptides are mapped to the UniProtKB protein entries. To delineate the structural correlation and consensus motif of these SNO sites, the dbSNO database also provides structural and functional analyses, including the motifs of substrate sites, solvent accessibility, protein secondary and tertiary structures, protein domains and gene ontology. AVAILABILITY: The dbSNO is now freely accessible via http://dbSNO.mbc.nctu.edu.tw. The database content is regularly updated upon collecting new data obtained from continuously surveying research articles. Tzong-Yi Lee, Yi-Ju Chen, Cheng-Tsung Lu, Wei-Chieh Ching, Yu-Chuan Teng, Hsien-Da Huang, Yu-Ju Chen |
Bioinform. | 6 |
| 2011 | miRTar: an integrated system for identifying miRNA-target interactions in HumanabstractBACKGROUND: MicroRNAs (miRNAs) are small non-coding RNA molecules that are ~22-nt-long sequences capable of suppressing protein synthesis. Previous research has suggested that miRNAs regulate 30% or more of the human protein-coding genes. The aim of this work is to consider various analyzing scenarios in the identification of miRNA-target interactions, as well as to provide an integrated system that will aid in facilitating investigation on the influence of miRNA targets by alternative splicing and the biological function of miRNAs in biological pathways. RESULTS: This work presents an integrated system, miRTar, which adopts various analyzing scenarios to identify putative miRNA target sites of the gene transcripts and elucidates the biological functions of miRNAs toward their targets in biological pathways. The system has three major features. First, the prediction system is able to consider various analyzing scenarios (1 miRNA:1 gene, 1:N, N:1, N:M, all miRNAs:N genes, and N miRNAs: genes involved in a pathway) to easily identify the regulatory relationships between interesting miRNAs and their targets, in 3'UTR, 5'UTR and coding regions. Second, miRTar can analyze and highlight a group of miRNA-regulated genes that participate in particular KEGG pathways to elucidate the biological roles of miRNAs in biological pathways. Third, miRTar can provide further information for elucidating the miRNA regulation, i.e., miRNA-target interactions, affected by alternative splicing. CONCLUSIONS: In this work, we developed an integrated resource, miRTar, to enable biologists to easily identify the biological functions and regulatory relationships between a group of known/putative miRNAs and protein coding genes. miRTar is now available at http://miRTar.mbc.nctu.edu.tw/. Justin Bo-Kai Hsu, Chih-Min Chiu, Sheng-Da Hsu, Wei-Yun Huang, Chia-Hung Chien, Tzong-Yi Lee, Hsien-Da Huang |
BMC Bioinform. | 7 |
| 2010 | Prediction of small non-coding RNA in bacterial genomes using support vector machines
Tzu-Hao Chang, Li-Ching Wu, Hsien-Da Huang, Baw-Jhiune Liu, Kuang-Fu Cheng, Jorng-Tzong Horng |
Expert Syst. Appl. | 4 |
| 2010 | An expert system to identify co-regulated gene groups from time-lagged gene clusters using cell cycle expression data
Li-Ching Wu, Jhih-Long Huang, Jorng-Tzong Horng, Hsien-Da Huang |
Expert Syst. Appl. | 4 |
| 2009 | RiboSW abstractRiboswitches are cis-acting genetic regulatory elements within a specific mRNA, and can regulate both transcription and translation by interacting with their corresponding metabolites. Recently, more and more riboswitches were identified and investigated about their roles in regulatory functions in different species. Both of the sequence contexts and structural conformations are important characteristics of riboswitches. None of previous developed tools, such Covariance Models (CMs), Riboswitch finder, and RibEx, provides a web server for efficiently searching homologous instances to known riboswitches and considers two crucial characteristics of each riboswitch, such as structural conformations and sequence contexts of functional regions. Therefore, we developed a systematic method to identify twelve kinds of riboswitches. The method is implemented and provided as a web server, RiboSW, to efficiently and conveniently identify riboswitches within messenger RNA sequences. RiboSW is now available on the web at http://bioinfo.csie.ncu.edu.tw/RiboSW/. The predictive accuracy of the proposed method is comparable with other previous tools. The efficiency of the proposed method for identifying riboswitches was improved in order to achieve a reasonable computational time required for the prediction. That makes it possible to have an accurate and convenient web server for biologists to obtain their analyzing results in a given mRNA sequence. Tzu-Hao Chang, Li-Ching Wu, Chi-Ta Yeh, Cheng-Wei Chang, Baw-Jhiune Liu, Hsien-Da Huang, Jorng-Tzong Horng |
BIBE | 6 |
| 2009 | A Human DNA Methylation Site Predictor Based on SVMabstractDuring gene expression, transcription factors are unable to bind to a transcription binding site (TFBS) involved in regulation if DNA methylation has occurred at the TFBS. Methyl-CpG-binding proteins may also occupy the TFBS and prevent the functioning of a transcription factor. Thus, the methylation status of CpG sites is an important issue when trying to understand gene regulation and shows strong correlation with the TFBS involved. In addition, CpG islands would seem to undergo cell-specific and tissue-specific me-thylation. Such differential methylation is presented at numerous genetic loci that are essential for development. Current DNA methylation site prediction tools need to be improved so that they include TFBS features and have greater accuracy in terms of the DNA region that is involved in methylation. We developed models that compare the differences across these regions and tissues. The TFBSs, DNA properties and DNA distribution were used as features for this classification. From the results, we found some TFBSs that were able to discriminate whether a sequence was methylated or not. The sensitivity, specificity and accuracy estimated using 10-fold cross validation were 90.8%, 80.54%, and 86.07%, respectively. Thus, for these four regions and twelve tissues, the performance levels (ACC) were all greater than 80%. We propose that the differential features or methylations vary between the different regions because the features common to each DNA region made up only 50% of the top 70 features. An online predictor based on EpiMeP is available at http://140.115.51.41/EpiMeP/. Supplementary file is available at http://140.115.51.41/EpiMeP/supplementary.doc. Yi-Ming Sun, Wei-Li Liao, Hsien-Da Huang, Baw-Jhiune Liu, Cheng-Wei Chang, Jorng-Tzong Horng, Li-Ching Wu |
BIBE | 3 |
| 2009 | miRExpress: Analyzing high-throughput sequencing data for profiling microRNA expressionabstractBACKGROUND: MicroRNAs (miRNAs), small non-coding RNAs of 19 to 25 nt, play important roles in gene regulation in both animals and plants. In the last few years, the oligonucleotide microarray is one high-throughput and robust method for detecting miRNA expression. However, the approach is restricted to detecting the expression of known miRNAs. Second-generation sequencing is an inexpensive and high-throughput sequencing method. This new method is a promising tool with high sensitivity and specificity and can be used to measure the abundance of small-RNA sequences in a sample. Hence, the expression profiling of miRNAs can involve use of sequencing rather than an oligonucleotide array. Additionally, this method can be adopted to discover novel miRNAs. RESULTS: This work presents a systematic approach, miRExpress, for extracting miRNA expression profiles from sequencing reads obtained by second-generation sequencing technology. A stand-alone software package is implemented for generating miRNA expression profiles from high-throughput sequencing of RNA without the need for sequenced genomes. The software is also a database-supported, efficient and flexible tool for investigating miRNA regulation. Moreover, we demonstrate the utility of miRExpress in extracting miRNA expression profiles from two Illumina data sets constructed for the human and a plant species. CONCLUSION: We develop miRExpress, which is a database-supported, efficient and flexible tool for detecting miRNA expression profile. The analysis of two Illumina data sets constructed from human and plant demonstrate the effectiveness of miRExpress to obtain miRNA expression profiles and show the usability in finding novel miRNAs. Wei-Chi Wang, Feng-Mao Lin, Wen-Chi Chang, Hsien-Da Huang, Na-Sheng Lin |
BMC Bioinform. | 5 |
| 2009 | Detecting LTR structures in human genomic sequences using profile hidden Markov models
Li-Ching Wu, Hsien-Da Huang, Yu-Chung Chang, Ying-Chun Lee, Jorng-Tzong Horng |
Expert Syst. Appl. | 2 |
| 2009 | An expert system to predict protein thermostability using decision tree
Li-Cheng Wu, Jian-Xin Lee, Hsien-Da Huang, Baw-Juine Liu, Jorng-Tzong Horng |
Expert Syst. Appl. | 3 |
| 2008 | Identifying Discriminative Amino Acids Within the Hemagglutinin of Human Influenza A H5N1 Virus Using a Decision TreeabstractRecently, the H5N1 virus has had an increasingly important impact on human life. This is because more and more people are becoming infected with this virus, and the possibility of a serious pandemic with human to human transmission is looming. This might occur if the genome of this influenza virus mutates either by antigenic drift or by antigenic shift, especially if there is a mutation of the hemagglutinin (HA) glycoprotein. The HA is the surface glycoprotein, and it binds to sialic acid of the host cell surface receptor. Thus, the combination of HA and sialic acid are central to whether influenza virus infects humans. In this study, we selected 497 HA protein sequences from the National Center for Biotechnology Information (NCBI) Influenza Resource database, and used a decision tree method to identify discriminative amino acids in the HA protein sequences that may possibly influence the binding of HA to sialic acid. Four such amino acid positions at 54, 55, 241, and 281 were identified and these may play an important role in infection by H5N1 influenza virus. Li-Ching Wu, Jorng-Tzong Horng, Hsien-Da Huang, Wei-Long Chen |
IEEE Trans. Inf. Technol. Biomed. | 3 |
| 2007 | Primer design for multiplex PCR using a genetic algorithm
Li-Cheng Wu, Jorng-Tzong Horng, Hsi-Yuan Huang, Feng-Mao Lin, Hsien-Da Huang, Meng-Feng Tsai |
Soft Comput. | 5 |
| 2006 | Database to Dynamically Aid Probe Design for Virus IdentificationabstractViral infection poses a major problem for public health, horticulture, and animal husbandry, possibly causing severe health crises and economic losses. Viral infections can be identified by the specific detection of viral sequences in many ways. The microarray approach not only tolerates sequence variations of newly evolved virus strains, but can also simultaneously diagnose many viral sequences. Many chips have so far been designed for clinical use. Most are designed for special purposes, such as typing enterovirus infection, and compare fewer than 30 different viral sequences. None considers primer design, increasing the likelihood of cross hybridization to similar sequences from other viruses. To prevent this possibility, this work establishes a platform and database that provides users with specific probes of all known viral genome sequences to facilitate the design of diagnostic chips. This work develops a system for designing probes online. A user can select any number of different viruses and set the experimental conditions such as melting temperature and length of probe. The system then returns the optimal sequences from the database. We have also developed a heuristic algorithm to calculate the probe correctness and show the correctness of the algorithm. (The system that supports probe design for identifying viruses has been published on our web page http://bioinfo.csie.ncu.edu.tw/.) Feng-Mao Lin, Hsien-Da Huang, Ann-Ping Tsou, P.-L. Chan, L.-C. Wu, Meng-Feng Tsai, Jorng-Tzong Horng |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2006 | Biological Data Warehousing System for Identifying Transcriptional Regulatory Sites From Gene Expressions of Microarray DataabstractIdentification of transcriptional regulatory sites plays an important role in the investigation of gene regulation. For this propose, we designed and implemented a data warehouse to integrate multiple heterogeneous biological data sources with data types such as text-file, XML, image, MySQL database model, and Oracle database model. The utility of the biological data warehouse in predicting transcriptional regulatory sites of coregulated genes was explored using a synexpression group derived from a microarray study. Both of the binding sites of known transcription factors and predicted over-represented (OR) oligonucleotides were demonstrated for the gene group. The potential biological roles of both known nucleotides and one OR nucleotide were demonstrated using bioassays. Therefore, the results from the wet-lab experiments reinforce the power and utility of the data warehouse as an approach to the genome-wide search for important transcription regulatory elements that are the key to many complex biological systems. Ann-Ping Tsou, Yi-Ming Sun, Chia-Lin Liu, Hsien-Da Huang, Jorng-Tzong Horng, Meng-Feng Tsai, Baw-Jhiune Liu |
IEEE Trans. Inf. Technol. Biomed. | 4 |
| 2005 | A database to aid probe design for virus identification
Feng-Mao Lin, Hsien-Da Huang, Yu-Chung Chang, Pak-Leong Chan, Jorng-Tzong Horng, Ming-Tat Ko |
APBC | 2 |
| 2005 | Primer design for multiplex PCR using a genetic algorithmabstractMultiplex Polymerase Chain Reaction (PCR) experiments are used for amplifying several segments of the target DNA simultaneously and thereby to conserve template DNA, reduce the experimental time, and minimize the experimental expense. The success of the experiment is dependent on primer design. However, this can be a dreary task as there are many constrains such as melting temperatures, primer length, GC content and complementarity that need to be optimized to obtain a good PCR product. Motivated by the lack of primer design tools for multiplex PCR genotypic assay, we propose a multiplex PCR primer design tool using a genetic algorithm, which is a stochastic approach based on the concept of biological evolution, biological genetics and genetic operations on chromosomes, to find an optimal selection of primer pairs for multiplex PCR experiments. The presented experimental results indicate that the proposed algorithm is capable of finding a series of primer pairs that obeies the design properties in the same tube. Feng-Mao Lin, Hsien-Da Huang, Hsi-Yuan Huang, Jorng-Tzong Horng |
GECCO | 2 |
| 2004 | Identifying the Combination of Genetic Factors that Determine Susceptibility to Cervical CancerabstractCervical cancer is common among women all over the world. Although infection with high-risk types of human papillomavirus (HPV) has been identified as the primary cause of cervical cancer, only some of those infected go on to develop cervical cancer. Obviously, the progression from HPV infection to cancer involves other environmental and host factors. Recent population-based twin and family studies have demonstrated the importance of the hereditary component of cervical cancer, associated with genetic susceptibility. Consequently, SNP markers and microsatellites should be considered genetic factors for determining what combinations of genetic factors are involved in precancerous changes to cervical cancer. This study employs a Bayesian network and four different decision tree algorithms, and compares the performance of these learning algorithms. The results of this study raise the possibility of investigations that could identify combinations of genetic factors, such as SNPs and microsatellites, that influence the risk associated with common complex multifactorial diseases, such as cervical cancer. The web site associated with this study is http://dblab8.csie.ncu.edu.tw/FactorAnalysis/. Jorng-Tzong Horng, Kai-Chih Hu, Li-Cheng Wu, Hsien-Da Huang, Horn-Cheng Lai, Ton-Yuen Chu |
BIBE | 4 |
| 2004 | RgS-Miner: A Biological Data Warehousing, Analyzing and Mining System for Identifying Transcriptional Regulatory Sites in Human Genome
Yi-Ming Sun, Hsien-Da Huang, Jorng-Tzong Horng, Shir-Ly Huang, Ann-Ping Tsou |
DEXA | 2 |
| 2004 | Identifying the combination of genetic factors that determine susceptibility to cervical cancerabstractCervical cancer is common among women all over the world. Although infection with high-risk types of human papillomavirus (HPV) has been identified as the primary cause of cervical cancer, only some of those infected go on to develop cervical cancer. Obviously, the progression from HPV infection to cancer involves other environmental and host factors. Recent population-based twin and family studies have demonstrated the importance of the hereditary component of cervical cancer, associated with genetic susceptibility. Consequently, single-nucleotide polymorphism (SNP) markers and microsatellites should be considered genetic factors for determining what combinations of genetic factors are involved in precancerous changes to cervical cancer. This study employs a Bayesian network and four different decision tree algorithms, and compares the performance of these learning algorithms. The results of this study raise the possibility of investigations that could identify combinations of genetic factors, such as SNPs and microsatellites, that influence the risk associated with common complex multifactorial diseases, such as cervical cancer. Jorng-Tzong Horng, Kai-Chih Hu, Li-Cheng Wu, Hsien-Da Huang, Feng-Mao Lin, Shir-Ly Huang, Horn-Cheng Lai, Ton-Yuen Chu |
IEEE Trans. Inf. Technol. Biomed. | 4 |
| 2003 | A Data Mining Method to Predict Transcriptional Regulatory Sites Based on Differentially Expressed Genes in Human GenomeabstractVery large-scale gene expression analysis, i.e., UniGene and dbEST, are provided to find those genes with significantly differential expression in specific tissues. The differentially expressed genes in a specific tissue are potentially regulated concurrently by a combination of transcription factors. This study attempts to mine putative binding sites on how combinations of the known regulatory sites homologs and over-represented repetitive elements are distributed in the promoter regions of considered groups of differentially expressed genes. We propose a data mining approach to statistically discover the significantly tissue-specific combinations of known site homologs and over-represented repetitive sequences, which are distributed in the promoter regions of differential gene groups. The association rules mined would facilitate to predict putative regulatory elements and identify genes potentially co-regulated by the putative regulatory elements. Hsien-Da Huang, Huei-Lin Chang, Tsung-Shan Tsou, Baw-Jhiune Liu, Cheng-Yan Kao, Jorng-Tzong Horng |
BIBE | 1 |
| 2003 | Database of repetitive elements in complete genomes and data mining using transcription factor binding sitesabstractApproximately 43% of the human genome is occupied by repetitive elements. Even more, around 51% of the rice genome is occupied by repetitive elements. The analysis presented here indicates that repetitive elements in complete genomes may have been very important in the evolutionary genomics. In this study, a database, called the Repeat Sequence Database, is first designed and implemented to store complete and comprehensive repetitive sequences. See http://rsdb.csie.ncu.edu.tw for more information. The database contains direct, inverted and palindromic repetitive sequences, and each repetitive sequence has a variable length ranging from seven to many hundred nucleotides. The repetitive sequences in the database are explored using a mathematical algorithm to mine rules on how combinations of individual binding sites are distributed among repetitive sequences in the database. Combinations of transcription factor binding sites in the repetitive sequences are obtained and then data mining techniques are applied to mine association rules from these combinations. The discovered associations are further pruned to remove insignificant associations and obtain a set of associations. The mined association rules facilitate efforts to identify gene classes regulated by similar mechanisms and accurately predict regulatory elements. Experiments are performed on several genomes including C. elegans, human chromosome 22, and yeast. Jorng-Tzong Horng, Feng-Mao Lin, J. H. Lin, Hsien-Da Huang, Baw-Jhiune Liu |
IEEE Trans. Inf. Technol. Biomed. | 4 |
| 2002 | Devising a cost effective baseball scheduling by evolutionary algorithmsabstractWe discuss the scheduling problems of a sports league and propose a new approach to solve these problems by applying evolution strategy. A schedule in a sports league must satisfy many constraints on timing, such as the number of games played between every pair of teams, the bounds on the number of consecutive home (or away) games for each team, every pair of teams must have played each other in the first half of the season, and so on. In addition to finding a feasible schedule that meets all the timing restrictions, the problem addressed has the additional complexity of having the objective of minimizing travel costs and every team having a balanced number of games at home. We formalize the scheduling problem into an optimization problem and adopt the concept of evolution strategy to solve it. We define the travel cost and distance cost for teams in the sports league by referring to Major League Baseball (MLB) in the United States and focus on the scheduling problem in MLB. Using the new method, it is more efficient at finding better results than previous approaches. Jih Tsung Yang, Hsien-Da Huang, Jorng-Tzong Horng |
IEEE Congress on Evolutionary Computation | 2 |
| 2001 | Discovering Common Structural Motifs from SSU 16 S Ribosomal RNA Secondary StructuresabstractSome structural motifs, like tetra-loops, in ribosomal RNA are known to functionally implicate in virtually every aspect of protein synthesis. Our aim in this study is to discover common structural motifs (CSMs), which are related to specific domains or functions, within the secondary structures of ribosomal RNAs in a data set constructed. After applying data mining techniques to mine the common structural motifs, a machine learning approach is used to find significant discriminating common structural motifs from groups of organisms. By applying to several data sets constructed in this study, it suggests that the CSMs can provide effective information to classify organisms and help biologists understand the functions of ribosomal RNA. From the experiments of the classification of organisms and the construction of phylogenetic trees by CSMs mined, we find our approach is promising. Hsien-Da Huang, Shu-Fen Fang, Jorng-Tzong Horng, Cheng-Yan Kao |
BIBE | 1 |