Weida Tong

dblp:53/3185 · DBLP profile ↗
← Back
44ranked-venue papers
0as first author
4since 2021 · last 2023
0000-0003-3488-6148ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 43 · 4 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2023 PLM-ARG: antibiotic resistance gene identification using a pretrained protein language model
abstract
MOTIVATION: Antibiotic resistance presents a formidable global challenge to public health and the environment. While considerable endeavors have been dedicated to identify antibiotic resistance genes (ARGs) for assessing the threat of antibiotic resistance, recent extensive investigations using metagenomic and metatranscriptomic approaches have unveiled a noteworthy concern. A significant fraction of proteins defies annotation through conventional sequence similarity-based methods, an issue that extends to ARGs, potentially leading to their under-recognition due to dissimilarities at the sequence level. RESULTS: Herein, we proposed an Artificial Intelligence-powered ARG identification framework using a pretrained large protein language model, enabling ARG identification and resistance category classification simultaneously. The proposed PLM-ARG was developed based on the most comprehensive ARG and related resistance category information (>28K ARGs and associated 29 resistance categories), yielding Matthew's correlation coefficients (MCCs) of 0.983 ± 0.001 by using a 5-fold cross-validation strategy. Furthermore, the PLM-ARG model was verified using an independent validation set and achieved an MCC of 0.838, outperforming other publicly available ARG prediction tools with an improvement range of 51.8%-107.9%. Moreover, the utility of the proposed PLM-ARG model was demonstrated by annotating resistance in the UniProt database and evaluating the impact of ARGs on the Earth's environmental microbiota. AVAILABILITY AND IMPLEMENTATION: PLM-ARG is available for academic purposes at https://github.com/Junwu302/PLM-ARG, and a user-friendly webserver (http://www.unimd.org/PLM-ARG) is also provided.
Jian Ouyang, Haipeng Qing, Jiajia Zhou 0004, Ruth Roberts, Rania Siam, Weida Tong, Tieliu Shi
Bioinform.8
2022 Best practice and reproducible science are required to advance artificial intelligence in real-world applications
abstract
Drug-induced liver injury (DILI) is one of the most significant concerns in medical practice but yet it still cannot be fully recapitulated with existing in vivo, in vitro and in silico approaches. To address this challenge, Chen et al. [ 1] developed a deep learning-based DILI prediction model based on chemical structure information alone. The reported model yielded an outstanding prediction performance (i.e. 0.958, 0.976, 0.935, 0.947, 0.926 and 0.913 for AUC, accuracy, recall, precision, F1-score and specificity, respectively, on a test set), far outperforming all publicly available and similar in silico DILI models. This extraordinary model performance is counter-intuitive to what we know about the underlying biology of DILI and the principles and hypothesis behind this type of in silico approach. In this Letter to the Editor, we raise awareness of several issues concerning data curation, model validation and comparison practices, and data and model reproducibility.
Skylar Connor, Shraddha Thakkar, Ruth Roberts, Weida Tong
Briefings Bioinform.6
2021 Text Fingerprinting and Topic Mining in the Prescription Opioid Use Literature
abstract
Prescription opioids are powerful pain-reducing medications. Thousands of articles that focus on prescription opioid use (POU) and its associated medical disorders have been published. However, it is time-consuming and labor-intensive to extract and understand the information of all POU-related published articles. In this study, we applied the well-adapted topic modeling method, Latent Dirichlet Allocation (LDA), to perform text mining on POU-related literature. We have collected six large academic abstract datasets by searching PubMed using the Medical Subject Headings (MeSH): prescription opioid, codeine, morphine, hydrocodone, oxycodone, and methadone. We then applied topic modeling to identify topics and analyze topic similarities/differences in these six datasets. Word clouds and histograms were used to depict the distribution of vocabularies over each topic in which the most prevalent words conveyed a topic’s meaning. TreeMap and trend analysis were performed to fingerprint abstracts and explore the prevalent topic dynamics in the POU-related literature. Results showed the ability of topic modeling as a computational tool to segregate a vast quantity of articles into different themes that provide a systematic literature overview. The LDA topics recaptured the search keywords in PubMed and revealed further relevant themes by comparison analysis between different datasets.
Huyen Le, Junxiu Zhou, Weizhong Zhao, Roger Perkins, Weigong Ge, Beverly Lyn-Cook, Henry Francis, Huixiao Hong, Weida Tong, Wen Zou
BIBM9
2021 Discovering Drug-Drug Associations in the FDA Adverse Event Reporting System Database with Data Mining Approaches
abstract
Objective: To classify causal associations among drugs and adverse events by the identified drug safety signals from the US Food and Drug Administration (FDA) Adverse Event Reporting System (FAERS).Material and Methods: FAERS reports were collected for a period between 2004 and 2014. Empirical Bayes Geometric Mean was applied to model associations between drugs and adverse events. Identified signals were evaluated using Reporting Ratio values and Chi-Square test. Based on the identified drug-associated adverse events, we constructed a drug-drug association network, and applied a random walk algorithm to find drug communities with similar adverse event patterns. We developed 14 clusters for comparison with the 14 main groups in the first level of the Anatomical Therapeutic Chemical (ATC) classification system to evaluate relationships between the two classification systems.Results: The retrieved FAERS dataset included 981 drugs and 16,179 adverse events, from which we identified 63,083 significant drug-adverse event pairs. We found new potential safety signals when comparing the drug-adverse event pairs with information from relevant sources. Network analysis of the constructed drug communities revealed connections among drugs, adverse events, and ATC codes, suggesting that drug adverse events may be used to predict the ATC codes of unclassified drugs. For the ATC-classified drugs, the network analysis revealed potential relationships among the drugs by calculating the similarities of the adverse events.Conclusion: The generated drug groups and the drug-drug associations derived from the network analysis might have predictive potential for adverse events, as well as provide information for drug review and development.
Weizhong Zhao, Huyen Le, James J. Chen, Hesha J. Duggirala, Richard Forshee, Taxiarchis Botsis, Henry Francis, Huixiao Hong, Weida Tong, Yi-Ting Hwang, Wen Zou
BIBM9
2019 Similarities and differences between variants called with human reference genome HG19 or HG38
abstract
BACKGROUND: Reference genome selection is a prerequisite for successful analysis of next generation sequencing (NGS) data. Current practice employs one of the two most recent human reference genome versions: HG19 or HG38. To date, the impact of genome version on SNV identification has not been rigorously assessed. METHODS: We conducted analysis comparing the SNVs identified based on HG19 vs HG38, leveraging whole genome sequencing (WGS) data from the genome-in-a-bottle (GIAB) project. First, SNVs were called using 26 different bioinformatics pipelines with either HG19 or HG38. Next, two tools were used to convert the called SNVs between HG19 and HG38. Lastly we calculated conversion rates, analyzed discordant rates between SNVs called with HG19 or HG38, and characterized the discordant SNVs. RESULTS: The conversion rates from HG38 to HG19 (average 95%) were lower than the conversion rates from HG19 to HG38 (average 99%). The conversion rates varied slightly among the various calling pipelines. Around 1.5% SNVs were discordantly converted between HG19 or HG38. The conversions from HG38 to HG19 had more SNVs which failed conversion and more discordant SNVs than the opposite conversion (HG19 to HG38). Most of the discordant SNVs had low read depth, were low confidence SNVs as defined by GIAB, and/or were predominated by G/C alleles (52% observed versus 42% expected). CONCLUSION: A significant number of SNVs could not be converted between HG19 and HG38. Based on careful review of our comparisons, we recommend HG38 (the newer version) for NGS SNV analysis. To summarize, our findings suggest caution when translating identified SNVs between different versions of the human reference genome.
Bohu Pan, Rebecca Kusko, Wenming Xiao, Yuanting Zheng, Chunlin Xiao, Sugunadevi Sakkiah, Ping Gong 0001, Weigong Ge, Leming Shi, Weida Tong, Huixiao Hong
BMC Bioinform.13
2019 Correction to: Similarities and differences between variants called with human reference genome HG19 or HG38
abstract
After publication of this supplement article.
Bohu Pan, Rebecca Kusko, Wenming Xiao, Yuanting Zheng, Chunlin Xiao, Sugunadevi Sakkiah, Ping Gong 0001, Weigong Ge, Leming Shi, Weida Tong, Huixiao Hong
BMC Bioinform.13
2019 Study of serious adverse drug reactions using FDA-approved drug labeling and MedDRA
abstract
BACKGROUND: Adverse Drug Reactions (ADRs) are of great public health concern. FDA-approved drug labeling summarizes ADRs of a drug product mainly in three sections, i.e., Boxed Warning (BW), Warnings and Precautions (WP), and Adverse Reactions (AR), where the severity of ADRs are intended to decrease in the order of BW > WP > AR. Several reported studies have extracted ADRs from labeling documents, but most, if not all, did not discriminate the severity of the ADRs by the different labeling sections. Such a practice could overstate or underestimate the impact of certain ADRs to the public health. In this study, we applied the Medical Dictionary for Regulatory Activities (MedDRA) to drug labeling and systematically analyzed and compared the ADRs from the three labeling sections with a specific emphasis on analyzing serious ADRs presented in BW, which is of most drug safety concern. RESULTS: This study investigated New Drug Application (NDA) labeling documents for 1164 single-ingredient drugs using Oracle Text search to extract MedDRA terms. We found that only a small portion of MedDRA Preferred Terms (PTs), 3819 out of 21,920 or 17.42%, were observed in a whole set of documents. In detail, 466/3819 (12.0%) PTs were in BW, 2023/3819 (53.0%) were in WP, and 2961/3819 (77.5%) were in AR sections. We also found a higher overlap of top 20 occurring BW PTs with WP sections compared to AR sections. Within the MedDRA System Organ Class levels, serious ADRs (sADRs) from BW were prevalent in Nervous System disorders and Vascular disorders. A Hierarchical Cluster Analysis (HCA) revealed that drugs within the same therapeutic category shared the same ADR patterns in BW (e.g., nervous system drug class is highly associated with drug abuse terms such as dependence, substance abuse, and respiratory depression). CONCLUSIONS: This study demonstrated that combining MedDRA standard terminologies with data mining techniques facilitated computer-aided ADR analysis of drug labeling. We also highlighted the importance of labeling sections that differ in seriousness and application in drug safety. Using sADRs primarily related to BW sections, we illustrated a prototype approach for computer-aided ADR monitoring and studies which can be applied to other public health documents.
Leihong Wu, Taylor Ingle, Anna Zhao-Wong, Stephen C. Harris, Shraddha Thakkar, Guangxu Zhou, Junshuang Yang, Joshua Xu, Darshan Mehta, Weigong Ge, Weida Tong
BMC Bioinform.12
2017 Comprehensive analysis of pulmonary adenocarcinoma in situ (AIS) revealed new insights into lung cancer progression
abstract
Pulmonary adenocarcinoma in situ (AIS) is an intermediate subtype of lung adenocarcinoma that exhibits non-invasive growth patterns, but can develop into invasive. Almost 100% of AIS patients can be cured with complete resection. In contrast, the five-year survival rate for those diagnosed with invasive lung adenocarcinoma is only about 4%. In order to get a better understanding of adenocarcinoma and identify marks indicating its evaluation and progression, what needs to be done is a genome-wide evaluation of the disease. In this study, we used RNA-seq data from normal, AIS, and invasive lung cancer samples to identify gene module with differential gene expressions that represents the properties of AIS, distinct from normal and invasive tumor. Utilizing the differential expression patterns of protein-coding genes and long non-coding RNAs (lncRNAs) between AIS and other conditions, we obtained a small group of 72 AIS-specific genes consisting 41 protein-coding genes, 5 annotated lncRNAs, and 26 novel lncRNA transcripts that expressed specifically in AIS samples. We consider that twelve of the protein-coding genes are lung cancer driver genes with located driver somatic mutations. Moreover, these AIS-specific genes show capabilities (98% accuracy in an independent data set) for indicating early stage lung cancer and normal situations. These genes are determined based on cell-adhesion functioning, including angiogenesis and fibronectin, etc., that are highly related to cancer development. The comparison of AIS with normal and invasive tumor provides the gene list revealing the mechanisms that accounts for AIS progression to invasive cancer. Furthermore we identified important signatures contributing to the early diagnosis of lung cancer, providing auxiliary studies to our ongoing precision medicine research (http://americancse.org/events/csce2017/keynotes_lectures/yang_talk).
William Yang, Jack Y. Yang, Weida Tong, Renchu Guan, Mary Yang
BIBM5
2017 Toward precision breast cancer survival prediction utilizing combined whole genome-wide expression and somatic mutation analysis
abstract
Breast cancer is the most common type of invasive cancer in females. It accounts for 18.2% of all cancer deaths worldwide. Although somatic mutations play important roles in cancer development and prognosis, the outcome predictions are largely based on the expression of marker genes. We submit that developing an innovative prognostic model incorporating somatic mutations with gene expression can improve survival prediction of cancer. We studied the whole genome-wide gene expression and somatic mutations of 1091 breast invasive carcinoma cases from The Cancer Genome Atlas (TCGA). We identified expression of 118 genes that could be used to build the predictor for breast cancer survival risks (log rank p <; 0.0001 and c-index=0.63627122). Multiple breast cancer survival-analysis-related genes are found in this gene set, such as FOXR2, FOXD1, MTNR1B, SDC1, PF4, IGF2BP1, ZIC3, OXT2, and PID1. We then selected between different survival-risk groups of 2000 mutated genes with different mutation rates. We applied enrichment analysis to the mutated gene list and identified 25 functional annotations, 15 gene ontology, and 14 gene pathway enriched terms that were related the cancer outcomes. We built the novel predictor and our results showed that combining the different features helps improved performance of survival prediction (c-index = 0. 64033769). Thus our model can be used to facilitate the advancement of our going precision medicine research (http://americancse.org/events/csce2017/keynotes_lectures/yang_talk).
William Yang, Jack Y. Yang, Renchu Guan, Weida Tong, Mary Yang
BIBM6
2017 Scaling bioinformatics applications on HPC
abstract
BACKGROUND: Recent breakthroughs in molecular biology and next generation sequencing technologies have led to the expenential growh of the sequence databases. Researchrs use BLAST for processing these sequences. However traditional software parallelization techniques (threads, message passing interface) applied in newer versios of BLAST are not adequate for processing these sequences in timely manner. METHODS: A new method for array job parallelization has been developed which offers O(T) theoretical speed-up in comparison to multi-threading and MPI techniques. Here T is the number of array job tasks. (The number of CPUs that will be used to complete the job equals the product of T multiplied by the number of CPUs used by a single task.) The approach is based on segmentation of both input datasets to the BLAST process, combining partial solutions published earlier (Dhanker and Gupta, Int J Comput Sci Inf Technol_5:4818-4820, 2014), (Grant et al., Bioinformatics_18:765-766, 2002), (Mathog, Bioinformatics_19:1865-1866, 2003). It is accordingly referred to as a "dual segmentation" method. In order to implement the new method, the BLAST source code was modified to allow the researcher to pass to the program the number of records (effective number of sequences) in the original database. The team also developed methods to manage and consolidate the large number of partial results that get produced. Dual segmentation allows for massive parallelization, which lifts the scaling ceiling in exciting ways. RESULTS: BLAST jobs that hitherto failed or slogged inefficiently to completion now finish with speeds that characteristically reduce wallclock time from 27 days on 40 CPUs to a single day using 4104 tasks, each task utilizing eight CPUs and taking less than 7 minutes to complete. CONCLUSIONS: The massive increase in the number of tasks when running an analysis job with dual segmentation reduces the size, scope and execution time of each task. Besides significant speed of completion, additional benefits include fine-grained checkpointing and increased flexibility of job submission. "Trickling in" a swarm of individual small tasks tempers competition for CPU time in the shared HPC environment, and jobs submitted during quiet periods can complete in extraordinarily short time frames. The smaller task size also allows the use of older and less powerful hardware. The CDRH workhorse cluster was commissioned in 2010, yet its eight-core CPUs with only 24GB RAM work well in 2017 for these dual segmentation jobs. Finally, these techniques are excitingly friendly to budget conscious scientific research organizations where probabilistic algorithms such as BLAST might discourage attempts at greater certainty because single runs represent a major resource drain. If a job that used to take 24 days can now be completed in less than an hour or on a space available basis (which is the case at CDRH), repeated runs for more exhaustive analyses can be usefully contemplated.
Mike Mikailov, Fu-Jyh Luo, Stuart Barkley, Lohit Valleru, Stephen Whitney, Shraddha Thakkar, Weida Tong, Nicholas Petrick
BMC Bioinform.8
2016 An Image Analysis Environment for Species Identification of Food Contaminating Beetles
abstract
Food safety is vital to the well-being of society; therefore, it is important to inspect food products to ensure minimal health risks are present. The presence of certain species of insects, especially storage beetles, is a reliable indicator of possible contamination during storage and food processing. However, the current approach of identifying species by visual examination of insect fragments is rather subjective and time-consuming. To aid this inspection process, we have developed in collaboration with FDA food analysts some image analysis-based machine intelligence to achieve species identification with up to 90% accuracy. The current project is a continuation of this development effort. Here we present an image analysis environment that allows practical deployment of the machine intelligence on computers with limited processing power and memory. Using this environment, users can prepare input sets by selecting images for analysis, and inspect these images through the integrated panning and zooming capabilities. After species analysis, the results panel allows the user to compare the analyzed images with reference images of the proposed species. Further additions to this environment should include a log of previously analyzed images, and eventually extend to interaction with a central cloud repository of images through a web-based interface.
Hongjian Ding, Leihong Wu, Howard Semey, Amy Barnes, Darryl Langley, Su Inn Park, Weida Tong, Joshua Xu
AAAI9
2016 Application of dynamic topic models to toxicogenomics data
abstract
BACKGROUND: All biological processes are inherently dynamic. Biological systems evolve transiently or sustainably according to sequential time points after perturbation by environment insults, drugs and chemicals. Investigating the temporal behavior of molecular events has been an important subject to understand the underlying mechanisms governing the biological system in response to, such as, drug treatment. The intrinsic complexity of time series data requires appropriate computational algorithms for data interpretation. In this study, we propose, for the first time, the application of dynamic topic models (DTM) for analyzing time-series gene expression data. RESULTS: A large time-series toxicogenomics dataset was studied. It contains over 3144 microarrays of gene expression data corresponding to rat livers treated with 131 compounds (most are drugs) at two doses (control and high dose) in a repeated schedule containing four separate time points (4-, 8-, 15- and 29-day). We analyzed, with DTM, the topics (consisting of a set of genes) and their biological interpretations over these four time points. We identified hidden patterns embedded in this time-series gene expression profiles. From the topic distribution for compound-time condition, a number of drugs were successfully clustered by their shared mode-of-action such as PPARɑ agonists and COX inhibitors. The biological meaning underlying each topic was interpreted using diverse sources of information such as functional analysis of the pathways and therapeutic uses of the drugs. Additionally, we found that sample clusters produced by DTM are much more coherent in terms of functional categories when compared to traditional clustering algorithms. CONCLUSIONS: We demonstrated that DTM, a text mining technique, can be a powerful computational approach for clustering time-series gene expression profiles with the probabilistic representation of their dynamic features along sequential time frames. The method offers an alternative way for uncovering hidden patterns embedded in time series gene expression profiles to gain enhanced understanding of dynamic behavior of gene regulation in the biological system.
Mikyung Lee, Ruili Huang, Weida Tong
BMC Bioinform.4
2016 A novel procedure on next generation sequencing data analysis using text mining algorithm
abstract
BACKGROUND: Next-generation sequencing (NGS) technologies have provided researchers with vast possibilities in various biological and biomedical research areas. Efficient data mining strategies are in high demand for large scale comparative and evolutional studies to be performed on the large amounts of data derived from NGS projects. Topic modeling is an active research field in machine learning and has been mainly used as an analytical tool to structure large textual corpora for data mining. METHODS: We report a novel procedure to analyse NGS data using topic modeling. It consists of four major procedures: NGS data retrieval, preprocessing, topic modeling, and data mining using Latent Dirichlet Allocation (LDA) topic outputs. The NGS data set of the Salmonella enterica strains were used as a case study to show the workflow of this procedure. The perplexity measurement of the topic numbers and the convergence efficiencies of Gibbs sampling were calculated and discussed for achieving the best result from the proposed procedure. RESULTS: The output topics by LDA algorithms could be treated as features of Salmonella strains to accurately describe the genetic diversity of fliC gene in various serotypes. The results of a two-way hierarchical clustering and data matrix analysis on LDA-derived matrices successfully classified Salmonella serotypes based on the NGS data. The implementation of topic modeling in NGS data analysis procedure provides a new way to elucidate genetic information from NGS data, and identify the gene-phenotype relationships and biomarkers, especially in the era of biological and medical big data. CONCLUSION: The implementation of topic modeling in NGS data analysis provides a new way to elucidate genetic information from NGS data, and identify the gene-phenotype relationships and biomarkers, especially in the era of biological and medical big data.
Weizhong Zhao, James J. Chen, Roger Perkins, Huixiao Hong, Weida Tong, Wen Zou
BMC Bioinform.7
2016 Erratum to: A novel procedure on next generation sequencing data analysis using text mining algorithm
abstract
After publication of the original article [1] it was brought to our attention that the following was incorrectly placed under subheading '3.Classification analysis and comparison' of subsection 'Evaluation of topic modeling performance' of the 'Methods' section:Topic model-derived clustering method [33] was applied, in which LDA was utilized as a feature reduction approach for cluster analysis.The LDAderived topics were considered as the new features of datasets.The sampletopic matrix (Fig. 1(f)) was treated as a new representation of the original dataset.Based on the sample-topic matrix (topic number was chosen as 5 and 30, respectively), conventional clustering algorithms, such as k-means, was used for the clustering analysis.The number of clusters was set as 7 in the k-means method due to 7 different serotypes in the dataset.While in comparison, k-means algorithm was also applied on VSM matrix using Hamming Distance similarities.For further comparison, due to the dimension reduction of topic modeling approach, the traditional tool of PCA was used to reduce features (Numbers of 2, 5, 10 and 30 were randomly selected as the reduced features, respectively) of VSM matrix followed by the k-means cluster analysis.Moreover, clustering by only LDA referred as "highest probable topic assignment" [33] (5 and 30 topics were used) was also used for comparison.In "highest probable topic assignment", the LDA-derived topics were made as the clusters of the dataset.Then, each sample was assigned to the cluster (Topic) with the highest probability in the row of the sample-topic matrix.To interpret the clustering results obtained by the k-means algorithm, samples in each cluster were labeled as the dominant serotype of the samples in the cluster.
Weizhong Zhao, James J. Chen, Roger Perkins, Huixiao Hong, Weida Tong, Wen Zou
BMC Bioinform.7
2015 Understanding and predicting binding between human leukocyte antigens (HLAs) and peptides by network analysis
abstract
BACKGROUND: As the major histocompatibility complex (MHC), human leukocyte antigens (HLAs) are one of the most polymorphic genes in humans. Patients carrying certain HLA alleles may develop adverse drug reactions (ADRs) after taking specific drugs. Peptides play an important role in HLA related ADRs as they are the necessary co-binders of HLAs with drugs. Many experimental data have been generated for understanding HLA-peptide binding. However, efficiently utilizing the data for understanding and accurately predicting HLA-peptide binding is challenging. Therefore, we developed a network analysis based method to understand and predict HLA-peptide binding. METHODS: Qualitative Class I HLA-peptide binding data were harvested and prepared from four major databases. An HLA-peptide binding network was constructed from this dataset and modules were identified by the fast greedy modularity optimization algorithm. To examine the significance of signals in the yielded models, the modularity was compared with the modularity values generated from 1,000 random networks. The peptides and HLAs in the modules were characterized by similarity analysis. The neighbor-edges based and unbiased leverage algorithm (Nebula) was developed for predicting HLA-peptide binding. Leave-one-out (LOO) validations and two-fold cross-validations were conducted to evaluate the performance of Nebula using the constructed HLA-peptide binding network. RESULTS: Nine modules were identified from analyzing the HLA-peptide binding network with a highest modularity compared to all the random networks. Peptide length and functional side chains of amino acids at certain positions of the peptides were different among the modules. HLA sequences were module dependent to some extent. Nebula archived an overall prediction accuracy of 0.816 in the LOO validations and average accuracy of 0.795 in the two-fold cross-validations and outperformed the method reported in the literature. CONCLUSIONS: Network analysis is a useful approach for analyzing large and sparse datasets such as the HLA-peptide binding dataset. The modules identified from the network analysis clustered peptides and HLAs with similar sequences and properties of amino acids. Nebula performed well in the predictions of HLA-peptide binding. We demonstrated that network analysis coupled with Nebula is an efficient approach to understand and predict HLA-peptide binding interactions and thus, could further our understanding of ADRs.
Heng Luo 0002, Hui Wen Ng, Leming Shi, Weida Tong, William Mattes, Donna Mendrick, Huixiao Hong
BMC Bioinform.5
2014 A phenome-guided drug repositioning through a latent variable model
abstract
BACKGROUND: The phenome represents a distinct set of information in the human population. It has been explored particularly in its relationship with the genome to identify correlations for diseases. The phenome has been also explored for drug repositioning with efforts focusing on the search space for the most similar candidate drugs. For a comprehensive analysis of the phenome, we assumed that all phenotypes (indications and side effects) were inter-connected with a probabilistic distribution and this characteristic may offer an opportunity to identify new therapeutic indications for a given drug. Correspondingly, we employed Latent Dirichlet Allocation (LDA), which introduces latent variables (topics) to govern the phenome distribution. RESULTS: We developed our model on the phenome information in Side Effect Resource (SIDER). We first developed a LDA model optimized based on its recovery potential through perturbing the drug-phenotype matrix for each of the drug-indication pairs where each drug-indication relationship was switched to "unknown" one at the time and then recovered based on the remaining drug-phenotype pairs. Of the probabilistically significant pairs, 70% was successfully recovered. Next, we applied the model on the whole phenome to narrow down repositioning candidates and suggest alternative indications. We were able to retrieve approved indications of 6 drugs whose indications were not listed in SIDER. For 908 drugs that were present with their indication information, our model suggested alternative treatment options for further investigations. Several of the suggested new uses can be supported with information from the scientific literature. CONCLUSIONS: The results demonstrated that the phenome can be further analyzed by a generative model, which can discover probabilistic associations between drugs and therapeutic uses. In this regard, LDA serves as an enrichment tool to explore new uses of existing drugs by narrowing down the search space.
Halil Bisgin, Reagan J. Kelly, Xiaowei Xu 0001, Weida Tong
BMC Bioinform.6
2014 Competitive molecular docking approach for predicting estrogen receptor subtype α agonists and antagonists
abstract
BACKGROUND: Endocrine disrupting chemicals (EDCs) are exogenous compounds that interfere with the endocrine system of vertebrates, often through direct or indirect interactions with nuclear receptor proteins. Estrogen receptors (ERs) are particularly important protein targets and many EDCs are ER binders, capable of altering normal homeostatic transcription and signaling pathways. An estrogenic xenobiotic can bind ER as either an agonist or antagonist to increase or inhibit transcription, respectively. The receptor conformations in the complexes of ER bound with agonists and antagonists are different and dependent on interactions with co-regulator proteins that vary across tissue type. Assessment of chemical endocrine disruption potential depends not only on binding affinity to ERs, but also on changes that may alter the receptor conformation and its ability to subsequently bind DNA response elements and initiate transcription. Using both agonist and antagonist conformations of the ERα, we developed an in silico approach that can be used to differentiate agonist versus antagonist status of potential binders. METHODS: The approach combined separate molecular docking models for ER agonist and antagonist conformations. The ability of this approach to differentiate agonists and antagonists was first evaluated using true agonists and antagonists extracted from the crystal structures available in the protein data bank (PDB), and then further validated using a larger set of ligands from the literature. The usefulness of the approach was demonstrated with enrichment analysis in data sets with a large number of decoy ligands. RESULTS: The performance of individual agonist and antagonist docking models was found comparable to similar models in the literature. When combined in a competitive docking approach, they provided the ability to discriminate agonists from antagonists with good accuracy, as well as the ability to efficiently select true agonists and antagonists from decoys during enrichment analysis. CONCLUSION: This approach enables evaluation of potential ER biological function changes caused by chemicals bound to the receptor which, in turn, allows the assessment of a chemical's endocrine disrupting potential. The approach can be used not only by regulatory authorities to perform risk assessments on potential EDCs but also by the industry in drug discovery projects to screen for potential agonists and antagonists.
Hui Wen Ng, Mao Shu, Heng Luo 0002, Weigong Ge, Roger Perkins, Weida Tong, Huixiao Hong
BMC Bioinform.7
2014 Advances in translational bioinformatics facilitate revealing the landscape of complex disease mechanisms
abstract
Advances of high-throughput technologies have rapidly produced more and more data from DNAs and RNAs to proteins, especially large volumes of genome-scale data. However, connection of the genomic information to cellular functions and biological behaviours relies on the development of effective approaches at higher systems level. In particular, advances in RNA-Seq technology has helped the studies of transcriptome, RNA expressed from the genome, while systems biology on the other hand provides more comprehensive pictures, from which genes and proteins actively interact to lead to cellular behaviours and physiological phenotypes. As biological interactions mediate many biological processes that are essential for cellular function or disease development, it is important to systematically identify genomic information including genetic mutations from GWAS (genome-wide association study), differentially expressed genes, bidirectional promoters, intrinsic disordered proteins (IDP) and protein interactions to gain deep insights into the underlying mechanisms of gene regulations and networks. Furthermore, bidirectional promoters can co-regulate many biological pathways, where the roles of bidirectional promoters can be studied systematically for identifying co-regulating genes at interactive network level. Combining information from different but related studies can ultimately help revealing the landscape of molecular mechanisms underlying complex diseases such as cancer.
Jack Y. Yang, A. Keith Dunker, Jun S. Liu, Xiang Qin, Hamid R. Arabnia, William Yang, Andrzej Niemierko, Zhongxue Chen, Zuojie Luo, Liangjiang Wang, Youping Deng, Weida Tong, Mary Yang
BMC Bioinform.14
2014 Identification of genes and pathways involved in kidney renal clear cell carcinoma
abstract
BACKGROUND: Kidney Renal Clear Cell Carcinoma (KIRC) is one of fatal genitourinary diseases and accounts for most malignant kidney tumours. KIRC has been shown resistance to radiotherapy and chemotherapy. Like many types of cancers, there is no curative treatment for metastatic KIRC. Using advanced sequencing technologies, The Cancer Genome Atlas (TCGA) project of NIH/NCI-NHGRI has produced large-scale sequencing data, which provide unprecedented opportunities to reveal new molecular mechanisms of cancer. We combined differentially expressed genes, pathways and network analyses to gain new insights into the underlying molecular mechanisms of the disease development. RESULTS: Followed by the experimental design for obtaining significant genes and pathways, comprehensive analysis of 537 KIRC patients' sequencing data provided by TCGA was performed. Differentially expressed genes were obtained from the RNA-Seq data. Pathway and network analyses were performed. We identified 186 differentially expressed genes with significant p-value and large fold changes (P < 0.01, |log(FC)| > 5). The study not only confirmed a number of identified differentially expressed genes in literature reports, but also provided new findings. We performed hierarchical clustering analysis utilizing the whole genome-wide gene expressions and differentially expressed genes that were identified in this study. We revealed distinct groups of differentially expressed genes that can aid to the identification of subtypes of the cancer. The hierarchical clustering analysis based on gene expression profile and differentially expressed genes suggested four subtypes of the cancer. We found enriched distinct Gene Ontology (GO) terms associated with these groups of genes. Based on these findings, we built a support vector machine based supervised-learning classifier to predict unknown samples, and the classifier achieved high accuracy and robust classification results. In addition, we identified a number of pathways (P < 0.04) that were significantly influenced by the disease. We found that some of the identified pathways have been implicated in cancers from literatures, while others have not been reported in the cancer before. The network analysis leads to the identification of significantly disrupted pathways and associated genes involved in the disease development. Furthermore, this study can provide a viable alternative in identifying effective drug targets. CONCLUSIONS: Our study identified a set of differentially expressed genes and pathways in kidney renal clear cell carcinoma, and represents a comprehensive computational approach to analysis large-scale next-generation sequencing data. The pathway and network analyses suggested that information from distinctly expressed genes can be utilized in the identification of aberrant upstream regulators. Identification of distinctly expressed genes and altered pathways are important in effective biomarker identification for early cancer diagnosis and treatment planning. Combining differentially expressed genes with pathway and network analyses using intelligent computational approaches provide an unprecedented opportunity to identify upstream disease causal genes and effective drug targets.
William Yang, Kenji Yoshigoe, Xiang Qin, Jun S. Liu, Jack Y. Yang, Andrzej Niemierko, Youping Deng, A. Keith Dunker, Zhongxue Chen, Liangjiang Wang, Hamid R. Arabnia, Weida Tong, Mary Yang
BMC Bioinform.14
2014 Mining hidden knowledge for drug safety assessment: topic modeling of LiverTox as a case study
abstract
BACKGROUND: Given the significant impact on public health and drug development, drug safety has been a focal point and research emphasis across multiple disciplines in addition to scientific investigation, including consumer advocates, drug developers and regulators. Such a concern and effort has led numerous databases with drug safety information available in the public domain and the majority of them contain substantial textual data. Text mining offers an opportunity to leverage the hidden knowledge within these textual data for the enhanced understanding of drug safety and thus improving public health. METHODS: In this proof-of-concept study, topic modeling, an unsupervised text mining approach, was performed on the LiverTox database developed by National Institutes of Health (NIH). The LiverTox structured one document per drug that contains multiple sections summarizing clinical information on drug-induced liver injury (DILI). We hypothesized that these documents might contain specific textual patterns that could be used to address key DILI issues. We placed the study on drug-induced acute liver failure (ALF) which was a severe form of DILI with limited treatment options. RESULTS: After topic modeling of the "Hepatotoxicity" sections of the LiverTox across 478 drug documents, we identified a hidden topic relevant to Hy's law that was a widely-accepted rule incriminating drugs with high risk of causing ALF in humans. Using this topic, a total of 127 drugs were further implicated, 77 of which had clear ALF relevant terms in the "Outcome and management" sections of the LiverTox. For the rest of 50 drugs, evidence supporting risk of ALF was found for 42 drugs from other public databases. CONCLUSION: In this case study, the knowledge buried in the textual data was extracted for identification of drugs with potential of causing ALF by applying topic modeling to the LiverTox database. The knowledge further guided identification of drugs with the similar potential and most of them could be verified and confirmed. This study highlights the utility of topic modeling to leverage information within textual drug safety databases, which provides new opportunities in the big data era to assess drug safety.
Minjun Chen, Xiaowei Xu 0001, Ayako Suzuki, Katarina Ilic, Weida Tong
BMC Bioinform.7
2014 Whole genome sequencing of 35 individuals provides insights into the genetic architecture of Korean population
abstract
BACKGROUND: Due to a significant decline in the costs associated with next-generation sequencing, it has become possible to decipher the genetic architecture of a population by sequencing a large number of individuals to a deep coverage. The Korean Personal Genomes Project (KPGP) recently sequenced 35 Korean genomes at high coverage using the Illumina Hiseq platform and made the deep sequencing data publicly available, providing the scientific community opportunities to decipher the genetic architecture of the Korean population. METHODS: In this study, we used two single nucleotide variant (SNV) calling pipelines: mapping the raw reads obtained from whole genome sequencing of 35 Korean individuals in KPGP using BWA and SOAP2 followed by SNV calling using SAMtools and SOAPsnp, respectively. The consensus SNVs obtained from the two SNV pipelines were used to represent the SNVs of the Korean population. We compared these SNVs to those from 17 other populations provided by the HapMap consortium and the 1000 Genomes Project (1KGP) and identified SNVs that were only present in the Korean population. We studied the mutation spectrum and analyzed the genes of non-synonymous SNVs only detected in the Korean population. RESULTS: We detected a total of 8,555,726 SNVs in the 35 Korean individuals and identified 1,213,613 SNVs detected in at least one Korean individual (SNV-1) and 12,640 in all of 35 Korean individuals (SNV-35) but not in 17 other populations. In contrast with the SNVs common to other populations in HapMap and 1KGP, the Korean only SNVs had high percentages of non-silent variants, emphasizing the unique roles of these Korean only SNVs in the Korean population. Specifically, we identified 8,361 non-synonymous Korean only SNVs, of which 58 SNVs existed in all 35 Korean individuals. The 5,754 genes of non-synonymous Korean only SNVs were highly enriched in some metabolic pathways. We found adhesion is the top disease term associated with SNV-1 and Nelson syndrome is the only disease term associated with SNV-35. We found that a significant number of Korean only SNVs are in genes that are associated with the drug term of adenosine. CONCLUSION: We identified the SNVs that were found in the Korean population but not seen in other populations, and explored the corresponding genes and pathways as well as the associated disease terms and drug terms. The results expand our knowledge of the genetic architecture of the Korean population, which will benefit the implementation of personalized medicine for the Korean population.
Joe Meehan, Zhenqiang Su, Hui Wen Ng, Mao Shu, Heng Luo 0002, Weigong Ge, Roger Perkins, Weida Tong, Huixiao Hong
BMC Bioinform.9
2013 A systems approach for analysis of high content screening assay data with topic modeling
abstract
BACKGROUND: High Content Screening (HCS) has become an important tool for toxicity assessment, partly due to its advantage of handling multiple measurements simultaneously. This approach has provided insight and contributed to the understanding of systems biology at cellular level. To fully realize this potential, the simultaneously measured multiple endpoints from a live cell should be considered in a probabilistic relationship to assess the cell's condition to response stress from a treatment, which poses a great challenge to extract hidden knowledge and relationships from these measurements. METHOD: In this work, we applied a text mining method of Latent Dirichlet Allocation (LDA) to analyze cellular endpoints from in vitro HCS assays and related to the findings to in vivo histopathological observations. We measured multiple HCS assay endpoints for 122 drugs. Since LDA requires the data to be represented in document-term format, we first converted the continuous value of the measurements to the word frequency that can processed by the text mining tool. For each of the drugs, we generated a document for each of the 4 time points. Thus, we ended with 488 documents (drug-hour) each having different values for the 10 endpoints which are treated as words. We extracted three topics using LDA and examined these to identify diagnostic topics for 45 common drugs located in vivo experiments from the Japanese Toxicogenomics Project (TGP) observing their necrosis findings at 6 and 24 hours after treatment. RESULTS: We found that assay endpoints assigned to particular topics were in concordance with the histopathology observed. Drugs showing necrosis at 6 hour were linked to severe damage events such as Steatosis, DNA Fragmentation, Mitochondrial Potential, and Lysosome Mass. DNA Damage and Apoptosis were associated with drugs causing necrosis at 24 hours, suggesting an interplay of the two pathways in these drugs. Drugs with no sign of necrosis we related to the Cell Loss and Nuclear Size assays, which is suggestive of hepatocyte regeneration. CONCLUSIONS: The evidence from this study suggests that topic modeling with LDA can enable us to interpret relationships of endpoints of in vitro assays along with an in vivo histological finding, necrosis. Effectiveness of this approach may add substantially to our understanding of systems biology.
Halil Bisgin, Minjun Chen, Reagan J. Kelly, Xiaowei Xu 0001, Weida Tong
BMC Bioinform.7
2013 Homology modeling, molecular docking, and molecular dynamics simulations elucidated α-fetoprotein binding modes
abstract
BACKGROUND: An important mechanism of endocrine activity is chemicals entering target cells via transport proteins and then interacting with hormone receptors such as the estrogen receptor (ER). α-Fetoprotein (AFP) is a major transport protein in rodent serum that can bind and sequester estrogens, thus preventing entry to the target cell and where they could otherwise induce ER-mediated endocrine activity. Recently, we reported rat AFP binding affinities for a large set of structurally diverse chemicals, including 53 binders and 72 non-binders. However, the lack of three-dimensional (3D) structures of rat AFP hinders further understanding of the structural dependence for binding. Therefore, a 3D structure of rat AFP was built using homology modeling in order to elucidate rat AFP-ligand binding modes through docking analyses and molecular dynamics (MD) simulations. METHODS: Homology modeling was first applied to build a 3D structure of rat AFP. Molecular docking and Molecular Mechanics-Generalized Born Surface Area (MM-GBSA) scoring were then used to examine potential rat AFP ligand binding modes. MD simulations and free energy calculations were performed to refine models of binding modes. RESULTS: A rat AFP tertiary structure was first obtained using homology modeling and MD simulations. The rat AFP-ligand binding modes of 13 structurally diverse, representative binders were calculated using molecular docking, (MM-GBSA) ranking and MD simulations. The key residues for rat AFP-ligand binding were postulated through analyzing the binding modes. CONCLUSION: The optimized 3D rat AFP structure and associated ligand binding modes shed light on rat AFP-ligand binding interactions that, in turn, provide a means to estimate binding affinity of unknown chemicals. Our results will assist in the evaluation of the endocrine disruption potential of chemicals.
Roger Perkins, Weida Tong, Huixiao Hong
BMC Bioinform.5
2012 Investigating drug repositioning opportunities in FDA drug labels through topic modeling
abstract
BACKGROUND: Drug repositioning offers an opportunity to revitalize the slowing drug discovery pipeline by finding new uses for currently existing drugs. Our hypothesis is that drugs sharing similar side effect profiles are likely to be effective for the same disease, and thus repositioning opportunities can be identified by finding drug pairs with similar side effects documented in U.S. Food and Drug Administration (FDA) approved drug labels. The safety information in the drug labels is usually obtained in the clinical trial and augmented with the observations in the post-market use of the drug. Therefore, our drug repositioning approach can take the advantage of more comprehensive safety information comparing with conventional de novo approach. METHOD: A probabilistic topic model was constructed based on the terms in the Medical Dictionary for Regulatory Activities (MedDRA) that appeared in the Boxed Warning, Warnings and Precautions, and Adverse Reactions sections of the labels of 870 drugs. Fifty-two unique topics, each containing a set of terms, were identified by using topic modeling. The resulting probabilistic topic associations were used to measure the distance (similarity) between drugs. The success of the proposed model was evaluated by comparing a drug and its nearest neighbor (i.e., a drug pair) for common indications found in the Indications and Usage Section of the drug labels. RESULTS: Given a drug with more than three indications, the model yielded a 75% recall, meaning 75% of drug pairs shared one or more common indications. This is significantly higher than the 22% recall rate achieved by random selection. Additionally, the recall rate grows rapidly as the number of drug indications increases and reaches 84% for drugs with 11 indications. The analysis also demonstrated that 65 drugs with a Boxed Warning, which indicates significant risk of serious and possibly life-threatening adverse effects, might be replaced with safer alternatives that do not have a Boxed Warning. In addition, we identified two therapeutic groups of drugs (Musculo-skeletal system and Anti-infective for systemic use) where over 80% of the drugs have a potential replacement with high significance. CONCLUSION: Topic modeling can be a powerful tool for the identification of repositioning opportunities by examining the adverse event terms in FDA approved drug labels. The proposed framework not only suggests drugs that can be repurposed, but also provides insight into the safety of repositioned drugs.
Halil Bisgin, Reagan J. Kelly, Xiaowei Xu 0001, Weida Tong
BMC Bioinform.6
2011 Mining FDA drug labels using an unsupervised learning technique - topic modeling
abstract
BACKGROUND: The Food and Drug Administration (FDA) approved drug labels contain a broad array of information, ranging from adverse drug reactions (ADRs) to drug efficacy, risk-benefit consideration, and more. However, the labeling language used to describe these information is free text often containing ambiguous semantic descriptions, which poses a great challenge in retrieving useful information from the labeling text in a consistent and accurate fashion for comparative analysis across drugs. Consequently, this task has largely relied on the manual reading of the full text by experts, which is time consuming and labor intensive. METHOD: In this study, a novel text mining method with unsupervised learning in nature, called topic modeling, was applied to the drug labeling with a goal of discovering "topics" that group drugs with similar safety concerns and/or therapeutic uses together. A total of 794 FDA-approved drug labels were used in this study. First, the three labeling sections (i.e., Boxed Warning, Warnings and Precautions, Adverse Reactions) of each drug label were processed by the Medical Dictionary for Regulatory Activities (MedDRA) to convert the free text of each label to the standard ADR terms. Next, the topic modeling approach with latent Dirichlet allocation (LDA) was applied to generate 100 topics, each associated with a set of drugs grouped together based on the probability analysis. Lastly, the efficacy of the topic modeling was evaluated based on known information about the therapeutic uses and safety data of drugs. RESULTS: The results demonstrate that drugs grouped by topics are associated with the same safety concerns and/or therapeutic uses with statistical significance (P<0.05). The identified topics have distinct context that can be directly linked to specific adverse events (e.g., liver injury or kidney injury) or therapeutic application (e.g., antiinfectives for systemic use). We were also able to identify potential adverse events that might arise from specific medications via topics. CONCLUSIONS: The successful application of topic modeling on the FDA drug labeling demonstrates its potential utility as a hypothesis generation means to infer hidden relationships of concepts such as, in this study, drug safety and therapeutic use in the study of biomedical documents.
Halil Bisgin, Xiaowei Xu 0001, Weida Tong
BMC Bioinform.5
2011 Selecting a single model or combining multiple models for microarray-based classifier development? - A comparative analysis based on large and diverse datasets generated from the MAQC-II project
abstract
BACKGROUND: Genomic biomarkers play an increasing role in both preclinical and clinical application. Development of genomic biomarkers with microarrays is an area of intensive investigation. However, despite sustained and continuing effort, developing microarray-based predictive models (i.e., genomics biomarkers) capable of reliable prediction for an observed or measured outcome (i.e., endpoint) of unknown samples in preclinical and clinical practice remains a considerable challenge. No straightforward guidelines exist for selecting a single model that will perform best when presented with unknown samples. In the second phase of the MicroArray Quality Control (MAQC-II) project, 36 analysis teams produced a large number of models for 13 preclinical and clinical endpoints. Before external validation was performed, each team nominated one model per endpoint (referred to here as 'nominated models') from which MAQC-II experts selected 13 'candidate models' to represent the best model for each endpoint. Both the nominated and candidate models from MAQC-II provide benchmarks to assess other methodologies for developing microarray-based predictive models. METHODS: We developed a simple ensemble method by taking a number of the top performing models from cross-validation and developing an ensemble model for each of the MAQC-II endpoints. We compared the ensemble models with both nominated and candidate models from MAQC-II using blinded external validation. RESULTS: For 10 of the 13 MAQC-II endpoints originally analyzed by the MAQC-II data analysis team from the National Center for Toxicological Research (NCTR), the ensemble models achieved equal or better predictive performance than the NCTR nominated models. Additionally, the ensemble models had performance comparable to the MAQC-II candidate models. Most ensemble models also had better performance than the nominated models generated by five other MAQC-II data analysis teams that analyzed all 13 endpoints. CONCLUSIONS: Our findings suggest that an ensemble method can often attain a higher average predictive performance in an external validation set than a corresponding "optimized" model method. Using an ensemble method to determine a final model is a potentially important supplement to the good modeling practices recommended by the MAQC-II project for developing microarray-based genomic biomarkers.
Minjun Chen, Leming Shi, Reagan J. Kelly, Roger Perkins, Weida Tong
BMC Bioinform.6
2011 Constructing a robust protein-protein interaction network by integrating multiple public databases
abstract
BACKGROUND: Protein-protein interactions (PPIs) are a critical component for many underlying biological processes. A PPI network can provide insight into the mechanisms of these processes, as well as the relationships among different proteins and toxicants that are potentially involved in the processes. There are many PPI databases publicly available, each with a specific focus. The challenge is how to effectively combine their contents to generate a robust and biologically relevant PPI network. METHODS: In this study, seven public PPI databases, BioGRID, DIP, HPRD, IntAct, MINT, REACTOME, and SPIKE, were used to explore a powerful approach to combine multiple PPI databases for an integrated PPI network. We developed a novel method called k-votes to create seven different integrated networks by using values of k ranging from 1-7. Functional modules were mined by using SCAN, a Structural Clustering Algorithm for Networks. Overall module qualities were evaluated for each integrated network using the following statistical and biological measures: (1) modularity, (2) similarity-based modularity, (3) clustering score, and (4) enrichment. RESULTS: Each integrated human PPI network was constructed based on the number of votes (k) for a particular interaction from the committee of the original seven PPI databases. The performance of functional modules obtained by SCAN from each integrated network was evaluated. The optimal value for k was determined by the functional module analysis. Our results demonstrate that the k-votes method outperforms the traditional union approach in terms of both statistical significance and biological meaning. The best network is achieved at k = 2, which is composed of interactions that are confirmed in at least two PPI databases. In contrast, the traditional union approach yields an integrated network that consists of all interactions of seven PPI databases, which might be subject to high false positives. CONCLUSIONS: We determined that the k-votes method for constructing a robust PPI network by integrating multiple public databases outperforms previously reported approaches and that a value of k=2 provides the best results. The developed strategies for combining databases show promise in the advancement of network construction and modeling.
Martha VenkataSwamy, Li Guo 0009, Zhenqiang Su, Yanbin Ye, Don Ding, Weida Tong, Xiaowei Xu 0001
BMC Bioinform.8
2011 Translating Clinical Findings into Knowledge in Drug Safety Evaluation - Drug Induced Liver Injury Prediction System (DILIps)
abstract
Drug-induced liver injury (DILI) is a significant concern in drug development due to the poor concordance between preclinical and clinical findings of liver toxicity. We hypothesized that the DILI types (hepatotoxic side effects) seen in the clinic can be translated into the development of predictive in silico models for use in the drug discovery phase. We identified 13 hepatotoxic side effects with high accuracy for classifying marketed drugs for their DILI potential. We then developed in silico predictive models for each of these 13 side effects, which were further combined to construct a DILI prediction system (DILIps). The DILIps yielded 60-70% prediction accuracy for three independent validation sets. To enhance the confidence for identification of drugs that cause severe DILI in humans, the "Rule of Three" was developed in DILIps by using a consensus strategy based on 13 models. This gave high positive predictive value (91%) when applied to an external dataset containing 206 drugs from three independent literature datasets. Using the DILIps, we screened all the drugs in DrugBank and investigated their DILI potential in terms of protein targets and therapeutic categories through network modeling. We demonstrated that two therapeutic categories, anti-infectives for systemic use and musculoskeletal system drugs, were enriched for DILI, which is consistent with current knowledge. We also identified protein targets and pathways that are related to drugs that cause DILI by using pathway analysis and co-occurrence text mining. While marketed drugs were the focus of this study, the DILIps has a potential as an evaluation tool to screen and prioritize new drug candidates or chemicals, such as environmental chemicals, to avoid those that might cause liver toxicity. We expect that the methodology can be also applied to other drug safety endpoints, such as renal or cardiovascular toxicity.
Don Ding, Reagan J. Kelly, Weida Tong
PLoS Comput. Biol.6
2010 ISA software suite: supporting standards-compliant experimental annotation and enabling curation at the community level
abstract
UNLABELLED: The first open source software suite for experimentalists and curators that (i) assists in the annotation and local management of experimental metadata from high-throughput studies employing one or a combination of omics and other technologies; (ii) empowers users to uptake community-defined checklists and ontologies; and (iii) facilitates submission to international public repositories. AVAILABILITY AND IMPLEMENTATION: Software, documentation, case studies and implementations at http://www.isa-tools.org.
Philippe Rocca-Serra, Marco Brandizi, Eamonn Maguire, Nataliya Sklyar, Chris F. Taylor, Kimberly Begley, Dawn Field, Stephen C. Harris, Winston Hide, Oliver Hofmann 0001, Steffen Neumann, Peter Sterk, Weida Tong, Susanna-Assunta Sansone
Bioinform.13
2010 The EDKB: an established knowledge base for endocrine disrupting chemicals
abstract
BACKGROUND: Endocrine disruptors (EDs) and their broad range of potential adverse effects in humans and other animals have been a concern for nearly two decades. Many putative EDs are widely used in commercial products regulated by the Food and Drug Administration (FDA) such as food packaging materials, ingredients of cosmetics, medical and dental devices, and drugs. The Endocrine Disruptor Knowledge Base (EDKB) project was initiated in the mid 1990's by the FDA as a resource for the study of EDs. The EDKB database, a component of the project, contains data across multiple assay types for chemicals across a broad structural diversity. This paper demonstrates the utility of EDKB database, an integral part of the EDKB project, for understanding and prioritizing EDs for testing. RESULTS: The EDKB database currently contains 3,257 records of over 1,800 EDs from different assays including estrogen receptor binding, androgen receptor binding, uterotropic activity, cell proliferation, and reporter gene assays. Information for each compound such as chemical structure, assay type, potency, etc. is organized to enable efficient searching. A user-friendly interface provides rapid navigation, Boolean searches on EDs, and both spreadsheet and graphical displays for viewing results. The search engine implemented in the EDKB database enables searching by one or more of the following fields: chemical structure (including exact search and similarity search), name, molecular formula, CAS registration number, experiment source, molecular weight, etc. The data can be cross-linked to other publicly available and related databases including TOXNET, Cactus, ChemIDplus, ChemACX, Chem Finder, and NCI DTP. CONCLUSION: The EDKB database enables scientists and regulatory reviewers to quickly access ED data from multiple assays for specific or similar compounds. The data have been used to categorize chemicals according to potential risks for endocrine activity, thus providing a basis for prioritizing chemicals for more definitive but expensive testing. The EDKB database is publicly available and can be found online at http://edkb.fda.gov/webstart/edkb/index.html.
Don Ding, Huixiao Hong, Roger Perkins, Steve Harris, Edward D. Bearden, Leming Shi, Weida Tong
BMC Bioinform.9
2010 An FDA bioinformatics tool for microbial genomics research on molecular characterization of bacterial foodborne pathogens using microarrays
abstract
BACKGROUND: Advances in microbial genomics and bioinformatics are offering greater insights into the emergence and spread of foodborne pathogens in outbreak scenarios. The Food and Drug Administration (FDA) has developed a genomics tool, ArrayTrack™, which provides extensive functionalities to manage, analyze, and interpret genomic data for mammalian species. ArrayTrack™ has been widely adopted by the research community and used for pharmacogenomics data review in the FDA's Voluntary Genomics Data Submission program. RESULTS: ArrayTrack™ has been extended to manage and analyze genomics data from bacterial pathogens of human, animal, and food origin. It was populated with bioinformatics data from public databases such as NCBI, Swiss-Prot, KEGG Pathway, and Gene Ontology to facilitate pathogen detection and characterization. ArrayTrack™'s data processing and visualization tools were enhanced with analysis capabilities designed specifically for microbial genomics including flag-based hierarchical clustering analysis (HCA), flag concordance heat maps, and mixed scatter plots. These specific functionalities were evaluated on data generated from a custom Affymetrix array (FDA-ECSG) previously developed within the FDA. The FDA-ECSG array represents 32 complete genomes of Escherichia coli and Shigella. The new functions were also used to analyze microarray data focusing on antimicrobial resistance genes from Salmonella isolates in a poultry production environment using a universal antimicrobial resistance microarray developed by the United States Department of Agriculture (USDA). CONCLUSION: The application of ArrayTrack™ to different microarray platforms demonstrates its utility in microbial genomics research, and thus will improve the capabilities of the FDA to rapidly identify foodborne bacteria and their genetic traits (e.g., antimicrobial resistance, virulence, etc.) during outbreak investigations. ArrayTrack™ is free to use and available to public, private, and academic researchers at http://www.fda.gov/ArrayTrack.
Joshua Xu, Don Ding, Scott A. Jackson, Isha R. Patel, Jonathan G. Frye, Wen Zou, Rajesh Nayak, Steven L. Foley, James J. Chen, Zhenqiang Su, Yanbin Ye, Steve Turner, Steve Harris, Guangxu Zhou, Carl Cerniglia, Weida Tong
BMC Bioinform.17
2010 Evaluation of gene expression data generated from expired Affymetrix GeneChip® microarrays using MAQC reference RNA samples
abstract
BACKGROUND: The Affymetrix GeneChip® system is a commonly used platform for microarray analysis but the technology is inherently expensive. Unfortunately, changes in experimental planning and execution, such as the unavailability of previously anticipated samples or a shift in research focus, may render significant numbers of pre-purchased GeneChip® microarrays unprocessed before their manufacturer's expiration dates. Researchers and microarray core facilities wonder whether expired microarrays are still useful for gene expression analysis. In addition, it was not clear whether the two human reference RNA samples established by the MAQC project in 2005 still maintained their transcriptome integrity over a period of four years. Experiments were conducted to answer these questions. RESULTS: Microarray data were generated in 2009 in three replicates for each of the two MAQC samples with either expired Affymetrix U133A or unexpired U133Plus2 microarrays. These results were compared with data obtained in 2005 on the U133Plus2 microarray. The percentage of overlap between the lists of differentially expressed genes (DEGs) from U133Plus2 microarray data generated in 2009 and in 2005 was 97.44%. While there was some degree of fold change compression in the expired U133A microarrays, the percentage of overlap between the lists of DEGs from the expired and unexpired microarrays was as high as 96.99%. Moreover, the microarray data generated using the expired U133A microarrays in 2009 were highly concordant with microarray and TaqMan® data generated by the MAQC project in 2005. CONCLUSIONS: Our results demonstrated that microarray data generated using U133A microarrays, which were more than four years past the manufacturer's expiration date, were highly specific and consistent with those from unexpired microarrays in identifying DEGs despite some appreciable fold change compression and decrease in sensitivity. Our data also suggested that the MAQC reference RNA samples, stored at -80°C, were stable over a time frame of at least four years.
Zhining Wen, Zhenqiang Su, Huixiao Hong, Weida Tong, Leming Shi
BMC Bioinform.7
2010 Two new ArrayTrack libraries for personalized biomedical research
abstract
BACKGROUND: Recent advances in high-throughput genotyping technology are paving the way for research in personalized medicine and nutrition. However, most of the genetic markers identified from association studies account for a small contribution to the total risk/benefit of the studied phenotypic trait. Testing whether the candidate genes identified by association studies are causal is critically important to the development of personalized medicine and nutrition. An efficient data mining strategy and a set of sophisticated tools are necessary to help better understand and utilize the findings from genetic association studies. DESCRIPTION: SNP (single nucleotide polymorphism) and QTL (quantitative trait locus) libraries were constructed and incorporated into ArrayTrack, with user-friendly interfaces and powerful search features. Data from several public repositories were collected in the SNP and QTL libraries and connected to other domain libraries (genes, proteins, metabolites, and pathways) in ArrayTrack. Linking the data sets within ArrayTrack allows searching of SNP and QTL data as well as their relationships to other biological molecules. The SNP library includes approximately 15 million human SNPs and their annotations, while the QTL library contains publically available QTLs identified in mouse, rat, and human. The QTL library was developed for finding the overlap between the map position of a candidate or metabolic gene and QTLs from these species. Two use cases were included to demonstrate the utility of these tools. The SNP and QTL libraries are freely available to the public through ArrayTrack at http://www.fda.gov/ArrayTrack. CONCLUSIONS: These libraries developed in ArrayTrack contain comprehensive information on SNPs and QTLs and are further cross-linked to other libraries. Connecting domain specific knowledge is a cornerstone of systems biology strategies and allows for a better understanding of the genetic and biological context of the findings from genetic association studies.
Joshua Xu, Carolyn Wise, Vijayalakshmi Varma, Baitang Ning, Huixiao Hong, Weida Tong, Jim Kaput
BMC Bioinform.7
2008 Assessing batch effects of genotype calling algorithm BRLMM for the Affymetrix GeneChip Human Mapping 500 K array set using 270 HapMap samples
abstract
BACKGROUND: Genome-wide association studies (GWAS) aim to identify genetic variants (usually single nucleotide polymorphisms [SNPs]) across the entire human genome that are associated with phenotypic traits such as disease status and drug response. Highly accurate and reproducible genotype calling are paramount since errors introduced by calling algorithms can lead to inflation of false associations between genotype and phenotype. Most genotype calling algorithms currently used for GWAS are based on multiple arrays. Because hundreds of gigabytes (GB) of raw data are generated from a GWAS, the samples are typically partitioned into batches containing subsets of the entire dataset for genotype calling. High call rates and accuracies have been achieved. However, the effects of batch size (i.e., number of chips analyzed together) and of batch composition (i.e., the choice of chips in a batch) on call rate and accuracy as well as the propagation of the effects into significantly associated SNPs identified have not been investigated. In this paper, we analyzed both the batch size and batch composition for effects on the genotype calling algorithm BRLMM using raw data of 270 HapMap samples analyzed with the Affymetrix Human Mapping 500 K array set. RESULTS: Using data from 270 HapMap samples interrogated with the Affymetrix Human Mapping 500 K array set, three different batch sizes and three different batch compositions were used for genotyping using the BRLMM algorithm. Comparative analysis of the calling results and the corresponding lists of significant SNPs identified through association analysis revealed that both batch size and composition affected genotype calling results and significantly associated SNPs. Batch size and batch composition effects were more severe on samples and SNPs with lower call rates than ones with higher call rates, and on heterozygous genotype calls compared to homozygous genotype calls. CONCLUSION: Batch size and composition affect the genotype calling results in GWAS using BRLMM. The larger the differences in batch sizes, the larger the effect. The more homogenous the samples in the batches, the more consistent the genotype calls. The inconsistency propagates to the lists of significantly associated SNPs identified in downstream association analysis. Thus, uniform and large batch sizes should be used to make genotype calls for GWAS. In addition, samples of high homogeneity should be placed into the same batch.
Huixiao Hong, Zhenqiang Su, Weigong Ge, Leming Shi, Roger Perkins, Joshua Xu, James J. Chen, Tao Han 0007, Jim Kaput, James C. Fuscoe, Weida Tong
BMC Bioinform.12
2008 Very Important Pool (VIP) genes - an application for microarray-based molecular signatures
abstract
BACKGROUND: Advances in DNA microarray technology portend that molecular signatures from which microarray will eventually be used in clinical environments and personalized medicine. Derivation of biomarkers is a large step beyond hypothesis generation and imposes considerably more stringency for accuracy in identifying informative gene subsets to differentiate phenotypes. The inherent nature of microarray data, with fewer samples and replicates compared to the large number of genes, requires identifying informative genes prior to classifier construction. However, improving the ability to identify differentiating genes remains a challenge in bioinformatics. RESULTS: A new hybrid gene selection approach was investigated and tested with nine publicly available microarray datasets. The new method identifies a Very Important Pool (VIP) of genes from the broad patterns of gene expression data. The method uses a bagging sampling principle, where the re-sampled arrays are used to identify the most informative genes. Frequency of selection is used in a repetitive process to identify the VIP genes. The putative informative genes are selected using two methods, t-statistic and discriminatory analysis. In the t-statistic, the informative genes are identified based on p-values. In the discriminatory analysis, disjoint Principal Component Analyses (PCAs) are conducted for each class of samples, and genes with high discrimination power (DP) are identified. The VIP gene selection approach was compared with the p-value ranking approach. The genes identified by the VIP method but not by the p-value ranking approach are also related to the disease investigated. More importantly, these genes are part of the pathways derived from the common genes shared by both the VIP and p-ranking methods. Moreover, the binary classifiers built from these genes are statistically equivalent to those built from the top 50 p-value ranked genes in distinguishing different types of samples. CONCLUSION: The VIP gene selection approach could identify additional subsets of informative genes that would not always be selected by the p-value ranking method. These genes are likely to be additional true positives since they are a part of pathways identified by the p-value ranking method and expected to be related to the relevant biology. Therefore, these additional genes derived from the VIP method potentially provide valuable biological insights.
Zhenqiang Su, Huixiao Hong, Leming Shi, Roger Perkins, Weida Tong
BMC Bioinform.6
2006 Integrating time-course microarray gene expression profiles with cytotoxicity for identification of biomarkers in primary rat hepatocytes exposed to cadmium
abstract
MOTIVATION: DNA microarrays can provide information about the expression levels of thousands of genes simultaneously at the transcriptomic level, while conventional cell viability and cytotoxicity measurement methods provide information about the biological functions at the cellular level. Integrating these data at different levels provides a promising approach for evaluating or predicting how cells respond to chemical exposure. It is important to investigate the multi-scale biological system in a systematic way to better understand the gene regulation networks and signal transduction pathways involved in the cellular responses to environmental factors. RESULTS: Primary rat hepatocytes were exposed to cadmium acetate at 0, 1.25 and 2 microM. mRNA expression profiles at 0, 3, 6, 12 and 24 h were measured using the Affymetrix RatTox U34 GeneChip arrays. Simultaneously, cytotoxicity was assessed by lactase dehydrogenase leakage assay. Gene expression profiles at different time points were used to evaluate cytotoxicity at subsequent time points using partial least squares, and it was found that gene expression profiles at 0 h had the best prediction accuracy for the cytotoxicity observed at 12 h. Some biomarkers whose expression profiles showed strong relationship with cytotoxicity were identified and the underlying pathways were reconstructed to illustrate how hepatocytes respond to cadmium exposure. Permutation studies were also applied to assess the reliability of the predictive models. AVAILABILITY: Matlab source code is available upon request and DNA microarray data are available at GEO (http://www.ncbi.nlm.nih.gov/geo).
Yongxi Tan, Leming Shi, Saber M. Hussain, Weida Tong, John M. Frazier
Bioinform.5
2006 Differential gene expression in mouse primary hepatocytes exposed to the peroxisome proliferator-activated receptor alpha agonists
abstract
BACKGROUND: Fibrates are a unique hypolipidemic drugs that lower plasma triglyceride and cholesterol levels through their action as peroxisome proliferator-activated receptor alpha (PPARalpha) agonists. The activation of PPARalpha leads to a cascade of events that result in the pharmacological (hypolipidemic) and adverse (carcinogenic) effects in rodent liver. RESULTS: To understand the molecular mechanisms responsible for the pleiotropic effects of PPARalpha agonists, we treated mouse primary hepatocytes with three PPARalpha agonists (bezafibrate, fenofibrate, and WY-14,643) at multiple concentrations (0, 10, 30, and 100 microM) for 24 hours. When primary hepatocytes were exposed to these agents, transactivation of PPARalpha was elevated as measured by luciferase assay. Global gene expression profiles in response to PPARalpha agonists were obtained by microarray analysis. Among differentially expressed genes (DEGs), there were 4, 8, and 21 genes commonly regulated by bezafibrate, fenofibrate, and WY-14,643 treatments across 3 doses, respectively, in a dose-dependent manner. Treatments with 100 muM of bezafibrate, fenofibrate, and WY-14,643 resulted in 151, 149, and 145 genes altered, respectively. Among them, 121 genes were commonly regulated by at least two drugs. Many genes are involved in fatty acid metabolism including oxidative reaction. Some of the gene changes were associated with production of reactive oxygen species, cell proliferation of peroxisomes, and hepatic disorders. In addition, 11 genes related to the development of liver cancer were observed. CONCLUSION: Our results suggest that treatment of PPARalpha agonists results in the production of oxidative stress and increased peroxisome proliferation, thus providing a better understanding of mechanisms underlying PPARalpha agonist-induced hepatic disorders and hepatocarcinomas.
Lei Guo 0006, Jim Collins, Stacey L. Dial, Kshama Mehta, Ernice Blann, Leming Shi, Weida Tong, Yvonne P. Dragan
BMC Bioinform.10
2006 Microarray analysis distinguishes differential gene expression patterns from large and small colony Thymidine kinase mutants of L5178Y mouse lymphoma cells
abstract
BACKGROUND: The Thymidine kinase (Tk) mutants generated from the widely used L5178Y mouse lymphoma assay fall into two categories, small colony and large colony. Cells from the large colonies grow at a normal rate while cells from the small colonies grow slower than normal. The relative proportion of large and small colonies after mutagen treatment is associated with a mutagen's ability to induce point mutations and/or chromosomal mutations. The molecular distinction between large and small colony mutants, however, is not clear. RESULTS: To gain insights into the underlying mechanisms responsible for the mutant colony phenotype, microarray gene expression analysis was carried out on 4 small and 4 large colony Tk mutant samples. NCTR-fabricated long-oligonucleotide microarrays of 20,000 mouse genes were used in a two-color reference design experiment. The data were analyzed within ArrayTrack software that was developed at the NCTR. Principal component analysis and hierarchical clustering of the gene expression profiles showed that the samples were clearly separated into two groups based on their colony size phenotypes. The Welch T-test was used for determining significant changes in gene expression between the large and small colony groups and 90 genes whose expression was significantly altered were identified (p < 0.01; fold change > 1.5). Using Ingenuity Pathways Analysis (IPA), 50 out of the 90 significant genes were found in the IPA database and mapped to four networks associated with cell growth. Eleven percent of the 90 significant genes were located on chromosome 11 where the Tk gene resides while only 5.6% of the genes on the microarrays mapped to chromosome 11. All of the chromosome 11 significant genes were expressed at a higher level in the small colony mutants compared to the large colony mutants. Also, most of the significant genes located on chromosome 11 were disproportionally concentrated on the distal end of chromosome 11 where the Tk mutations occurred. CONCLUSION: The results indicate that microarray analysis can define cellular phenotypes and identify genes that are related to the colony size phenotypes. The findings suggest that genes in the DNA segment altered by the Tk mutations were significantly up-regulated in the small colony mutants, but not in the large colony mutants, leading to differential expression of a set of growth regulation genes that are related to cell apoptosis and other cellular functions related to the restriction of cell growth.
Tao Han 0007, Jianyong Wang 0001, Weida Tong, Martha M. Moore, James C. Fuscoe
BMC Bioinform.3
2006 GOFFA: Gene Ontology For Functional Analysis - A FDA Gene Ontology Tool for Analysis of Genomic and Proteomic Data
abstract
BACKGROUND: Gene Ontology (GO) characterizes and categorizes the functions of genes and their products according to biological processes, molecular functions and cellular components, facilitating interpretation of data from high-throughput genomics and proteomics technologies. The most effective use of GO information is achieved when its rich and hierarchical complexity is retained and the information is distilled to the biological functions that are most germane to the phenomenon being investigated. RESULTS: Here we present a FDA GO tool named Gene Ontology for Functional Analysis (GOFFA). GOFFA first ranks GO terms in the order of prevalence for a list of selected genes or proteins, and then it allows the user to interactively select GO terms according to their significance and specific biological complexity within the hierarchical structure. GOFFA provides five interactive functions (Tree view, Terms View, Genes View, GO Path and GO TreePrune) to analyze the GO data. Among the five functions, GO Path and GO TreePrune are unique. The GO Path simultaneously displays the ranks that order GOFFA Tree Paths based on statistical analysis. The GO TreePrune provides a visual display of a reduced GO term set based on a user's statistical cut-offs. Therefore, the GOFFA visual display can provide an intuitive depiction of the most likely relevant biological functions. CONCLUSION: With GOFFA, the user can dynamically interact with the GO data to interpret gene expression results in the context of biological plausibility, which can lead to new discoveries or identify new hypotheses. AVAILABILITY: GOFFA is available through ArrayTrack softwarehttp://edkb.fda.gov/webstart/arraytrack/.
Hongmei Sun, Roger Perkins, Weida Tong
BMC Bioinform.5
2005 Bioinformatics approaches for cross-species liver cancer analysis based on microarray gene expression profiling
abstract
BACKGROUND: The completion of the sequencing of human, mouse and rat genomes and knowledge of cross-species gene homologies enables studies of differential gene expression in animal models. These types of studies have the potential to greatly enhance our understanding of diseases such as liver cancer in humans. Genes co-expressed across multiple species are most likely to have conserved functions. We have used various bioinformatics approaches to examine microarray expression profiles from liver neoplasms that arise in albumin-SV40 transgenic rats to elucidate genes, chromosome aberrations and pathways that might be associated with human liver cancer. RESULTS: In this study, we first identified 2223 differentially expressed genes by comparing gene expression profiles for two control, two adenoma and two carcinoma samples using an F-test. These genes were subsequently mapped to the rat chromosomes using a novel visualization tool, the Chromosome Plot. Using the same plot, we further mapped the significant genes to orthologous chromosomal locations in human and mouse. Many genes expressed in rat 1q that are amplified in rat liver cancer map to the human chromosomes 10, 11 and 19 and to the mouse chromosomes 7, 17 and 19, which have been implicated in studies of human and mouse liver cancer. Using Comparative Genomics Microarray Analysis (CGMA), we identified regions of potential aberrations in human. Lastly, a pathway analysis was conducted to predict altered human pathways based on statistical analysis and extrapolation from the rat data. All of the identified pathways have been known to be important in the etiology of human liver cancer, including cell cycle control, cell growth and differentiation, apoptosis, transcriptional regulation, and protein metabolism. CONCLUSION: The study demonstrates that the hepatic gene expression profiles from the albumin-SV40 transgenic rat model revealed genes, pathways and chromosome alterations consistent with experimental and clinical research in human liver cancer. The bioinformatics tools presented in this paper are essential for cross species extrapolation and mapping of microarray data, its analysis and interpretation.
Weida Tong, Roger Perkins, Leming Shi, Huixiao Hong, X. Cao, Qian Xie 0004, S. H. Yim, J. M. Ward, Henry C. Pitot, Yvonne P. Dragan
BMC Bioinform.2
2005 Quality control and quality assessment of data from surface-enhanced laser desorption/ionization (SELDI) time-of flight (TOF) mass spectrometry (MS)
abstract
BACKGROUND: Proteomic profiling of complex biological mixtures by the ProteinChip technology of surface-enhanced laser desorption/ionization time-of-flight (SELDI-TOF) mass spectrometry (MS) is one of the most promising approaches in toxicological, biological, and clinic research. The reliable identification of protein expression patterns and associated protein biomarkers that differentiate disease from health or that distinguish different stages of a disease depends on developing methods for assessing the quality of SELDI-TOF mass spectra. The use of SELDI data for biomarker identification requires application of rigorous procedures to detect and discard low quality spectra prior to data analysis. RESULTS: The systematic variability from plates, chips, and spot positions in SELDI experiments was evaluated using biological and technical replicates. Systematic biases on plates, chips, and spots were not found. The reproducibility of SELDI experiments was demonstrated by examining the resulting low coefficient of variances of five peaks presented in all 144 spectra from quality control samples that were loaded randomly on different spots in the chips of six bioprocessor plates. We developed a method to detect and discard low quality spectra prior to proteomic profiling data analysis, which uses a correlation matrix to measure the similarities among SELDI mass spectra obtained from similar biological samples. Application of the correlation matrix to our SELDI data for liver cancer and liver toxicity study and myeloma-associated lytic bone disease study confirmed this approach as an efficient and reliable method for detecting low quality spectra. CONCLUSION: This report provides evidence that systematic variability between plates, chips, and spots on which the samples were assayed using SELDI based proteomic procedures did not exist. The reproducibility of experiments in our studies was demonstrated to be acceptable and the profiling data for subsequent data analysis are reliable. Correlation matrix was developed as a quality control tool to detect and discard low quality spectra prior to data analysis. It proved to be a reliable method to measure the similarities among SELDI mass spectra and can be used for quality control to decrease noise in proteomic profiling data prior to data analysis.
Huixiao Hong, Yvonne P. Dragan, Joshua Epstein, Candee Teitel, Bangzheng Chen, Qian Xie 0004, Leming Shi, Roger Perkins, Weida Tong
BMC Bioinform.10
2005 Cross-platform comparability of microarray technology: Intra-platform consistency and appropriate data analysis procedures are essential
abstract
BACKGROUND: The acceptance of microarray technology in regulatory decision-making is being challenged by the existence of various platforms and data analysis methods. A recent report (E. Marshall, Science, 306, 630-631, 2004), by extensively citing the study of Tan et al. (Nucleic Acids Res., 31, 5676-5684, 2003), portrays a disturbingly negative picture of the cross-platform comparability, and, hence, the reliability of microarray technology. RESULTS: We reanalyzed Tan's dataset and found that the intra-platform consistency was low, indicating a problem in experimental procedures from which the dataset was generated. Furthermore, by using three gene selection methods (i.e., p-value ranking, fold-change ranking, and Significance Analysis of Microarrays (SAM)) on the same dataset we found that p-value ranking (the method emphasized by Tan et al.) results in much lower cross-platform concordance compared to fold-change ranking or SAM. Therefore, the low cross-platform concordance reported in Tan's study appears to be mainly due to a combination of low intra-platform consistency and a poor choice of data analysis procedures, instead of inherent technical differences among different platforms, as suggested by Tan et al. and Marshall. CONCLUSION: Our results illustrate the importance of establishing calibrated RNA samples and reference datasets to objectively assess the performance of different microarray platforms and the proficiency of individual laboratories as well as the merits of various data analysis procedures. Thus, we are progressively coordinating the MAQC project, a community-wide effort for microarray quality control.
Leming Shi, Weida Tong, Uwe Scherf, Jing Han 0003, Raj K. Puri, Felix W. Frueh, Federico M. Goodsaid, Lei Guo 0006, Zhenqiang Su, Tao Han 0007, James C. Fuscoe, Z. Alex Xu, Tucker A. Patterson, Huixiao Hong, Qian Xie 0004, Roger Perkins, James J. Chen, Daniel A. Casciano
BMC Bioinform.2
2005 Microarray scanner calibration curves: characteristics and implications
abstract
BACKGROUND: Microarray-based measurement of mRNA abundance assumes a linear relationship between the fluorescence intensity and the dye concentration. In reality, however, the calibration curve can be nonlinear. RESULTS: By scanning a microarray scanner calibration slide containing known concentrations of fluorescent dyes under 18 PMT gains, we were able to evaluate the differences in calibration characteristics of Cy5 and Cy3. First, the calibration curve for the same dye under the same PMT gain is nonlinear at both the high and low intensity ends. Second, the degree of nonlinearity of the calibration curve depends on the PMT gain. Third, the two PMTs (for Cy5 and Cy3) behave differently even under the same gain. Fourth, the background intensity for the Cy3 channel is higher than that for the Cy5 channel. The impact of such characteristics on the accuracy and reproducibility of measured mRNA abundance and the calculated ratios was demonstrated. Combined with simulation results, we provided explanations to the existence of ratio underestimation, intensity-dependence of ratio bias, and anti-correlation of ratios in dye-swap replicates. We further demonstrated that although Lowess normalization effectively eliminates the intensity-dependence of ratio bias, the systematic deviation from true ratios largely remained. A method of calculating ratios based on concentrations estimated from the calibration curves was proposed for correcting ratio bias. CONCLUSION: It is preferable to scan microarray slides at fixed, optimal gain settings under which the linearity between concentration and intensity is maximized. Although normalization methods improve reproducibility of microarray measurements, they appear less effective in improving accuracy.
Leming Shi, Weida Tong, Zhenqiang Su, Tao Han 0007, Jing Han 0003, Raj K. Puri, Felix W. Frueh, Federico M. Goodsaid, Lei Guo 0006, William S. Branham, James J. Chen, Z. Alex Xu, Stephen C. Harris, Huixiao Hong, Qian Xie 0004, Roger Perkins, James C. Fuscoe
BMC Bioinform.2
2005 Decision Forest Analysis of 61 Single Nucleotide Polymorphisms in a Case-Control Study of Esophageal Cancer; a novel method
abstract
BACKGROUND: Systematic evaluation and study of single nucleotide polymorphisms (SNPs) made possible by high throughput genotyping technologies and bioinformatics promises to provide breakthroughs in the understanding of complex diseases. Understanding how the millions of SNPs in the human genome are involved in conferring susceptibility or resistance to disease, or in rendering a drug efficacious or toxic in the individual is a major goal of the relatively new fields of pharmacogenomics. Esophageal squamous cell carcinoma is a high-mortality cancer with complex etiology and progression involving both genetic and environmental factors. We examined the association between esophageal cancer risk and patterns of 61 SNPs in a case-control study for a population from Shanxi Province in North Central China that has among the highest rates of esophageal squamous cell carcinoma in the world. METHODS: High-throughput Masscode mass spectrometry genotyping was done on genomic DNA from 574 individuals (394 cases and 180 age-frequency matched controls). SNPs were chosen from among genes involving DNA repair enzymes, and Phase I and Phase II enzymes. We developed a novel adaptation of the Decision Forest pattern recognition method named Decision Forest for SNPs (DF-SNPs). The method was designated to analyze the SNP data. RESULTS: The classifier in separating the cases from the controls developed with DF-SNPs gave concordance, sensitivity and specificity, of 94.7%, 99.0% and 85.1%, respectively; suggesting its usefulness for hypothesizing what SNPs or combinations of SNPs could be involved in susceptibility to esophageal cancer. Importantly, the DF-SNPs algorithm incorporated a randomization test for assessing the relevance (or importance) of individual SNPs, SNP types (Homozygous common, heterozygous and homozygous variant) and patterns of SNP types (SNP patterns) that differentiate cases from controls. For example, we found that the different genotypes of SNP GADD45B E1122 are all associated with cancer risk. CONCLUSION: The DF-SNPs method can be used to differentiate esophageal squamous cell carcinoma cases from controls based on individual SNPs, SNP types and SNP patterns. The method could be useful to identify potential biomarkers from the SNP data and complement existing methods for genotype analyses.
Qian Xie 0004, Luke D. Ratnasinghe, Huixiao Hong, Roger Perkins, Ze-Zhong Tang, Philip R. Taylor, Weida Tong
BMC Bioinform.8