EDBT 2026 Demo / reviewers in the wild / expert
W. Jim Zheng
dblp:82/6434 · also Wenjin Jim Zheng
· DBLP profile ↗
27ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0001-7411-6047ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 26 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CeLLTra: aligning cell names with gene expression via a pathway-informed transformerabstractMOTIVATION: Single-cell RNA sequencing (scRNA-Seq) technology enables detailed exploration of gene expression at the individual cell level, crucial for annotating cell types and understanding cellular diversity. Traditional methods for cell type annotation often rely on marker genes and manual labeling, posing challenges due to low data quality and incomplete reference datasets. RESULTS: We developed CeLLTra, a novel contrastive learning framework that leverages a Transformer-based model integrating biological pathway information to group genes into super tokens, effectively capturing comprehensive gene expression from scRNA-Seq data. By combining this pathway-informed Transformer with a pretrained domain-specific language model, CeLLTra accurately aligns cell-type annotations with gene expression profiles. Evaluations on a large-scale human scRNA-Seq dataset showed that CeLLTra significantly outperformed state-of-the-art methods in supervised and zero-shot cell-type prediction. Additionally, CeLLTra generalized well to external datasets, improving clustering performance and enabling better characterization of cancerous cell states in tumor-infiltrating myeloid cells from non-small cell lung cancer patients. AVAILABILITY AND IMPLEMENTATION: CeLLTra is freely available on GitHub (https://github.com/WJZheng-group/CeLLTra) and Zenodo (https://doi.org/10.5281/zenodo.17666735). The datasets underlying this article are the following: GSE201333 and GSE127465. All these datasets are publicly available and can be freely accessed on the Gene Expression Omnibus repository. Zaiyi Zheng, Rongbin Li, Wen Chen 0001, Yuntao Yang, Meer A Ali, Jundong Li, W. Jim Zheng |
Bioinform. | 8 |
| 2024 | SAFER: sub-hypergraph attention-based neural network for predicting effective responses to dose combinationsabstractBACKGROUND: The potential benefits of drug combination synergy in cancer medicine are significant, yet the risks must be carefully managed due to the possibility of increased toxicity. Although artificial intelligence applications have demonstrated notable success in predicting drug combination synergy, several key challenges persist: (1) Existing models often predict average synergy values across a restricted range of testing dosages, neglecting crucial dose amounts and the mechanisms of action of the drugs involved. (2) Many graph-based models rely on static protein-protein interactions, failing to adapt to dynamic and higher-order relationships. These limitations constrain the applicability of current methods. RESULTS: We introduce SAFER, a Sub-hypergraph Attention-based graph model, addressing these issues by incorporating complex relationships among biological knowledge networks and considering dosing effects on subject-specific networks. SAFER outperformed previous models on the benchmark and the independent test set. The analysis of subgraph attention weight for the lung cancer cell line highlighted JAK-STAT signaling pathway, PRDM12, ZNF781, and CDC5L that have been implicated in lung fibrosis. CONCLUSIONS: SAFER presents an interpretable framework designed to identify drug-responsive signals. Tailored for comprehending dose effects on subject-specific molecular contexts, our model uniquely captures dose-level drug combination responses. This capability unlocks previously inaccessible avenues of investigation compared to earlier models. Furthermore, the SAFER framework can be leveraged by future inquiries to investigate molecular networks that uniquely characterize individual patients and can be applied to prioritize personalized effective treatment based on safe dose combinations. Yi-Ching Tang, Rongbin Li, Jing Tang 0002, W. Jim Zheng, Xiaoqian Jiang |
BMC Bioinform. | 4 |
| 2024 | Ensemble pretrained language models to extract biomedical knowledge from literatureabstractOBJECTIVES: The rapid expansion of biomedical literature necessitates automated techniques to discern relationships between biomedical concepts from extensive free text. Such techniques facilitate the development of detailed knowledge bases and highlight research deficiencies. The LitCoin Natural Language Processing (NLP) challenge, organized by the National Center for Advancing Translational Science, aims to evaluate such potential and provides a manually annotated corpus for methodology development and benchmarking. MATERIALS AND METHODS: For the named entity recognition (NER) task, we utilized ensemble learning to merge predictions from three domain-specific models, namely BioBERT, PubMedBERT, and BioM-ELECTRA, devised a rule-driven detection method for cell line and taxonomy names and annotated 70 more abstracts as additional corpus. We further finetuned the T0pp model, with 11 billion parameters, to boost the performance on relation extraction and leveraged entites' location information (eg, title, background) to enhance novelty prediction performance in relation extraction (RE). RESULTS: Our pioneering NLP system designed for this challenge secured first place in Phase I-NER and second place in Phase II-relation extraction and novelty prediction, outpacing over 200 teams. We tested OpenAI ChatGPT 3.5 and ChatGPT 4 in a Zero-Shot setting using the same test set, revealing that our finetuned model considerably surpasses these broad-spectrum large language models. DISCUSSION AND CONCLUSION: Our outcomes depict a robust NLP system excelling in NER and RE across various biomedical entities, emphasizing that task-specific models remain superior to generic large ones. Such insights are valuable for endeavors like knowledge graph development and hypothesis formulation in biomedical research. Qiang Wei 0002, Liang-Chin Huang, Jianfu Li, Yao-Shun Chuang, Jianping He 0002, Avisha Das, Vipina Kuttichi Keloth, Yuntao Yang, Chiamaka S. Diala, Kirk Roberts, Cui Tao, Xiaoqian Jiang, W. Jim Zheng, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 15 |
| 2024 | Developing deep learning-based strategies to predict the risk of hepatocellular carcinoma among patients with nonalcoholic fatty liver disease from electronic health recordsabstractOBJECTIVE: The accuracy of deep learning models for many disease prediction problems is affected by time-varying covariates, rare incidence, covariate imbalance and delayed diagnosis when using structured electronic health records data. The situation is further exasperated when predicting the risk of one disease on condition of another disease, such as the hepatocellular carcinoma risk among patients with nonalcoholic fatty liver disease due to slow, chronic progression, the scarce of data with both disease conditions and the sex bias of the diseases. The goal of this study is to investigate the extent to which the aforementioned issues influence deep learning performance, and then devised strategies to tackle these challenges. These strategies were applied to improve hepatocellular carcinoma risk prediction among patients with nonalcoholic fatty liver disease. METHODS: We evaluated two representative deep learning models in the task of predicting the occurrence of hepatocellular carcinoma in a cohort of patients with nonalcoholic fatty liver disease (n = 220,838) from a national EHR database. The disease prediction task was carefully formulated as a classification problem while taking censorship and the length of follow-up into consideration. RESULTS: We developed a novel backward masking scheme to deal with the issue of delayed diagnosis which is very common in EHR data analysis and evaluate how the length of longitudinal information after the index date affects disease prediction. We observed that modeling time-varying covariates improved the performance of the algorithms and transfer learning mitigated reduced performance caused by the lack of data. In addition, covariate imbalance, such as sex bias in data impaired performance. Deep learning models trained on one sex and evaluated in the other sex showed reduced performance, indicating the importance of assessing covariate imbalance while preparing data for model training. CONCLUSIONS: The strategies developed in this work can significantly improve the performance of hepatocellular carcinoma risk prediction among patients with nonalcoholic fatty liver disease. Furthermore, our novel strategies can be generalized to apply to other disease risk predictions using structured electronic health records, especially for disease risks on condition of another disease. Yujia Zhou 0003, Ruoxing Li, Kenneth D. Chavin, Hua Xu 0001, Liang Li 0026, David J. H. Shih, W. Jim Zheng |
J. Biomed. Informatics | 9 |
| 2022 | An evidence-based lexical pattern approach for quality assurance of Gene Ontology relationsabstractGene Ontology (GO) is widely used in the biological domain. It is the most comprehensive ontology providing formal representation of gene functions (GO concepts) and relations between them. However, unintentional quality defects (e.g. missing or erroneous relations) in GO may exist due to the large size of GO concepts and complexity of GO structures. Such quality defects would impact the results of GO-based analyses and applications. In this work, we introduce a novel evidence-based lexical pattern approach for quality assurance of GO relations. We leverage two layers of evidence to suggest potentially missing relations in GO as follows. We first utilize related concept pairs (i.e. existing relations) in GO to extract relationship-specific lexical patterns, which serve as the first layer evidence to automatically suggest potentially missing relations between unrelated concept pairs. For each suggested missing relation, we further identify two other existing relations as the second layer of evidence that resemble the difference between the missing relation and the existing relation based on which the missing relation is suggested. Applied to the 15 December 2021 release of GO, this approach suggested a total of 866 potentially missing relations. Local domain experts evaluated the entire set of potentially missing relations, and identified 821 as missing relations and 45 indicate erroneous existing relations. We submitted these findings to the GO consortium for further validation and received encouraging feedback. These indicate that our evidence-based approach can be utilized to uncover missing relations and erroneous existing relations in GO. Rashmie Abeysinghe, Yuntao Yang, Mason Bartels, W. Jim Zheng, Licong Cui |
Briefings Bioinform. | 4 |
| 2022 | CNGPLD: case-control copy-number analysis using Gaussian process latent differenceabstractMOTIVATION: Cross-sectional analyses of primary cancer genomes have identified regions of recurrent somatic copy-number alteration, many of which result from positive selection during cancer formation and contain driver genes. However, no effective approach exists for identifying genomic loci under significantly different degrees of selection in cancers of different subtypes, anatomic sites or disease stages. RESULTS: CNGPLD is a new tool for performing case-control somatic copy-number analysis that facilitates the discovery of differentially amplified or deleted copy-number aberrations in a case group of cancer compared with a control group of cancer. This tool uses a Gaussian process statistical framework in order to account for the covariance structure of copy-number data along genomic coordinates and to control the false discovery rate at the region level. AVAILABILITY AND IMPLEMENTATION: CNGPLD is freely available at https://bitbucket.org/djhshih/cngpld as an R package. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. David J. H. Shih, Ruoxing Li, W. Jim Zheng, Kim-Anh Do, Shiaw-Yih Lin, Scott L. Carter |
Bioinform. | 4 |
| 2021 | Anticancer drug synergy prediction in understudied tissues using transfer learningabstractOBJECTIVE: Drug combination screening has advantages in identifying cancer treatment options with higher efficacy without degradation in terms of safety. A key challenge is that the accumulated number of observations in in-vitro drug responses varies greatly among different cancer types, where some tissues are more understudied than the others. Thus, we aim to develop a drug synergy prediction model for understudied tissues as a way of overcoming data scarcity problems. MATERIALS AND METHODS: We collected a comprehensive set of genetic, molecular, phenotypic features for cancer cell lines. We developed a drug synergy prediction model based on multitask deep neural networks to integrate multimodal input and multiple output. We also utilized transfer learning from data-rich tissues to data-poor tissues. RESULTS: We showed improved accuracy in predicting synergy in both data-rich tissues and understudied tissues. In data-rich tissue, the prediction model accuracy was 0.9577 AUROC for binarized classification task and 174.3 mean squared error for regression task. We observed that an adequate transfer learning strategy significantly increases accuracy in the understudied tissues. CONCLUSIONS: Our synergy prediction model can be used to rank synergistic drug combinations in understudied tissues and thus help to prioritize future in-vitro experiments. Code is available at https://github.com/yejinjkim/synergy-transfer. Yejin Kim 0001, Jing Tang 0002, W. Jim Zheng, Xiaoqian Jiang |
J. Am. Medical Informatics Assoc. | 4 |
| 2021 | Deep representation learning of patient data from Electronic Health Records (EHR): A systematic review
Yuqi Si, Jingcheng Du, Xiaoqian Jiang, Timothy A. Miller, Fei Wang 0001, W. Jim Zheng, Kirk Roberts |
J. Biomed. Informatics | 7 |
| 2020 | Deep Representation Learning of Patient Data from Electronic Health Records: A Systematic Review
Yuqi Si, Jingcheng Du, Xiaoqian Jiang, Timothy A. Miller, Fei Wang 0001, W. Jim Zheng, Kirk Roberts |
AMIA | 7 |
| 2020 | A transformation-based method for auditing the IS-A hierarchy of biomedical terminologies in the Unified Medical Language SystemabstractOBJECTIVE: The Unified Medical Language System (UMLS) integrates various source terminologies to support interoperability between biomedical information systems. In this article, we introduce a novel transformation-based auditing method that leverages the UMLS knowledge to systematically identify missing hierarchical IS-A relations in the source terminologies. MATERIALS AND METHODS: Given a concept name in the UMLS, we first identify its base and secondary noun chunks. For each identified noun chunk, we generate replacement candidates that are more general than the noun chunk. Then, we replace the noun chunks with their replacement candidates to generate new potential concept names that may serve as supertypes of the original concept. If a newly generated name is an existing concept name in the same source terminology with the original concept, then a potentially missing IS-A relation between the original and the new concept is identified. RESULTS: Applying our transformation-based method to English-language concept names in the UMLS (2019AB release), a total of 39 359 potentially missing IS-A relations were detected in 13 source terminologies. Domain experts evaluated a random sample of 200 potentially missing IS-A relations identified in the SNOMED CT (U.S. edition) and 100 in Gene Ontology. A total of 173 of 200 and 63 of 100 potentially missing IS-A relations were confirmed by domain experts, indicating that our method achieved a precision of 86.5% and 63% for the SNOMED CT and Gene Ontology, respectively. CONCLUSIONS: Our results showed that our transformation-based method is effective in identifying missing IS-A relations in the UMLS source terminologies. Fengbo Zheng, Jay Shi, Yuntao Yang, W. Jim Zheng, Licong Cui |
J. Am. Medical Informatics Assoc. | 4 |
| 2020 | A content-based literature recommendation system for datasets to improve data reusability - A case study on Gene Expression Omnibus (GEO) datasets
Braja Gopal Patra, Vahed Maroufy, Babak Soltanalizadeh, W. Jim Zheng, Kirk Roberts, Hulin Wu |
J. Biomed. Informatics | 5 |
| 2018 | Predict effective drug combination by deep belief network and ontology fingerprints
Guocai Chen, Alex Tsoi, Hua Xu 0001, W. Jim Zheng |
J. Biomed. Informatics | 4 |
| 2018 | A study of generalizability of recurrent neural network-based predictive models for heart failure onset risk using a large and heterogeneous EHR data set
Laila Rasmy, Yonghui Wu 0001, Ningtao Wang, W. Jim Zheng, Fei Wang 0001, Hulin Wu, Hua Xu 0001, Degui Zhi |
J. Biomed. Informatics | 5 |
| 2017 | HiCComp: Multiple-level comparative analysis of Hi-C data by triplet networkabstractHi-C technique is an important tool for the study of 3D genome organization. In the past few years, we have seen an explosion of Hi-C data in a variety of cell/tissue types. While these publicly available data presents an unprecedented opportunity to interrogate chromosomal architecture, how to quantitatively compare Hi-C data from different tissues and identify tissue-specific chromatin interactions remains challenging. Here, we present HiCComp, a comprehensive framework for comparing Hi-C data. HiCComp utilizes convolutional neural networks to extract key features in Hi-C interaction matrices in a fully automatic way. The core component of HiCComp is a triplet network, which contains three identical convolutional neural networks with shared parameters. The inputs to our network are three Hi-C matrices: two of them are biological replicates from the same cell type and the third one is from another cell type. The HiCComp network takes advantages of the two biological replicates to estimate the natural variation in the experiments and further use it to identify significant variations between Hi-C matrices from different cell types. Furthermore, we incorporate systematic occluding method into our framework so that we can identify the dynamic interaction regions from Hi-C maps. Finally, we show that the dynamic regions between two cell types are enriched for transcription factor binding sites and histone modifications that are associated with cis-regulatory functions, suggesting these variations in 3D genome structure are potentially gene regulatory events. W. Jim Zheng, Jijun Tang |
BIBM | 3 |
| 2017 | The International Conference on Intelligent Biology and Medicine (ICIBM) 2016: from big data to big analytical toolsabstractThe 2016 International Conference on Intelligent Biology and Medicine (ICIBM 2016) was held on December 8-10, 2016 in Houston, Texas, USA. ICIBM included eight scientific sessions, four tutorials, one poster session, four highlighted talks and four keynotes that covered topics on 3D genomics structural analysis, next generation sequencing (NGS) analysis, computational drug discovery, medical informatics, cancer genomics, and systems biology. Here, we present a summary of the nine research articles selected from ICIBM 2016 program for publishing in BMC Bioinformatics. Zhandong Liu, W. Jim Zheng, Genevera I. Allen, Jianhua Ruan, Zhongming Zhao |
BMC Bioinform. | 2 |
| 2015 | A comparative study of disease genes and drug targets in the human protein interactomeabstractBACKGROUND: Disease genes cause or contribute genetically to the development of the most complex diseases. Drugs are the major approaches to treat the complex disease through interacting with their targets. Thus, drug targets are critical for treatment efficacy. However, the interrelationship between the disease genes and drug targets is not clear. RESULTS: In this study, we comprehensively compared the network properties of disease genes and drug targets for five major disease categories (cancer, cardiovascular disease, immune system disease, metabolic disease, and nervous system disease). We first collected disease genes from genome-wide association studies (GWAS) for five disease categories and collected their corresponding drugs based on drugs' Anatomical Therapeutic Chemical (ATC) classification. Then, we obtained the drug targets for these five different disease categories. We found that, though the intersections between disease genes and drug targets were small, disease genes were significantly enriched in targets compared to their enrichment in human protein-coding genes. We further compared network properties of the proteins encoded by disease genes and drug targets in human protein-protein interaction networks (interactome). The results showed that the drug targets tended to have higher degree, higher betweenness, and lower clustering coefficient in cancer Furthermore, we observed a clear fraction increase of disease proteins or drug targets in the near neighborhood compared with the randomized genes. CONCLUSIONS: The study presents the first comprehensive comparison of the disease genes and drug targets in the context of interactome. The results provide some foundational network characteristics for further designing computational strategies to predict novel drug targets and drug repurposing. Jingchun Sun, Kevin W. Zhu, W. Jim Zheng, Hua Xu 0001 |
BMC Bioinform. | 3 |
| 2014 | An Integrative Framework for Drug Target Prediction and Repurposing
Jingchun Sun, Cui Tao, Kevin W. Zhu, W. Jim Zheng, Hua Xu 0001 |
AMIA | 4 |
| 2013 | Exploring genomes with a game engineabstractStudying genomes continues to be beneficial for evolutionary discover, and prognosticating/diagnosing many genetic disorders and diseases. Most of the these studies have used systems that view the DNA in a linear structure, but having this information is only a small part of fully understanding what they can reveal. Visualizing genomes in real time 3D can give researchers more insight, but this is fraught with hardware limitations. Each element contains vast amounts of information that cannot be processed at once. However, by using a game engine and sophisticated video game visualization techniques, we were able to construct a multi-platform real-time 3D genome viewer. Jeremiah J. Shepherd, Lingxi Zhou, W. Jim Zheng, Jijun Tang |
BIBM | 4 |
| 2013 | Exploring genomes with a game engine
Jeremiah J. Shepherd, Bill Arndt, W. Jim Zheng, Jijun Tang |
FDG | 3 |
| 2011 | Consistent Differential Expression Pattern (CDEP) on microarray to identify genes related to metastatic behaviorabstractBACKGROUND: To utilize the large volume of gene expression information generated from different microarray experiments, several meta-analysis techniques have been developed. Despite these efforts, there remain significant challenges to effectively increasing the statistical power and decreasing the Type I error rate while pooling the heterogeneous datasets from public resources. The objective of this study is to develop a novel meta-analysis approach, Consistent Differential Expression Pattern (CDEP), to identify genes with common differential expression patterns across different datasets. RESULTS: We combined False Discovery Rate (FDR) estimation and the non-parametric RankProd approach to estimate the Type I error rate in each microarray dataset of the meta-analysis. These Type I error rates from all datasets were then used to identify genes with common differential expression patterns. Our simulation study showed that CDEP achieved higher statistical power and maintained low Type I error rate when compared with two recently proposed meta-analysis approaches. We applied CDEP to analyze microarray data from different laboratories that compared transcription profiles between metastatic and primary cancer of different types. Many genes identified as differentially expressed consistently across different cancer types are in pathways related to metastatic behavior, such as ECM-receptor interaction, focal adhesion, and blood vessel development. We also identified novel genes such as AMIGO2, Gem, and CXCL11 that have not been shown to associate with, but may play roles in, metastasis. CONCLUSIONS: CDEP is a flexible approach that borrows information from each dataset in a meta-analysis in order to identify genes being differentially expressed consistently. We have shown that CDEP can gain higher statistical power than other existing approaches under a variety of settings considered in the simulation study, suggesting its robustness and insensitivity to data variation commonly associated with microarray experiments. AVAILABILITY: CDEP is implemented in R and freely available at: http://genomebioinfo.musc.edu/CDEP/. CONTACT: [email protected]. Lam C. Tsoi, Tingting Qin, Elizabeth H. Slate, W. Jim Zheng |
BMC Bioinform. | 4 |
| 2010 | Genome3D: A viewer-model framework for integrating and visualizing multi-scale epigenomic information within a three-dimensional genomeabstractBACKGROUND: New technologies are enabling the measurement of many types of genomic and epigenomic information at scales ranging from the atomic to nuclear. Much of this new data is increasingly structural in nature, and is often difficult to coordinate with other data sets. There is a legitimate need for integrating and visualizing these disparate data sets to reveal structural relationships not apparent when looking at these data in isolation. RESULTS: We have applied object-oriented technology to develop a downloadable visualization tool, Genome3D, for integrating and displaying epigenomic data within a prescribed three-dimensional physical model of the human genome. In order to integrate and visualize large volume of data, novel statistical and mathematical approaches have been developed to reduce the size of the data. To our knowledge, this is the first such tool developed that can visualize human genome in three-dimension. We describe here the major features of Genome3D and discuss our multi-scale data framework using a representative basic physical model. We then demonstrate many of the issues and benefits of multi-resolution data integration. CONCLUSIONS: Genome3D is a software visualization tool that explores a wide range of structural genomic and epigenetic data. Data from various sources of differing scales can be integrated within a hierarchical framework that is easily adapted to new developments concerning the structure of the physical genome. In addition, our tool has a simple annotation mechanism to incorporate non-structural information. Genome3D is unique is its ability to manipulate large amounts of multi-resolution data from diverse sources to uncover complex and new structural relationships within the genome. Thomas M. Asbury, Matt Mitman, Jijun Tang, W. Jim Zheng |
BMC Bioinform. | 4 |
| 2009 | Evaluation of genome-wide association study results through development of ontology fingerprintsabstractMOTIVATION: Genome-wide association (GWA) studies may identify multiple variants that are associated with a disease or trait. To narrow down candidates for further validation, quantitatively assessing how identified genes relate to a phenotype of interest is important. RESULTS: We describe an approach to characterize genes or biological concepts (phenotypes, pathways, diseases, etc.) by ontology fingerprint--the set of Gene Ontology (GO) terms that are overrepresented among the PubMed abstracts discussing the gene or biological concept together with the enrichment p-value of these terms generated from a hypergeometric enrichment test. We then quantify the relevance of genes to the trait from a GWA study by calculating similarity scores between their ontology fingerprints using enrichment p-values. We validate this approach by correctly identifying corresponding genes for biological pathways with a 90% average area under the ROC curve (AUC). We applied this approach to rank genes identified through a GWA study that are associated with the lipid concentrations in plasma as well as to prioritize genes within linkage disequilibrium (LD) block. We found that the genes with highest scores were: ABCA1, lipoprotein lipase (LPL) and cholesterol ester transfer protein, plasma for high-density lipoprotein; low-density lipoprotein receptor, APOE and APOB for low-density lipoprotein; and LPL, APOA1 and APOB for triglyceride. In addition, we identified genes relevant to lipid metabolism from the literature even in cases where such knowledge was not reflected in current annotation of these genes. These results demonstrate that ontology fingerprints can be used effectively to prioritize genes from GWA studies for experimental validation. Lam C. Tsoi, Michael Boehnke, Richard L. Klein, W. Jim Zheng |
Bioinform. | 4 |
| 2009 | Text-mining approach to evaluate terms for ontology development
Lam C. Tsoi, Wenle Zhao, W. Jim Zheng |
J. Biomed. Informatics | 4 |
| 2006 | Combining comparative genomics with de novo motif discovery to identify human transcription factor DNA-binding motifsabstractBACKGROUND: As more and more genomes are sequenced, comparative genomics approaches provide a methodology for identifying conserved regulatory elements that may be involved in gene regulations. RESULTS: We developed a novel method to combine comparative genomics with de novo motif discovery to identify human transcription factor binding motifs that are overrepresented and conserved in the upstream regions of a set of co-regulated genes. The method is validated by analyzing a well-characterized muscle specific gene set, and the results showed that our approach performed better than the existing programs in terms of sensitivity and prediction rate. CONCLUSION: The newly developed method can be used to extract regulatory signals in co-regulated genes, which can be derived from the microarray clustering analysis. Linyong Mao, W. Jim Zheng |
BMC Bioinform. | 2 |
| 2005 | Capturing biological information with class?Cresponsibility?Ccollaboration cardsabstractUNLABELLED: Class-responsibility-collaboration (CRC) cards have been used extensively in the software industry for defining complex object-oriented software requirements. We have adapted this tool to capture information about biological components, collaborators and responsibilities within these collaborations, which is not captured by current annotation tools. CRC cards should provide a common ground that will facilitate communication between biologist and computer scientists. AVAILABILITY: A CRC card template, XML representation and XML schema are freely available at http://people.musc.edu/~zhengw/CRCCard/CRC_Card_Index.html SUPPLEMENTARY INFORMATION: Supplemental Figures 1-4. Daniel Shegogue, W. Jim Zheng |
Bioinform. | 2 |
| 2005 | Object-oriented biological system integration: a SARS coronavirus exampleabstractMOTIVATION: The importance of studying biology at the system level has been well recognized, yet there is no well-defined process or consistent methodology to integrate and represent biological information at this level. To overcome this hurdle, a blending of disciplines such as computer science and biology is necessary. RESULTS: By applying an adapted, sequential software engineering process, a complex biological system (severe acquired respiratory syndrome-coronavirus viral infection) has been reverse-engineered and represented as an object-oriented software system. The scalability of this object-oriented software engineering approach indicates that we can apply this technology for the integration of large complex biological systems. AVAILABILITY: A navigable web-based version of the system is freely available at http://people.musc.edu/~zhengw/SARS/Software-Process.htm Daniel Shegogue, W. Jim Zheng |
Bioinform. | 2 |
| 2005 | Integration of the Gene Ontology into an object-oriented architectureabstractBACKGROUND: To standardize gene product descriptions, a formal vocabulary defined as the Gene Ontology (GO) has been developed. GO terms have been categorized into biological processes, molecular functions, and cellular components. However, there is no single representation that integrates all the terms into one cohesive model. Furthermore, GO definitions have little information explaining the underlying architecture that forms these terms, such as the dynamic and static events occurring in a process. In contrast, object-oriented models have been developed to show dynamic and static events. A portion of the TGF-beta signaling pathway, which is involved in numerous cellular events including cancer, differentiation and development, was used to demonstrate the feasibility of integrating the Gene Ontology into an object-oriented model. RESULTS: Using object-oriented models we have captured the static and dynamic events that occur during a representative GO process, "transforming growth factor-beta (TGF-beta) receptor complex assembly" (GO:0007181). CONCLUSION: We demonstrate that the utility of GO terms can be enhanced by object-oriented technology, and that the GO terms can be integrated into an object-oriented model by serving as a basis for the generation of object functions and attributes. Daniel Shegogue, W. Jim Zheng |
BMC Bioinform. | 2 |