EDBT 2026 Demo / reviewers in the wild / expert
Bairong Shen
dblp:48/4172
· DBLP profile ↗
25ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0003-2899-1531ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 24 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards precision oncology: unsupervised manifold learning for spatial molecular profiling in cancer tissuesabstractPrecision oncology relies on the accurate characterization of spatial molecular distributions in cancer tissues to uncover critical biomarkers and guide clinical decision-making. However, the high dimensionality and complexity of mass spectrometry imaging (MSI) data pose significant challenges for effective analysis. This study presents an unsupervised manifold learning framework to address these challenges by mapping high-dimensional MSI data into a low-dimensional space while preserving essential molecular patterns. This method enables efficient dimensionality reduction, clustering, and visualization of MSI data, facilitating the discovery of spatially resolved molecular features. Applied to datasets from prostate cancer and colorectal adenocarcinoma, the proposed method accurately identifies cancerous regions and reveals highly correlated molecular markers with Pearson correlation coefficients up to 0.79. These findings demonstrate the potential of unsupervised manifold learning to enhance the interpretability and utility of MSI data in precision oncology, paving the way for improved biomarker discovery and cancer diagnostics. Guoqing Jiang, Jingming He, Xuemeng Fan, Xiaoya Gao, Cong Wu 0005, Bairong Shen |
BMC Bioinform. | 7 |
| 2026 | Bridging the gaps: Utilizing unlabeled face recognition datasets to boost semi-supervised facial expression recognition
Mengqiao He, Jinhua Feng, Bairong Shen |
Neurocomputing | 4 |
| 2026 | Biobanking for intelligent medicine: assessment and evaluation with the SHARE principleabstractBACKGROUND: Biobanks are essential for intelligent medicine but face fragmentation and heterogeneity. No standardized framework exists for assessing biobank data value using public information; this study addresses this gap. MATERIALS AND METHODS: We systematically evaluated 94 global biobanks (2010-2024) through a literature review and structured data extraction. Based on 12 international standards, we developed the 5-dimensional SHARE principle (Standardization, Hierarchical structuring, Analytical compatibility, Regulatory compliance, Evolutionary adaptability), operationalized into 10 indicators with a 4-tier scoring system. GPT-4o provided AI-supported prescoring, which was validated by 10 experts and through case studies, including RARPKB. RESULTS: The SHARE principle and a classification map of 94 biobanks were generated. AI and expert scoring showed substantial consistency (κ = 0.62; 95% CI, 0.54-0.70). Biobanks were categorized into 4 tiers: Traditional (60-69), Data-Driven (70-79), Knowledge-Guided (80-89), and Generative and Reasoning-oriented Biobank (90-100). Case validation confirmed utility for disease-specific biobanks. DISCUSSION: We highlight the principal findings, critically examine the reliance on public documentation, propose mitigation strategies, and discuss indicator weighting and implications for translational informatics. CONCLUSIONS: The SHARE principle provides a scalable, standardized method for assessing biobank data value, supporting biobank development, resource discovery, and the development of AI-driven biomedical ecosystems for intelligent medicine. Amin Ullah, Yingbo Zhang, Hui Zong, Xingyun Liu, Bairong Shen |
J. Am. Medical Informatics Assoc. | 9 |
| 2025 | Comprehensive human respiratory genome catalogue underlies the high resolution and precision of the respiratory microbiomeabstractThe human respiratory microbiome plays a crucial role in respiratory health, but there is no comprehensive respiratory genome catalogue (RGC) for studying the microbiome. In this study, we collected whole-metagenome shotgun sequencing data from 4067 samples and sequenced long reads of 124 samples, yielding 9.08 and 0.42 Tbp of short- and long-read data, respectively. By submitting these data with a novel assembly algorithm, we obtained a comprehensive human RGC. This high-quality RGC contains 190,443 contigs over 1 kbps and an N50 length exceeding 13 kbps; it comprises 159 high-quality and 393 medium-quality genomes, including 117 previously uncharacterized respiratory bacteria. Moreover, the RGC contains 209 respiratory-specific species not captured by the unified human gastrointestinal genome. Using the RGC, we revisited a study on a pediatric pneumonia dataset and identified 17 pneumonia-specific respiratory pathogens, reversing an inaccurate etiological conclusion due to the previous incomplete reference. Furthermore, we applied the RGC to the data of 62 participants with a clinical diagnosis of infection. Compared to the Nucleotide database, the RGC yielded greater specificity (0 versus 0.444, respectively) and sensitivity (0.852 versus 0.881, respectively), suggesting that the RGC provides superior sensitivity and specificity for the clinical diagnosis of respiratory diseases. Yinhu Li, Guangze Pan, Shuai Wang 0036, Zhengtu Li, Yiqi Jiang, Shuaicheng Li 0001, Bairong Shen |
Briefings Bioinform. | 9 |
| 2025 | Expertise or Hallucination? A Comprehensive Evaluation of ChatGPT's Aptitude in Clinical GeneticsabstractWhether viewed as an expert or as a source of ‘knowledge hallucination’, the use of ChatGPT in medical practice has stirred ongoing debate. This study sought to evaluate ChatGPT's capabilities in the field of clinical genetics, focusing on tasks such as ‘Clinical genetics exams’, ‘Associations between genetic diseases and pathogenic genes’, and ‘Limitations and trends in clinical genetics’. Results indicated that ChatGPT performed exceptionally well in question-answering tasks, particularly in clinical genetics exams and diagnosing single-gene diseases. It also effectively outlined the current limitations and prospective trends in clinical genetics. However, ChatGPT struggled to provide comprehensive answers regarding multi-gene or epigenetic diseases, particularly with respect to genetic variations or chromosomal abnormalities. In terms of systematic summarization and inference, some randomness was evident in ChatGPT's responses. In summary, while ChatGPT possesses a foundational understanding of general knowledge in clinical genetics due to hyperparameter learning, it encounters significant challenges when delving into specialized knowledge and navigating the complexities of clinical genetics, particularly in mitigating ‘Knowledge Hallucination’. To optimize its performance and depth of expertise in clinical genetics, integration with specialized knowledge databases and knowledge graphs is imperative. Yingbo Zhang, Shumin Ren, Chaoying Zhan, Mengqiao He, Xingyun Liu, Cong Wu 0005, Chuanzhu Fan, Bairong Shen |
IEEE Trans. Big Data | 11 |
| 2024 | Elevated incidence of somatic mutations at prevalent genetic sitesabstractThe common loci represent a distinct set of the human genome sites that harbor genetic variants found in at least 1% of the population. Small somatic mutations occur at the common loci and non-common loci, i.e. csmVariants and ncsmVariants, are presumed with similar probabilities. However, our work revealed that within the coding region, common loci constituted only 1.03% of all loci, yet they accounted for 5.14% of TCGA somatic mutations. Furthermore, the small somatic mutation incidence rate at these common loci was 2.7 times that observed in the non-common. Notably, the csmVariants exhibited an impressive recurrent rate of 36.14%, which was 2.59 times of the ncsmVariants. The C-to-T transition at the CpG sites accounted for 32.41% of the csmVariants, which was 2.93 times for the ncsmVariants. Interestingly, the aging-related mutational signature contributed to 13.87% of the csmVariants, 5.5 times that of ncsmVariants. Moreover, 35.93% of the csmVariants contexts exhibited palindromic features, outperforming ncsmVariant contexts by 1.84 times. Notably, cancer patients with higher csmVariants rates had better progression-free survival. Furthermore, cancer patients with high-frequency csmVariants enriched with mismatch repair deficiency were also associated with better progression-free survival. The accumulation of csmVariants during cancerogenesis is a complex process influenced by various factors. These include the presence of a substantial percentage of palindromic sequences at csmVariants sites, the impact of aging and DNA mismatch repair deficiency. Together, these factors contribute to the higher somatic mutation incidence rates of common loci and the overall accumulation of csmVariants in cancer development. Shuaicheng Li 0001, Bairong Shen |
Briefings Bioinform. | 3 |
| 2024 | scFed: federated learning for cell type classification with scRNA-seqabstractThe advent of single-cell RNA sequencing (scRNA-seq) has revolutionized our understanding of cellular heterogeneity and complexity in biological tissues. However, the nature of large, sparse scRNA-seq datasets and privacy regulations present challenges for efficient cell identification. Federated learning provides a solution, allowing efficient and private data use. Here, we introduce scFed, a unified federated learning framework that allows for benchmarking of four classification algorithms without violating data privacy, including single-cell-specific and general-purpose classifiers. We evaluated scFed using eight publicly available scRNA-seq datasets with diverse sizes, species and technologies, assessing its performance via intra-dataset and inter-dataset experimental setups. We find that scFed performs well on a variety of datasets with competitive accuracy to centralized models. Though Transformer-based model excels in centralized training, its performance slightly lags behind single-cell-specific model within the scFed framework, coupled with a notable time complexity concern. Our study not only helps select suitable cell identification methods but also highlights federated learning's potential for privacy-preserving, collaborative biomedical research. Shuang Wang 0002, Bochen Shen, Lanting Guo, Mengqi Shang, Jinze Liu, Bairong Shen |
Briefings Bioinform. | 7 |
| 2024 | Identify gestational diabetes mellitus by deep learning model from cell-free DNA at the early gestation stageabstractGestational diabetes mellitus (GDM) is a common complication of pregnancy, which has significant adverse effects on both the mother and fetus. The incidence of GDM is increasing globally, and early diagnosis is critical for timely treatment and reducing the risk of poor pregnancy outcomes. GDM is usually diagnosed and detected after 24 weeks of gestation, while complications due to GDM can occur much earlier. Copy number variations (CNVs) can be a possible biomarker for GDM diagnosis and screening in the early gestation stage. In this study, we proposed a machine-learning method to screen GDM in the early stage of gestation using cell-free DNA (cfDNA) sequencing data from maternal plasma. Five thousand and eighty-five patients from north regions of Mainland China, including 1942 GDM, were recruited. A non-overlapping sliding window method was applied for CNV coverage screening on low-coverage (~0.2×) sequencing data. The CNV coverage was fed to a convolutional neural network with attention architecture for the binary classification. The model achieved a classification accuracy of 88.14%, precision of 84.07%, recall of 93.04%, F1-score of 88.33% and AUC of 96.49%. The model identified 2190 genes associated with GDM, including DEFA1, DEFA3 and DEFB1. The enriched gene ontology (GO) terms and KEGG pathways showed that many identified genes are associated with diabetes-related pathways. Our study demonstrates the feasibility of using cfDNA sequencing data and machine-learning methods for early diagnosis of GDM, which may aid in early intervention and prevention of adverse pregnancy outcomes. Zicheng Zhao, Yousheng Yan, Wentao Yue, Ruixia Liu, Hailong Feng, Yujiao Chen, Bairong Shen, Lijian Zhao, Chenghong Yin |
Briefings Bioinform. | 16 |
| 2024 | PCAO2: an ontology for integration of prostate cancer associated genotypic, phenotypic and lifestyle dataabstractDisease ontologies facilitate the semantic organization and representation of domain-specific knowledge. In the case of prostate cancer (PCa), large volumes of research results and clinical data have been accumulated and needed to be standardized for sharing and translational researches. A formal representation of PCa-associated knowledge will be essential to the diverse data standardization, data sharing and the future knowledge graph extraction, deep phenotyping and explainable artificial intelligence developing. In this study, we constructed an updated PCa ontology (PCAO2) based on the ontology development life cycle. An online information retrieval system was designed to ensure the usability of the ontology. The PCAO2 with a subclass-based taxonomic hierarchy covers the major biomedical concepts for PCa-associated genotypic, phenotypic and lifestyle data. The current version of the PCAO2 contains 633 concepts organized under three biomedical viewpoints, namely, epidemiology, diagnosis and treatment. These concepts are enriched by the addition of definition, synonym, relationship and reference. For the precision diagnosis and treatment, the PCa-associated genes and lifestyles are integrated in the viewpoint of epidemiological aspects of PCa. PCAO2 provides a standardized and systematized semantic framework for studying large amounts of heterogeneous PCa data and knowledge, which can be further, edited and enriched by the scientific community. The PCAO2 is freely available at https://bioportal.bioontology.org/ontologies/PCAO, http://pcaontology.net/ and http://pcaontology.net/mobile/. Chunjiang Yu, Hui Zong, Yalan Chen, Yibin Zhou, Xingyun Liu, Xiaonan Zheng, Hua Min, Bairong Shen |
Briefings Bioinform. | 10 |
| 2024 | Advancing Chinese biomedical text mining with community challengesabstractOBJECTIVE: This study aims to review the recent advances in community challenges for biomedical text mining in China. METHODS: We collected information of evaluation tasks released in community challenges of biomedical text mining, including task description, dataset description, data source, task type and related links. A systematic summary and comparative analysis were conducted on various biomedical natural language processing tasks, such as named entity recognition, entity normalization, attribute extraction, relation extraction, event extraction, text classification, text similarity, knowledge graph construction, question answering, text generation, and large language model evaluation. RESULTS: We identified 39 evaluation tasks from 6 community challenges that spanned from 2017 to 2023. Our analysis revealed the diverse range of evaluation task types and data sources in biomedical text mining. We explored the potential clinical applications of these community challenge tasks from a translational biomedical informatics perspective. We compared with their English counterparts, and discussed the contributions, limitations, lessons and guidelines of these community challenges, while highlighting future directions in the era of large language models. CONCLUSION: Community challenge evaluation competitions have played a crucial role in promoting technology innovation and fostering interdisciplinary collaboration in the field of biomedical text mining. These challenges provide valuable platforms for researchers to develop state-of-the-art solutions. Hui Zong, Jiaxue Cha, Weizhe Feng, Erman Wu, Aibin Shao, Zuofeng Li, Buzhou Tang, Bairong Shen |
J. Biomed. Informatics | 11 |
| 2023 | Translational informatics for human microbiota: data resources, models and applicationsabstractWith the rapid development of human intestinal microbiology and diverse microbiome-related studies and investigations, a large amount of data have been generated and accumulated. Meanwhile, different computational and bioinformatics models have been developed for pattern recognition and knowledge discovery using these data. Given the heterogeneity of these resources and models, we aimed to provide a landscape of the data resources, a comparison of the computational models and a summary of the translational informatics applied to microbiota data. We first review the existing databases, knowledge bases, knowledge graphs and standardizations of microbiome data. Then, the high-throughput sequencing techniques for the microbiome and the informatics tools for their analyses are compared. Finally, translational informatics for the microbiome, including biomarker discovery, personalized treatment and smart healthcare for complex diseases, are discussed. Ahmad Ud Din, Baivab Sinha, Fuliang Qian, Bairong Shen |
Briefings Bioinform. | 6 |
| 2023 | Next-Generation Sequencing Markup Language (NGSML): A Medium for the Representation and Exchange of NGS DataabstractWith the increasing demand for low-cost high-throughput sequencing of large genomes, next-generation sequencing (NGS) technology has developed rapidly. NGS can not only be used in basic scientific research but also in clinical diagnostics and healthcare. Numerous software systems and tools have been developed to analyze NGS data, and various data formats have been produced to accommodate different sequencing equipment providers or analytical software. However, the data interoperability between these tools brings great challenges to researchers. A generic format that could be shared by most of the software and tools in the NGS field would make data interoperability and sharing easier. In this paper, we defined a general XML-based NGS markup language (NGSML) format for the representation and exchange of NGS data. We also developed a user-friendly GUI tool, NGSMLEditor, for presenting, creating, editing, and converting NGSML files. By using NGSML, various types of NGS data can be saved in one unified format. Compared with the unstructured plain text file, a structured data format based on XML technology solves the incompatibility of various NGS data formats. The NGSML specifications are freely available from http://www.sysbio.org.cn/NGSML. NGSMLEditor is open source under GNU GPL and can be downloaded from the website. Chunjiang Yu, Wenying Yan, Wentao Wu 0004, Bairong Shen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2022 | CRPMKB: a knowledge base of cancer risk prediction models for systematic comparison and personalized applicationsabstractMOTIVATION: In the era of big data and precision medicine, accurate risk assessment is a prerequisite for the implementation of risk screening and preventive treatment. A large number of studies have focused on the risk of cancer, and related risk prediction models have been constructed, but there is a lack of effective resource integration for systematic comparison and personalized applications. Therefore, the establishment and analysis of the cancer risk prediction model knowledge base (CRPMKB) is of great significance. RESULTS: The current knowledge base contains 802 model data. The model comparison indicates that the accuracy of cancer risk prediction was greatly affected by regional differences, cancer types and model types. We divided the model variables into four categories: environment, behavioral lifestyle, biological genetics and clinical examination, and found that there are differences in the distribution of various variables among different cancer types. Taking 50 genes involved in the lung cancer risk prediction models as an example to perform pathway enrichment analyses and the results showed that these genes were significantly enriched in p53 Signaling and Aryl Hydrocarbon Receptor Signaling pathways which are associated with cancer and specific diseases. In addition, we verified the biological significance of overlapping lung cancer genes via STRING database. CRPMKB was established to provide researchers an online tool for the future personalized model application and developing. This study of CRPMKB suggests that developing more targeted models based on specific demographic characteristics and cancer types will further improve the accuracy of cancer risk model predictions. AVAILABILITY AND IMPLEMENTATION: CRPMKB is freely available at http://www.sysbio.org.cn/CRPMKB/. The data underlying this article are available in the article and in its online supplementary material. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shumin Ren, Yanwen Jin, Yalan Chen, Bairong Shen |
Bioinform. | 4 |
| 2021 | HFBD: a biomarker knowledge database for heart failure heterogeneity and personalized applicationsabstractMOTIVATION: Heart failure (HF) is a cardiovascular disease with a high incidence around the world. Accumulating studies have focused on the identification of biomarkers for HF precision medicine. To understand the HF heterogeneity and provide biomarker information for the personalized diagnosis and treatment of HF, a knowledge database collecting the distributed and multiple-level biomarker information is necessary. RESULTS: In this study, the HF biomarker knowledge database (HFBD) was established by manually collecting the data and knowledge from literature in PubMed. HFBD contains 2618 records and 868 HF biomarkers (731 single and 137 combined) extracted from 1237 original articles. The biomarkers were classified into proteins, RNAs, DNAs and the others at molecular, image, cellular and physiological levels. The biomarkers were annotated with biological, clinical and article information as well as the experimental methods used for the biomarker discovery. With its user-friendly interface, this knowledge database provides a unique resource for the systematic understanding of HF heterogeneity and personalized diagnosis and treatment of HF in the era of precision medicine. AVAILABILITY AND IMPLEMENTATION: The platform is openly available at http://sysbio.org.cn/HFBD/. Hongxin He, Manhong Shi, Chaoying Zhan, Xingyun Liu, Shumin Ren, Bairong Shen |
Bioinform. | 9 |
| 2020 | Decoding competing endogenous RNA networks for cancer biomarker discoveryabstractCrosstalk between competing endogenous RNAs (ceRNAs) is mediated by shared microRNAs (miRNAs) and plays important roles both in normal physiology and tumorigenesis; thus, it is attractive for systems-level decoding of gene regulation. As ceRNA networks link the function of miRNAs with that of transcripts sharing the same miRNA response elements (MREs), e.g. pseudogenes, competing mRNAs, long non-coding RNAs, and circular RNAs, the perturbation of crucial interactions in ceRNA networks may contribute to carcinogenesis by affecting the balance of cellular regulatory system. Therefore, discovering biomarkers that indicate cancer initiation, development, and/or therapeutic responses via reconstructing and analyzing ceRNA networks is of clinical significance. In this review, the regulatory function of ceRNAs in cancer and crucial determinants of ceRNA crosstalk are firstly discussed to gain a global understanding of ceRNA-mediated carcinogenesis. Then, computational and experimental approaches for ceRNA network reconstruction and ceRNA validation, respectively, are described from a systems biology perspective. We focus on strategies for biomarker identification based on analyzing ceRNA networks and highlight the translational applications of ceRNA biomarkers for cancer management. This article will shed light on the significance of miRNA-mediated ceRNA interactions and provide important clues for discovering ceRNA network-based biomarker in cancer biology, thereby accelerating the pace of precision medicine and healthcare for cancer patients. Bairong Shen |
Briefings Bioinform. | 4 |
| 2020 | iODA: An integrated tool for analysis of cancer pathway consistency from heterogeneous multi-omics data
Chunjiang Yu, Bairong Shen |
J. Biomed. Informatics | 5 |
| 2019 | Computer-aided biomarker discovery for precision medicine: data resources, models and applicationsabstractBiomarkers are a class of measurable and evaluable indicators with the potential to predict disease initiation and progression. In contrast to disease-associated factors, biomarkers hold the promise to capture the changeable signatures of biological states. With methodological advances, computer-aided biomarker discovery has now become a burgeoning paradigm in the field of biomedical science. In recent years, the 'big data' term has accumulated for the systematical investigation of complex biological phenomena and promoted the flourishing of computational methods for systems-level biomarker screening. Compared with routine wet-lab experiments, bioinformatics approaches are more efficient to decode disease pathogenesis under a holistic framework, which is propitious to identify biomarkers ranging from single molecules to molecular networks for disease diagnosis, prognosis and therapy. In this review, the concept and characteristics of typical biomarker types, e.g. single molecular biomarkers, module/network biomarkers, cross-level biomarkers, etc., are explicated on the guidance of systems biology. Then, publicly available data resources together with some well-constructed biomarker databases and knowledge bases are introduced. Biomarker identification models using mathematical, network and machine learning theories are sequentially discussed. Based on network substructural and functional evidences, a novel bioinformatics model is particularly highlighted for microRNA biomarker discovery. This article aims to give deep insights into the advantages and challenges of current computational approaches for biomarker detection, and to light up the future wisdom toward precision medicine and nation-wide healthcare. Fuliang Qian, Bairong Shen |
Briefings Bioinform. | 6 |
| 2019 | Modeling and Simulation Studies of Complex Biological Systems for Precision Medicine and HealthcareabstractHere present three articles on the private preservation of genome or electronic health record (EHR) data. Three articles focus on the algorithm developing for the analysis of GWAS and EHR-based phenotyping as well as electrocardiogram (ECG)-based disease recognition and classification. Here also include an article developing an algorithm for repositioning of old drugs for their new applications. The articles in this special section proposed several computational model and simulation methods to address diverse medical and healthcare issues which will be helpful to the promotion of the cross-disciplinary researches on the translational medicine and healthcare. Bairong Shen, Xiaoqian Jiang, Xingming Zhao |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2018 | NGS-FC: A Next-Generation Sequencing Data Format ConverterabstractWith the widespread implementation of next-generation sequencing (NGS) technologies, millions of sequences have been produced. A lot of databases were created to store and organize the high-throughput sequencing data. Numerous analysis software programs and tools have been developed over the past years. Most of them use specific formats for data representation and storage. Data interoperability becomes a crucial challenge and many tools have been developed to convert NGS data from one format to another. However, most of them were developed for specific and limited formats. Here, we present NGS-FC (Next-Generation Sequencing Format Converter), which provides a framework to support the conversion between several formats. It supports 14 formats now and provides interfaces to enable users to improve the existing converters and add new ones. Moreover, NGS-FC achieved the overall competitive performance in comparison with some existing converters in terms of RAM usage and running time. The software is written in Java and can be executed standalone. The source code and documentation are freely available at http://sysbio.suda.edu.cn/NGS-FC. Chunjiang Yu, Wentao Wu 0004, Fei Zhu 0003, Bairong Shen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 8 |
| 2016 | PON-Sol: prediction of effects of amino acid substitutions on protein solubilityabstractMOTIVATION: Solubility is one of the fundamental protein properties. It is of great interest because of its relevance to protein expression. Reduced solubility and protein aggregation are also associated with many diseases. RESULTS: We collected from literature the largest experimentally verified solubility affecting amino acid substitution (AAS) dataset and used it to train a predictor called PON-Sol. The predictor can distinguish both solubility decreasing and increasing variants from those not affecting solubility. PON-Sol has normalized correct prediction ratio of 0.491 on cross-validation and 0.432 for independent test set. The performance of the method was compared both to solubility and aggregation predictors and found to be superior. PON-Sol can be used for the prediction of effects of disease-related substitutions, effects on heterologous recombinant protein expression and enhanced crystallizability. One application is to investigate effects of all possible AASs in a protein to aid protein engineering. AVAILABILITY AND IMPLEMENTATION: PON-Sol is freely available at http://structure.bmc.lu.se/PON-Sol The training and test data are available at http://structure.bmc.lu.se/VariBench/ponsol.php CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Abhishek Niroula, Bairong Shen, Mauno Vihinen |
Bioinform. | 3 |
| 2015 | Deciphering oncogenic drivers: from single genes to integrated pathwaysabstractTechnological advances in next-generation sequencing have uncovered a wide spectrum of aberrations in cancer genomes. The extreme diversity in cancer mutations necessitates computational approaches to differentiate between the 'drivers' with vital function in cancer progression and those nonfunctional 'passengers'. Although individual driver mutations are routinely identified, mutational profiles of different tumors are highly heterogeneous. There is growing consensus that pathways rather than single genes are the primary target of mutations. Here we review extant bioinformatics approaches to identifying oncogenic drivers at different mutational levels, highlighting the strategies for discovering driver pathways and networks from cancer mutation data. These approaches will help reduce the mutation complexity, thus providing a simplified picture of cancer. Maomin Sun, Bairong Shen |
Briefings Bioinform. | 3 |
| 2014 | Protein-protein interaction network constructing based on text mining and reinforcement learning with application to prostate cancerabstractAs a notoriously lethal human disease, cancer has obtained much concern for a long time. There have accumulated huge amounts of literature and experimental data on cancer-related research. It is impossible for people to deal with these texts manually to discover novel information and knowledge. However, text mining has an advantage of extracting previously unknown and understandable knowledge from large amounts of texts, and forming well-defined knowledge, providing the possibility to fully taking use of the existed texts. With the proceeding of biomedical research, people have gradually realized that complex biological functions and the phenomenon of life are the results of complex interactions among a variety of biological entities, such as protein. Deeply studying protein interaction network is essential to understand life. We, adopting reinforcement learning idea, put forward an algorithm for protein interaction network constructing. With the algorithm, nodes are used to represent proteins and edges denote interactions. During the evolutionary process, a node selects with which nodes in the network it tends to interact. Keep selecting and carrying on iteration, until eventually attaining an optimal network. The network is the result of the dynamic nature of learning behavior. As a malignancy, prostate cancer has been concerned for a long time. We attain biological texts from PubMed and establish a prostate cancer protein interaction networks by the proposed methods. The results show that our proposed method is pretty good. Network topology analysis results also show that the network node degree distribution is scale-free. Fei Zhu 0003, Quan Liu 0004, Bairong Shen |
BIBM | 4 |
| 2013 | Biomedical text mining and its applications in cancer research
Fei Zhu 0003, Preecha Patumcharoenpol, Jonathan H. Chan, Asawin Meechai, Wanwipa Vongsangnak, Bairong Shen |
J. Biomed. Informatics | 8 |
| 2008 | Towards patterns tree of gene coexpression in eukaryotic speciesabstractMOTIVATION: Cellular pathways behave coordinated regulation activity, and some reported works also have affirmed that genes in the same pathway have similar expression pattern. However, the complexity of biological systems regulation actually causes expression relationships between genes to display multiple patterns, such as linear, non-linear, local, global, linear with time-delayed, non-linear with time-delayed, monotonic and non-monotonic, which should be the explicit representation of cellular inner regulation mechanism in mRNA level. To investigate the relationship between different patterns, our work aims to systematically reveal gene-expression relationship patterns in cellular pathways and to check for the existence of dominating gene-expression pattern. By a large scale analysis of genes expression in three eukaryotic species, Saccharomyces cerevisiae, Caenorhabditis elegans and Human, we constructed gene coexpression patterns tree to systematically and hierarchically illustrate the different patterns and their interrelations. RESULTS: The results show that the linear is the dominating expression pattern in the same pathway. The time-shifted pattern is another important relationship pattern. Many genes from the different pathway also present coexpression patterns. The non-linear, non-monotonic and time-delayed relationship patterns reflect the remote interactions between the genes in cellular processes. Gene coexpression phenomena in the same pathways are diverse in different species. Genes in S.cerevisiae and C.elegans present strong coexpression relationships, especially in C.elegans, coexpression is more universal and stronger due to its special array of genes. However in Human, gene coexpression is not apparent and the human genome involves more complicated functional relationships. In conclusion, different patterns corresponding to different coordinating behaviors coexist. The patterns trees of different species give us comprehensive insight and understanding of genes expression activity in the cellular society. Xia Li 0004, Bairong Shen, Min Ding 0006, Ziyin Shen |
Bioinform. | 4 |
| 2003 | RankViaContact: ranking and visualization of amino acid contactsabstractSUMMARY: RankViaContact is a web service for calculation of residue-residue contact energies in proteins based on a coarse-grained model, and for visualization of interactions. The service provides information about ranked contact energies of residues, coordination numbers and the relative solvent accessibility of selected residues, as well as sequence and structure information. The program can be used to design stabilizing mutations, to analyze residue-residue contacts and to study the consequences of mutations. AVAILABILITY: http://bioinf.uta.fi/Rank.htm. Bairong Shen, Mauno Vihinen |
Bioinform. | 1 |