EDBT 2026 Demo / reviewers in the wild / expert
Yongsheng Bai
dblp:56/2260
· DBLP profile ↗
22ranked-venue papers
4as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorComputer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Data-Centric Multimodal Pavement Distress Detection Using Intensity-Range ImageryabstractAutomated pavement distress detection across large-scale roadway networks generates massive volumes of multimodal image data that must be efficiently indexed, searched, and analyzed within infrastructure monitoring and multimedia retrieval pipelines. Leveraging paired intensity–range imagery is attractive because intensity encodes appearance and texture, while range captures complementary three-dimensional geometry for robust 2D–3D distress understanding. However, modality utility is highly class-dependent, and naive early fusion can entangle modalities and suppress class-critical evidence, degrading detection and downstream retrieval performance. We study nine asphalt concrete pavement (ACP) distress types in the TxDOT 2D/3D pavement dataset, observe heterogeneous modality preferences across categories, and propose a decoupled multimodal YOLO variant with a full-backbone split design that keeps intensity and range feature streams separate in the backbone and fuses them only in the neck. Building on this principle, we further develop a class- and modality-aware online augmentation strategy: geometric transforms are synchronized across channels to preserve 2D/3D alignment, while pixel-level perturbations are applied per modality, with strengths selected based on each class’s modality dominance. On the ACP benchmark, the split backbone improves validation AP50 from 0.766 to 0.812, and the proposed augmentation further boosts it to 0.869. On the held-out test set, our system achieves 0.833 AP50 and 0.428 AP50–95, outperforming early fusion and uniform two-channel augmentation baselines under the same sensing configuration. Overall, our results indicate that late-fusion split learning with modality-aligned augmentation is a simple, data-centric design pattern for robust multimodal detection on large-scale intensity–range imagery. Wenhan Tao, Yongsheng Bai, Jelena Tesic |
ICMR | 2 |
| 2025 | A Multi-Omics Bioinformatics Pipeline Framework for Discovering Conserved MiRNA Regulatory Networks Across Human CancersabstractMutations within miRNA binding regions can lead to dysregulation, causing the uncontrolled expression of disease risk genes and tumor formation. Since wet-lab clinical trials are both time and cost-intensive, we developed an integrative bioinformatics framework and subsequent pipeline to identify significant miRNA–gene interactions common across multiple cancers. Risk gene datasets for eight cancers (BLCA, BRCA, COAD, KICH, LIHC, LUAD, LUSC, and PRAD) were obtained through The Cancer Genome Atlas (TCGA) database, along with complementary miRNAs using TarBase. Functional enrichment was performed through the Database for Annotation, Visualization, and Integrated Discovery (DAVID). Overlapping terms and shared genes were detected, and common genes were functionally enriched through MIENTURNET KEGG, Reactome, and Disease Ontology pathways. Finally, a custom pipeline was built using Python, incorporating TarBase, DIANA-microT, and the UCSC Genome Browser. Cross-cancer comparison revealed six biological terms and fourteen multi-cancer gene intersections, while functional analysis revealed converging pathways between cancer pairs. Key genes, including TP53 and ZNF3, were widespread across several of the eight cancers and functional categories, indicating universal roles in transcriptional regulation and genomic stability. Pipeline development allowed for the automated extraction of genomic coordinates, nucleotide sequences, and amino acid annotations, eliminating the redundancy of manual data compiling. Alex J. Mi, Joyce Tang, Amy Xu, Kyle Yang, Yongsheng Bai |
BIBM | 5 |
| 2024 | Bioinformatics Analysis Reveals Candidate Biomarkers Served As Dendritic Cell-based Lung Squamous CellCarcinoma (LUSC) TreatmentabstractLung squamous cell carcinoma (LUSC) causes high mortality rates worldwide. The goal of the study is to use bioinformatics approaches to identify candidate genes associated with lung cancer and the ones that could serve as targets for specific immune infiltration.The candidate genes in two significant clusters reported for LUSC based on TCGA database were chosen for analysis. The pattern of gene expression in both clusters shows opposite to their genomic landscape signature: one cluster of genes tends to have longer average gene length and fewer average number of exons than the genes in the other cluster. Functional annotation using The Database for Annotation, Visualization, and Integrated Discovery (DAVID) was performed after genomic information analysis. Groups of genes for their enriched terms having the False Discovery Rate (FDR) value less than 0.1 were further studied. Cancer association for selected genes were confirmed using the Cancer Genetics Web tool. Three genes (RNASE6, CCR1, CMKLR1) from one cluster and 11 genes (DLL4, NID2, ROBO4, COL4A1, VASH1, VWF, BGN, KIT, PRND, OSBPL10, CD34) from the other cluster are prioritized to be top candidate genes.Correlation between gene expression and immune cell infiltration was examined using TIMER2.0. Positive correlations were observed between the expressions of these genes and six different types of immune infiltration, particularly with dendritic cells, according to TIMER2.0 Gene analysis. Mutation analysis revealed varied frequencies, with ROBO4 showing the highest mutation rate (6.6%) among our identified LUSC-associated genes. Five genes (DLL4, CMKLR1, CD34, ROBO4, and NID2) exhibit a strong positive correlation between their expressions with immune infiltration. This suggests that these genes may serve as potential targets for dendritic cell immune infiltration in the treatment of LUSC. The candidate genes prioritized in our study have medical significance. Yongsheng Bai |
BIBM | 2 |
| 2024 | Lightweight Deep Learning for Missing Data Imputation in Wastewater Treatment With Variational Residual Auto-EncoderabstractDeep variational residual auto-encoder (ResNet-VAE) has shown promising outcomes in missing imputation of wastewater quality data. Nevertheless, with its large storage size and computation overhead, it is of great difficulty to deploy wastewater treatment plant (WWTP) sensors for real-time missing imputation. To address this problem, we propose a novel approach called lightweight-ResNet-VAE to compress classical ResNet-VAE by network pruning, weight quantization, and relative indexing in this article. First, we develop a three-step network pruning method to sparsify the weight matrices by removing insignificant weights to reduce the time cost of model inference. Second, we develop weight quantization and use eight shared weights to compress the size of each weight from 32-bit to 3-bit. Finally, the relative indexing is adopted to further compress the size of the classical ResNet-VAE by compressed sparse row (CSR), which greatly accelerates the model computation and saves storage size. Experiments on the Beipai IoT influent quality data set demonstrate that lightweight-ResNet-VAE compresses the size of the classical ResNet-VAE from 301.88 to 26.24 kB with a compression rate of 11.50 times, and outperforms the baseline methods in terms of computation acceleration, storage saving and energy consumption with only a slight decrease on accuracy as 2.74% in MAPE of missing imputation for wastewater quality data due to pruning less significant weights and quantizing the remaining weights. Wen Zhang 0001, Rui Li 0108, Pei Quan, Yongsheng Bai, Bojun Su |
IEEE Internet Things J. | 5 |
| 2023 | Characterization of oncogenes and tumor suppressor genes targeted by onco-microRNAs and tumor suppressor microRNAs in kidney cancerabstractMutated oncogenes and tumor suppressor genes can significantly alter cell functions during the development of cancer. A thorough analysis of their mutation patterns may offer some insights into cancer emergence as well as some suggestions for potential drug targeting. Kidney renal clear cell carcinoma (KIRC) is the most common subtype of kidney cancer. KIRC has a poor prognosis and a high mortality rate. To date, immunotherapy based on immune checkpoints is the most promising treatment. MicroRNAs (miRNAs) down-regulate target genes by binding to the 3’ UTR of the respective targets. MiRNAs target genes at the post-transcriptional level after mRNA has been transcribed. The complexity of the miRNA-gene targeting interaction, as a result, has made this research more challenging. There are two classes of genes that this study focuses on which are oncogenes and tumor suppressor genes. Oncogenes function by deregulating cell proliferation and suppressing apoptosis. Tumor suppressor genes function by inhibiting cell proliferation and tumor development. There have been many studies that focus on the targeting relationship between genes and miRNAs, but few studies investigate the big picture of the relationship between oncogenes and tumor suppressor genes and how it can be regulated at the genome-wide level in the context of the miRNA targeting mechanism. We downloaded the significant miRNA-gene targeting cluster files for KIRC from a published study. We then separated and annotated the tumor suppressors and oncogenes based on their survival analysis. We grouped the data into four different categories based on the miRNA-targeted gene relationship (TS-TS, Onco-Onco, TS-Onco, Onco-TS). Based on the survival significance (p-value < .05) analysis, we have identified 28 oncogenes and 188 tumor suppressor genes. We have also identified 8 tumor suppressor miRNAs and 62 onco-miRNAs. Claire Shen, Yongsheng Bai |
BIBM | 2 |
| 2023 | Identification of Key Biomarkers Associated with Ductal Breast Cancer in Spatial Transcriptomics DataabstractSpatial transcriptomics (ST) is a collection of groundbreaking genomic technologies that allows the measurement of gene expression with spatial localization information on tissues. Recent developments in ST analysis have allowed a deep investigation of breast cancer environments and the cellular composition of such malignancies. Despite these advancements, a comprehensive spatial exploration of breast cancer-related genes and their involvement in oncogenic signaling pathways has not yet been fully realized. This limitation is partly due to a lack of refined methods capable of effectively extracting differential gene expression information from multiple ST datasets. In this study, we integrated three 10x genomics breast cancer ST datasets that offer high-resolution gene expression insights. We identified nine potential breast cancer marker genes that are differentially expressed in tumor regions and performed various downstream analyses such as survival analysis, protein-protein interactions, and conserved domain analysis to articulate the neoplastic tendencies of our identified genes. In summary, our analysis would provide guidance for researchers to better comprehend the biological functions of the detected candidate genes and investigate potential therapeutic targets. Ellie Xi, Lulu Shang, Tutu Hu, Chloe Yu, Yongsheng Bai |
BIBM | 6 |
| 2022 | A Bioinformatics Pipeline for the identification of disease-causing variants in humans that can change protein structureabstractOne active field of research in bioinformatics is how protein structures are altered by missense variants. The introduction of AlphaFold2 markedly facilitates the prediction of protein structure, however, whether the protein structure changes due to missense variants contribute to the clinically observable phenotype is still not well studied. In this study, we analyzed a dataset of pathogenic and benign variants from the ClinVar database, which allowed us to study their impact on protein structure alterations through accessing the phenotype possibly derived from these genetic variants. We reported that some pathogenic variants, such as APOE variants, are mostly clustering in the high pLDDT regions, have higher RMSD scores, and have gains and losses of the local amino acid interaction. This suggests that the pathogenic effect is likely caused by a disrupted protein structure. Our findings may help us better understand certain diseases and facilitate drug discovery. Mingjia Ma, Ian Hou, Jeslyn Gao, Yongsheng Bai, Xiaoming Liu 0021 |
BIBM | 5 |
| 2022 | Loci2Tissue: Ranking tissues by the e3xpression of disease-associated genes reveals insights of the underlying mechanisms of complex diseases and traitsabstractModern high throughput technologies routinely produce a large set of genomic loci of biological interest like genome-wide association studies (GWASs). Annotating the set of genomic loci may lead to new biological insights. However, available tools are limited.In this study, we developed a new bioinformatics software package named loci2tissue that aims to connect a set of input genomic loci to tissues and organs, thus providing annotation in terms of tissue-specific transcription regulation potential. This is achieved by utilizing multi-tissue expression quantitative trait loci (eQTLs) information provided by Genotype-Tissue Expression (GTEx) to connect genomic loci to genes in a tissue-specific manner. We then rank tissues based on the tissue-specific expression levels of these genes. When applying loci2tissue to sets of loci that harbor genetic variants linked to complex diseases, we are able to identify specific tissues involved in complex diseases or traits.Our analyses revealed interesting tissue-trait pairs. As examples, we found significant enrichment of visceral omental adipose tissue in Alzheimer’s disease and the hippocampus tissue in Parkinson’s disease. These results shed light on the underlying biology of many complex diseases and traits where the tissue is likely to be the source of pathogenesis. Boqi Wang, Daniel Lu, Catherine Zhang 0002, Nicole Xu, Steven Qiu, Yongsheng Bai, Brian Hu 0003, Zhaohui Qin |
BIBM | 6 |
| 2022 | Prioritizing Intellectual Disability Candidate Genes and Understanding Family Diseases Using Machine LearningabstractWith thousands of developmental disorders, new sequencing technologies such as whole exome sequencing (WES) have been developed to identify causal genes and variants within patients. Unfortunately, analysis of WES remains challenging as an individual’s exome spans over 30,000 complex variants. Such variants are often implicated in X-linked diseases that lead to developmental disorders (DD). In support of using WES to analyze DD-associated diseases, a bioinformatics software called Exomiser can be implemented. Exomiser is equipped with multiple filtering, prioritization, and ranking algorithms, making it the ideal software for scientists to identify disease-causing variants. We applied Exomiser to analyze all 17 family members of the CEPH family 1463, a reference panel that serves as the CEU (Central Europe) population of the HapMap project, making it the ideal basis for family-disease research. Through Exomiser’s combined score assessment of 0.95+, We identified 12 critical DD-associated disease genes. We used DAVID, Orphanet, and AMELIE to verify the candidate genes and elucidated common biological annotations, ICD-10 codes, and performed causality comparisons. Yongsheng Bai |
BIBM | 2 |
| 2021 | Bioinformatics analysis of miRNAs identifies enrichment of axon guidance pathway genes in ovarian cancer stem cellsabstractOvarian cancer ranks fifth in cancer deaths among women, accounting for more deaths than any other cancer of the female reproductive system. Cancer stem cells (CSCs) are a small population of cancer cells that are believed to be the reason for cancer initiation and tumor relapse. Recently, the important role of MicroRNAs (miRNAs) in maintaining and regulating CSCs through targeting multiple oncogenic signaling pathways has been reported. Previously, we have found differentially expressed miRNAs in ovarian CSCs compared with bulk cancer cells in SKOV3 and Kuramochi cell lines. Here using online miRNA target prediction tools combined with a web interface tool MMiRNA-Tar, we identified the downstream targets of these differentially expressed miRNAs. After filtering, we found 271 novel targets of downregulated miRNAs and 442 novel targets of upregulated miRNAs, which may potentially control cancer cell stemness. Among this gene list, we further validated the potentially important role of the axon guidance pathway genes by using multiple bioinformatic analyses. In summary, our study identified that dysregulated miRNA can facilitate maintenance of CSCs by regulating the axon guidance pathway genes, which will ultimately help us find novel cancer therapeutic targets. Shurui Cai, Renata Fu, Ellie Xi, Daniel Lin, Yongsheng Bai, Qi-En Wang |
BIBM | 8 |
| 2021 | A Deep Learning Model for Ancestry Estimation with Craniometric MeasurementsabstractAncestry estimation from human skeletal remains is a significant component in forensic anthropology studies. Although some tools have been developed, their performances are not very good in practice. In this paper, we evaluated the utility of the deep learning method in cranial ancestry estimation based on Howells craniometric data. Specifically, 2,524 cranial individuals from the Howells main datasets and 468 from the Howells test datasets were analyzed in the paper. The individuals of 82 craniometric measurements in the Howells datasets were clustered into six ancestry groups: African, Austro-Melanesian, Polynesian-Micronesia, East Asian, Native American, and European. After the data engineering process, the Howells datasets with 30 craniometric measurements were fed into a feedforward neural network (FNN) model. The FNN model was trained on 80% of the Howells main datasets, validated on 10% of the Howells main datasets, and tested on another 10% of the Howells main datasets. The model's prediction accuracy for all six ancestry groups reached 80.6%. Compared with popular ancestry estimation programs AncesTrees and Fordisc 3.1 using the Howells test dataset, the performance of the FNN deep learning method is on par overall and better for some ancestry groups. Yibo Dong 0003, Andrew Gao, Ian Hou, Kevin Ma, Ruoxian Huang, Yongsheng Bai, Xiaoming Liu 0021 |
BIBM | 6 |
| 2021 | Computational Prediction of Biological Signatures for Candidate Driver Genes Associated with Ovarian CancerabstractOvarian cancer is a prominent cause of cancer deaths among women, accounting for the most deaths of any cancer in the female reproductive system. The functions of a gene in ovarian tissue are affected by its expression level. Oncogenes and tumor suppressor genes are often differentially expressed between ovarian cancer and normal tissues. We selected 50 genes with the highest expression values based on TCGA ovarian cancer RNA-seq data as the input genes of this study. We then conducted protein-protein interaction and functional annotation analysis. The 43 interacted genes were enriched with 11 GO and 13 KEGG Pathway terms. The 22 out of 43 genes that are not highly expressed in the normal ovarian tissue, as well as 12 other genes from the list of input genes with ovarian cancer association reported from published literature, were selected for further investigation. The Multiple Sequence Comparison by Log-Expectation (MUSCLE) tool was run on those 34 genes, and all genes were associated with conserved domains. Finally, survival analysis was performed on 34 candidate genes using the Kaplan-Meier Plotter to evaluate the impact they have on ovarian cancer progression. We reported 21 genes (C3, CLU, COLIAI, COLlA2, COL3A1, CTCFL, FN1, HSPAS, IGF2, PSAP, SPARC, TUBB, ALDOA, GAPDH, HSP90AB1, PKM, RPLP0, RPLPI, RPS4, RPS6, UBC) that had domain conservation and significant p-values for survival analysis. Daniel Lin, Renata Fu, Ellie Xi, Yongsheng Bai |
BIBM | 4 |
| 2020 | Computational identification of key pathways and differentially-expressed gene signatures in ovarian cancer stem cellsabstractCancer stem cells (CSCs) are a tumor cell type capable of self-renewal and differentiation and believed to be responsible for metastasis, therapy resistance, and tumor relapse. Therefore, targeting CSCs could be a promising therapy strategy to improve the outcome of cancer patients, and this requires accurate identification of CSCs from the tumors and better characterize this cell population in terms of stemness maintenance. In the present study, we sought to establish a pathway and gene signature in epithelial ovarian CSCs. Using three publicly available datasets from the Gene expression omnibus (GEO), we identified several common, potentially regulatory pathways and differentially-expressed genes that could be used to characterize and identify ovarian CSCs. Several established signaling pathways, such as the MAPK pathway and the Wnt pathway, were found to be enriched in ovarian CSCs. In addition, we found that the axon guidance pathway is also enriched in ovarian CSCs, suggesting a novel regulatory mechanism in the maintenance of this cell population. We also identified four genes, ITGBl, TUBB2B, CXCL2, and SORBS2, that are differentially expressed in ovarian CSCs across all three GEO datasets, indicating that this gene expression signature might be used to identify CSCs in epithelial ovarian cancers. Renata Fu, Yongsheng Bai, Qi-En Wang |
BIBM | 2 |
| 2020 | Systematic and Comprehensive Survey of Genomic Loci Associated with Complex Diseases and TraitsabstractOver the past decade, Genome-wide association studies (GWAS) have been successfully developed and applied to identify many sequence variants that are significantly associated with common diseases and traits. Hundreds of thousands of such trait-associated variants have already been cataloged, making GWASs great resources for complex disease research. The main challenge ahead lies in elucidating disease mechanisms from these findings. In this study, we conducted a systematic survey of the relationships between genes which have been linked to complex diseases and previously researched genetic associations to explore whether any enrichment of known biological knowledge exists. To achieve this, we queried these genes against various collections of previously researched genomic associations through Enrichr. We uncovered many intriguing enrichment patterns between disease-associated genes and biomedical conditions such as the association between Obesity & Hot Drink Temperature, Alzheimer's and Nose Size, and Body Mass Index & the PCDHA10 Pathway. These findings could potentially shed light on the pathogenicity of these complex diseases and may help in the future development of novel therapeutic treatments. Isabella He, Tianli Jiang, Aryan Thakur, Yongsheng Bai, Zhaohui Qin |
BIBM | 4 |
| 2020 | Investigating Genetic Signatures for Sex-Biased miRNA targeted Genes related to Intellectual DisabilityabstractIntellectual Disability (ID) often co-occurs with ASD and other diseases, so investigation of Non-syndromic ID (NS-ID) likely provides more accurate insights into genetic causes. As ~80% of NS-ID genes reside on the X-chromosome, contributing to the high male-female ratio in ID patients, we investigated the gender bias in ID by obtaining targeted genes of 13 X-linked and sex-biased miRNAs associated with NS-ID and studying the genetic signatures of 73 brain-expressed (i.e. overexpressed in brain tissues) sex-biased miRNA-targeted genes. Genetic signatures, including Gene Ontology information, disease association, exon numbers, and Protein-Protein Interaction, all of which presented differences between male-biased and female-biased genes, suggest genes' pathogeny by reflecting possible mutations during alternative splicing and pathogenic disruptions of interactions/pathways. Specifically, our result suggests that while the proportion of brainexpressed male-biased genes targeted by X-linked miRNAs is higher, X-linked ID-related genes are more prevalent in females. We also found that sex-biased ID-related genes tend to have more exon numbers, and interactions among miRNA-targeted female-biased genes are significantly higher. Moreover, previously discovered prevalent symptoms in ID patients, such as osteogenesis imperfecta and osteoporosis, are explained by ID-related genes' signatures. Multiple associations between our prioritized genes and ID-related pathways, symptoms, and syndromes further indicate the role miRNA has in targeting plays in the context of ID. Junmeng Yang, Susan Huang, Yongsheng Bai |
BIBM | 4 |
| 2020 | End-to-end Deep Learning Methods for Automated Damage Detection in Extreme Events at Various ScalesabstractRobust Mask R-CNN (Mask Regional Convolutional Neural Network) methods are proposed and tested for automatic detection of cracks on structures or their components that may be damaged during extreme events, such as earthquakes. We curated a new dataset with 2,021 labeled images for training and validation and aimed to find end-to-end deep neural networks for crack detection in the field. With data augmentation and parameters fine-tuning, Path Aggregation Network (PANet) with spatial attention mechanisms and High- resolution Network (HRNet) are introduced into Mask R-CNNs. The tests on three public datasets with low- or high-resolution images demonstrate that the proposed methods can achieve a big improvement over alternative networks, so the proposed method may be sufficient for crack detection for a variety of scales in real applications. Yongsheng Bai, Halil Sezen, Alper Yilmaz 0001 |
ICPR | 1 |
| 2020 | MMiRNA-Viewer2, a bioinformatics tool for visualizing functional annotation for MiRNA and MRNA pairs in a networkabstractBACKGROUND: Although there are many studies on the characteristics of miRNA-mRNA interactions using miRNA and mRNA sequencing data, the complexity of the change of the correlation coefficients and expression values of the miRNA-mRNA pairs between tumor and normal samples is still not resolved, and this hinders the potential clinical applications. There is an urgent need to develop innovative methodologies and tools that can characterize and visualize functional consequences of cancer risk gene and miRNA pairs while analyzing the tumor and normal samples simultaneously. RESULTS: web server integrates and displays the mRNA and miRNA gene annotation information, signaling cascade pathways and direct cancer association between miRNAs and mRNAs. Functional annotation and gene regulatory information can be directly retrieved from our web server, which can help users quickly identify significant interaction sub-network and report possible disease or cancer association. The tool can identify pivotal miRNAs or mRNAs that contribute to the complexity of cancer, while engaging modern next-generation sequencing technology to analyze the tumor and normal samples concurrently. We compared our tools with other visualization tools. CONCLUSION: serves as a multitasking platform in which users can identify significant interaction clusters and retrieve functional and cancer-associated information for miRNA-mRNA pairs between tumor and normal samples. Our tool is applicable across a range of diseases and cancers and has advantages over existing tools. Yongsheng Bai, Steve Baker, Kevin Exoo, Xingqin Dai, Lizhong Ding 0002, Naureen Aslam Khattak, Hannah Liu, Xiaoming Liu 0021 |
BMC Bioinform. | 1 |
| 2020 | Proceedings of the 2019 MidSouth Computational Biology and Bioinformatics Society (MCBIOS) ConferenceabstractThe 16th Annual MidSouth Computational Biology and Bioinformatics Society (MCBIOS XVI) conference was held in Hilton Birmingham at University of Alabama Birmingham (UAB) conference center on March 28–30, 2019. The theme of the conference was Informatics for Precision Medicine. The co-chairs and conference hosts were Drs. Jake Y. Chen and Matthew Might from the University of Alabama Birmingham. Jonathan D. Wren, Yongsheng Bai, Zhaohui S. Qin, Da Yan 0001, Ramin Homayouni |
BMC Bioinform. | 2 |
| 2017 | Structure Modeling and Molecular Docking Studies of Schizophrenia Candidate Genes, Synapsins 2 (SYN2) and Trace Amino Acid Receptor (TAAR6)
Naureen Aslam Khattak, Sheikh Arslan Sehgal, Yongsheng Bai, Youping Deng |
ISBRA | 3 |
| 2017 | Identification of genome-wide non-canonical spliced regions and analysis of biological functions for spliced sequences using Read-Split-FlyabstractBACKGROUND: It is generally thought that most canonical or non-canonical splicing events involving U2- and U12 spliceosomes occur within nuclear pre-mRNAs. However, the question of whether at least some U12-type splicing occurs in the cytoplasm is still unclear. In recent years next-generation sequencing technologies have revolutionized the field. The "Read-Split-Walk" (RSW) and "Read-Split-Run" (RSR) methods were developed to identify genome-wide non-canonical spliced regions including special events occurring in cytoplasm. As the significant amount of genome/transcriptome data such as, Encyclopedia of DNA Elements (ENCODE) project, have been generated, we have advanced a newer more memory-efficient version of the algorithm, "Read-Split-Fly" (RSF), which can detect non-canonical spliced regions with higher sensitivity and improved speed. The RSF algorithm also outputs the spliced sequences for further downstream biological function analysis. RESULTS: We used open access ENCODE project RNA-Seq data to search spliced intron sequences against the U12-type spliced intron sequence database to examine whether some events could occur as potential signatures of U12-type splicing. The check was performed by searching spliced sequences against 5'ss and 3'ss sequences from the well-known orthologous U12-type spliceosomal intron database U12DB. Preliminary results of searching 70 ENCODE samples indicated that the presence of 5'ss with U12-type signature is more frequent than U2-type and prevalent in non-canonical junctions reported by RSF. The selected spliced sequences have also been further studied using miRBase to elucidate their functionality. Preliminary results from 70 samples of ENCODE datasets show that several miRNAs are prevalent in studied ENCODE samples. Two of these are associated with many diseases as suggested in the literature. Specifically, hsa-miR-1273 and hsa-miR-548 are associated with many diseases and cancers. CONCLUSIONS: Our RSF pipeline is able to detect many possible junctions (especially those with a high RPKM) with very high overall accuracy and relative high accuracy for novel junctions. We have incorporated useful parameter features into the pipeline such as, handling variable-length read data, and searching spliced sequences for splicing signatures and miRNA events. We suggest RSF, a tool for identifying novel splicing events, is applicable to study a range of diseases across biological systems under different experimental conditions. Yongsheng Bai, Jeff Kinne, Lizhong Ding 0002, Ethan Rath, Aaron Cox, Siva Dharman Naidu |
BMC Bioinform. | 1 |
| 2017 | Identification of streptococcal small RNAs that are putative targets of RNase III through bioinformatics analysis of RNA sequencing dataabstractBACKGROUND: Small noncoding regulatory RNAs (sRNAs) are post-transcriptional regulators, regulating mRNAs, proteins, and DNA in bacteria. One class of sRNAs, trans-acting sRNAs, are the most abundant sRNAs transcribed from the intergenic regions (IGRs) of the bacterial genome. In Streptococcus pyogenes, a common and potentially deadly pathogen, many sRNAs have been identified, but only a few have been studied. The goal of this study is to identify trans-acting sRNAs that can be substrates of RNase III. The endoribonuclease RNase III cleaves double stranded RNAs, which can be formed during the interaction between an sRNA and target mRNAs. RESULTS: For this study, we created an RNase III null mutant of Streptococcus pyogenes and its RNA sequencing (RNA-Seq) data were analyzed and compared to that of the wild-type. First, we developed a custom script that can detect intergenic regions of the S. pyogenes genome. A differential expression analysis with Cufflinks and Stringtie was then performed to identify the intergenic regions whose expression was influenced by the RNase III gene deletion. CONCLUSION: This analysis yielded 12 differentially expressed regions with >|2| fold change and p ≤ 0.05. Using Artemis and Bamview genome viewers, these regions were visually verified leaving 6 putative sRNAs. This study not only expanded our knowledge on novel sRNAs but would also give us new insight into sRNA degradation. Ethan Rath, Stephanie Pitman, Kyu Hong Cho, Yongsheng Bai |
BMC Bioinform. | 4 |
| 2016 | Dissecting the biological relationship between TCGA miRNA and mRNA sequencing data using MMiRNA-ViewerabstractBACKGROUND: MicroRNAs (miRNA) are short nucleotides that interact with their target genes through 3' untranslated regions (UTRs). The Cancer Genome Atlas (TCGA) harbors an increasing amount of cancer genome data for both tumor and normal samples. However, there are few visualization tools focusing on concurrently displaying important relationships and attributes between miRNAs and mRNAs of both cancer tumor and normal samples. Moreover, a deep investigation of miRNA-mRNA target and biological relationships across multiple cancer types by integrating web-based analysis has not been thoroughly conducted. RESULTS: We developed an interactive visualization tool called MMiRNA-Viewer that can concurrently present the co-relationships of expression between miRNA-mRNA pairs of both tumor and normal samples into a single graph. The input file of MMiRNA-Viewer contains the expression information including fold changes between normal and tumor samples for mRNAs and miRNAs, the correlation between mRNA and miRNA, and the predicted target relationship by a number of databases. Users can also load their own input data into MMiRNA-Viewer and visualize and compare detailed information about cancer-related gene expression changes, and also changes in the expression of transcription-regulating miRNAs. To validate the MMiRNA-Viewer, eight types of TCGA cancer datasets with both normal and control samples were selected in this study and three filter steps were applied subsequently. We performed Gene Ontology (GO) analysis for genes available in final selected 238 pairs and also for genes in the top 5 % (95 percentile) for each of eight cancer types to report a significant number of genes involved in various biological functions and pathways. We also calculated various centrality measurement matrices for the largest connected component(s) in each of eight cancers and reported top genes and miRNAs with high centrality measurements. CONCLUSIONS: With its user-friendly interface, dynamic visualization and advanced queries, we also believe MMiRNA-Viewer offers an intuitive approach for visualizing and elucidating co-relationships between miRNAs and mRNAs of both tumor and normal samples. We suggest that miRNA and mRNA pairs with opposite fold changes of their expression and with inverted correlation values between tumor and normal samples might be most relevant for explaining the decoupling of mRNAs and their targeting miRNAs in tumor samples for certain cancer types. Yongsheng Bai, Lizhong Ding 0002, Steve Baker, Jenny M. Bai, Ethan Rath, Jianghong Wu, Hui Jiang 0002, Gary W. Stuart |
BMC Bioinform. | 1 |