EDBT 2026 Demo / reviewers in the wild / expert
Pora Kim
dblp:38/2445
· DBLP profile ↗
14ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0002-8321-6864ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FusionPub, a therapeutic landscape of human fusion genesabstractTo advance the development of fusion protein-targeting therapeutics, it is crucial to understand how patients with fusion oncoproteins have been treated and what small molecules have been studied. To fill this gap, we developed FusionPub, a knowledgebase of the therapeutic landscape of human fusion genes. We searched PubMed abstracts for 107K human fusion genes with 14K drugs. We also searched PubMed abstracts for 18K human fusion proteins, considering all gene synonyms of individual partner genes with 14K drugs. After manual curation, we found 17 342 records from 9623 PubMed abstracts. For these fusion genes and drugs, we provide a summary sentence of their relationship using the generative AI, details of the drugs, clinical trial records, common drug records from patient treatment history, etc. FusionPub will be a valuable and unique resource providing knowledge of the development of fusion gene targeting therapeutics. Himansu Kumar, Abayomi Adegunlehin, Pora Kim |
Briefings Bioinform. | 4 |
| 2025 | KinaseFusionDB: an integrative knowledge of kinase fusion proteins in multi-scalesabstractKinase fusion genes were the most targeted fusion gene group among multiple major cellular gene groups. Kinase inhibitors disrupt aberrant signaling cascades and inhibit tumor progression, yet the specific mechanisms of action of the U.S. Food and Drug Administration (FDA)-approved inhibitors in the context of kinase fusion oncoproteins remain largely unknown. This gap limits our ability to develop personalized therapies and next-generation kinase inhibitors. To address this, we developed a novel in silico pipeline for predicting 3D structures of kinase fusion proteins and performing structure-based virtual screening. This approach enables large-scale structural annotation and drug screening across pan-cancer kinase fusions. We present KinaseFusionDB, available at https://compbio.uth.edu/KinaseFusionDB, a comprehensive knowledgebase providing functional annotation of 7680 kinase fusion genes, 1399 predicted fusion protein structures, predicted Local Distance Difference Test (pLDDT)-based confidence scoring, and virtual screening data using FDA-approved kinase inhibitors. Our analysis revealed that most predicted structures showed high pLDDT scores (pLDDT >70) within conserved kinase domains. Structural alignment with known Protein Data Banks demonstrated shared structural motifs despite variation in fusion breakpoints. Virtual screening results highlighted repurposing opportunities and isoform-specific binding preferences. KinaseFusionDB is a valuable resource for investigating kinase fusion structure-function relationships and guiding the design of personalized and next-generation kinase inhibitor therapies. Himansu Kumar, Abayomi Adegunlehin, Loren Trowbridge, Leonardo Aguilar, Pora Kim |
Briefings Bioinform. | 6 |
| 2025 | AllergenAI: a deep learning model predicting allergenicity based on protein sequenceabstractBACKGROUND: Innovations in protein engineering offer promising solutions for redesigning allergenic proteins to minimize adverse reactions in sensitive individuals. Earlier models for predicting allergenicity have relied on the knowledge of physicochemical properties and sequence homology to assess the potential risk. However, to better understand the allergenic proteins’ sequence features, we need a novel sequence-based deep learning model for predicting allergenicity. RESULTS: We present a novel AI-based tool, AllergenAI, to quantify the allergenic potential of a protein’s sequence without using any other known features. Our study utilized allergenic protein sequence data archived in the three well-established databases, SDAP 2.0, COMPARE, and AlgPred 2, to train a convolutional neural network and assessed its prediction performance by cross-validation. We then used AllergenAI to find novel potential proteins of the cupin family in date palm, spinach, maize, and red clover plants with a high allergenicity score that might have an adverse allergenic effect on sensitive individuals. By analyzing the feature importance scores (FIS) of vicilins, we identified a proline-alanine-rich (P-A) motif in the top 50% of FIS regions that overlapped with known IgE epitope regions of vicilin allergens. We then used the approximately 1600 allergen structures in our SDAP database, in a pilot study to show the potential to incorporate 3D information in a CNN model. The prediction quality was slightly increased. CONCLUSION: Our allergenicity prediction study through the development of AllergenAI provides a foundation for identifying the critical features that distinguish allergenic proteins. Surendra S. Negi, Chengyuan Yang, Xiaobo Zhou 0005, Catherine H. Schein, Werner Braun, Pora Kim |
BMC Bioinform. | 7 |
| 2024 | FusionNW, a potential clinical impact assessment of kinases in pan-cancer fusion gene networkabstractKinase fusion genes are the most active fusion gene group in human cancer fusion genes. To help choose the clinically significant kinase so that the cancer patients that have fusion genes can be better diagnosed, we need a metric to infer the assessment of kinases in pan-cancer fusion genes rather than relying on the sample frequency expressed fusion genes. Most of all, multiple studies assessed human kinases as the drug targets using multiple types of genomic and clinical information, but none used the kinase fusion genes in their study. The assessment studies of kinase without kinase fusion gene events can miss the effect of one of the mechanisms that enhance the kinase function in cancer. To fill this gap, in this study, we suggest a novel way of assessing genes using a network propagation approach to infer how likely individual kinases influence the kinase fusion gene network composed of ~5K kinase fusion gene pairs. To select a better seed of propagation, we chose the top genes via dimensionality reduction like a principal component or latent layer information of six features of individual genes in pan-cancer fusion genes. Our approach may provide a novel way to assess of human kinases in cancer. Chengyuan Yang, Himansu Kumar, Pora Kim |
Briefings Bioinform. | 3 |
| 2023 | Systematic investigation of the homology sequences around the human fusion gene breakpoints in pan-cancer - bioinformatics study for a potential link to MMEJabstractMicrohomology-mediated end joining (MMEJ), an error-prone DNA damage repair mechanism, frequently leads to chromosomal rearrangements due to its ability to engage in promiscuous end joining of genomic instability and also leads to increasing mutational load at the sequences flanking the breakpoints (BPs). In this study, we systematically investigated the homology sequences around the genomic breakpoint area of human fusion genes, which were formed by the chromosomal rearrangements initiated by DNA double-strand breakage. Since the RNA-seq data is the typical data set to check the fusion genes, for the known exon junction fusion breakpoints identified from RNA-seq data, we have to infer the high chance of genomic breakpoint regions. For this, we utilized the high feature importance score area calculated from our recently developed fusion BP prediction model, FusionAI and identified 151 K microhomologies among ~24 K fusion BPs in 20 K fusion genes. From our multiple bioinformatics studies, we found a relationship between sequence homologies and the immune system. This in-silico study will provide novel knowledge on the sequence homologies around the coded structural variants. Pora Kim, Himansu Kumar, Chengyuan Yang, Ruihan Luo |
Briefings Bioinform. | 1 |
| 2023 | Genetic control of RNA editing in neurodegenerative diseaseabstractA-to-I RNA editing diversifies human transcriptome to confer its functional effects on the downstream genes or regulations, potentially involving in neurodegenerative pathogenesis. Its variabilities are attributed to multiple regulators, including the key factor of genetic variants. To comprehensively investigate the potentials of neurodegenerative disease-susceptibility variants from the view of A-to-I RNA editing, we analyzed matched genetic and transcriptomic data of 1596 samples across nine brain tissues and whole blood from two large consortiums, Accelerating Medicines Partnership-Alzheimer's Disease and Parkinson's Progression Markers Initiative. The large-scale and genome-wide identification of 95 198 RNA editing quantitative trait loci revealed the preferred genetic effects on adjacent editing events. Furthermore, to explore the underlying mechanisms of the genetic controls of A-to-I RNA editing, several top RNA-binding proteins were pointed out, such as EIF4A3, U2AF2, NOP58, FBL, NOP56 and DHX9, since their regulations on multiple RNA-editing events were probably interfered by these genetic variants. Moreover, these variants may also contribute to the variability of other molecular phenotypes associated with RNA editing, including the functions of 3 proteins, expressions of 277 genes and splicing of 449 events. All the analyses results shown in NeuroEdQTL (https://relab.xidian.edu.cn/NeuroEdQTL/) constituted a unique resource for the understanding of neurodegenerative pathogenesis from genotypes to phenotypes related to A-to-I RNA editing. Sijia Wu, Qiuping Xue, Mengyuan Yang 0001, Pora Kim, Xiaobo Zhou 0005 |
Briefings Bioinform. | 5 |
| 2021 | Landscape of drug-resistance mutations in kinase regulatory hotspotsabstractMore than 48 kinase inhibitors (KIs) have been approved by Food and Drug Administration. However, drug-resistance (DR) eventually occurs, and secondary mutations have been found in the previously targeted primary-mutated cancer cells. Cancer and drug research communities recognize the importance of the kinase domain (KD) mutations for kinasopathies. So far, a systematic investigation of kinase mutations on DR hotspots has not been done yet. In this study, we systematically investigated four types of representative mutation hotspots (gatekeeper, G-loop, αC-helix and A-loop) associated with DR in 538 human protein kinases using large-scale cancer data sets (TCGA, ICGC, COSMIC and GDSC). Our results revealed 358 kinases harboring 3318 mutations that covered 702 drug resistance hotspot residues. Among them, 197 kinases had multiple genetic variants on each residue. We further computationally assessed and validated the epidermal growth factor receptor mutations on protein structure and drug-binding efficacy. This is the first study to provide a landscape view of DR-associated mutation hotspots in kinase's secondary structures, and its knowledge will help the development of effective next-generation KIs for better precision medicine. Pora Kim, Junmei Wang, Zhongming Zhao |
Briefings Bioinform. | 1 |
| 2021 | miRactDB characterizes miRNA-gene relation switch between normal and cancer tissues across pan-cancerabstractIt has been increasingly accepted that microRNA (miRNA) can both activate and suppress gene expression, directly or indirectly, under particular circumstances. Yet, a systematic study on the switch in their interaction pattern between activation and suppression and between normal and cancer conditions based on multi-omics evidences is not available. We built miRactDB, a database for miRNA-gene interaction, at https://ccsm.uth.edu/miRactDB, to provide a versatile resource and platform for annotation and interpretation of miRNA-gene relations. We conducted a comprehensive investigation on miRNA-gene interactions and their biological implications across tissue types in both tumour and normal conditions, based on TCGA, CCLE and GTEx databases. We particularly explored the genetic and epigenetic mechanisms potentially contributing to the positive correlation, including identification of miRNA binding sites in the gene coding sequence (CDS) and promoter regions of partner genes. Integrative analysis based on this resource revealed that top-ranked genes derived from TCGA tumour and adjacent normal samples share an overwhelming part of biological processes, which are quite different than those from CCLE and GTEx. The most active miRNAs predicted to target CDS and promoter regions are largely overlapped. These findings corroborate that adjacent normal tissues might have undergone significant molecular transformations towards oncogenesis before phenotypic and histological change; and there probably exists a small yet critical set of miRNAs that profoundly influence various cancer hallmark processes. miRactDB provides a unique resource for the cancer and genomics communities to screen, prioritize and rationalize their candidates of miRNA-gene interactions, in both normal and cancer scenarios. Hua Tan, Pora Kim, Peiqing Sun, Xiaobo Zhou 0005 |
Briefings Bioinform. | 2 |
| 2021 | ADeditome provides the genomic landscape of A-to-I RNA editing in Alzheimer's diseaseabstractA-to-I RNA editing, contributing to nearly 90% of all editing events in human, has been reported to involve in the pathogenesis of Alzheimer's disease (AD) due to its roles in brain development and immune regulation, such as the deficient editing of GluA2 Q/R related to cell death and memory loss. Currently, there are urgent needs for the systematic annotations of A-to-I RNA editing events in AD. Here, we built ADeditome, the annotation database of A-to-I RNA editing in AD available at https://ccsm.uth.edu/ADeditome, aiming to provide a resource and reference for functional annotation of A-to-I RNA editing in AD to identify therapeutically targetable genes in an individual. We detected 1676 363 editing sites in 1524 samples across nine brain regions from ROSMAP, MayoRNAseq and MSBB. For these editing events, we performed multiple functional annotations including identification of specific and disease stage associated editing events and the influence of editing events on gene expression, protein recoding, alternative splicing and miRNA regulation for all the genes, especially for AD-related genes in order to explore the pathology of AD. Combing all the analysis results, we found 108 010 and 26 168 editing events which may promote or inhibit AD progression, respectively. We also found 5582 brain region-specific editing events with potentially dual roles in AD across different brain regions. ADeditome will be a unique resource for AD and drug research communities to identify therapeutically targetable editing events. Significance: ADeditome is the first comprehensive resource of the functional genomics of individual A-to-I RNA editing events in AD, which will be useful for many researchers in the fields of AD pathology, precision medicine, and therapeutic researches. Sijia Wu, Mengyuan Yang 0001, Pora Kim, Xiaobo Zhou 0005 |
Briefings Bioinform. | 3 |
| 2021 | ExonSkipAD provides the functional genomic landscape of exon skipping events in Alzheimer's diseaseabstractExon skipping (ES), the most common alternative splicing event, has been reported to contribute to diverse human diseases due to the loss of functional domains/sites or frameshifting of the open reading frame (ORF) and noticed as therapeutic targets. Accumulating transcriptomic studies of aging brains show the splicing disruption is a widespread hallmark of neurodegenerative diseases such as Alzheimer's disease (AD). Here, we built ExonSkipAD, the ES annotation database aiming to provide a resource/reference for functional annotation of ES events in AD and identify therapeutic targets in exon units. We identified 16 414 genes that have ~156 K, ~ 69 K, ~ 231 K ES events from the three representative AD cohorts of ROSMAP, MSBB and Mayo, respectively. For these ES events, we performed multiple functional annotations relating to ES mechanisms or downstream. Specifically, through the functional feature retention studies followed by the open reading frames (ORFs), we identified 275 important cellular regulators that might lose their cellular regulator roles due to exon skipping in AD. ExonSkipAD provides twelve categories of annotations: gene summary, gene structures and expression levels, exon skipping events with PSIs, ORF annotation, exon skipping events in the canonical protein sequence, 3'-UTR located exon skipping events lost miRNA-binding sites, SNversus in the skipped exons with a depth of coverage, AD stage-associated exon skipping events, splicing quantitative trait loci (sQTLs) in the skipped exons, correlation with RNA-binding proteins, and related drugs & diseases. ExonSkipAD will be a unique resource of transcriptomic diversity research for understanding the mechanisms of neurodegenerative disease development and identifying potential therapeutic targets in AD. Significance AS the first comprehensive resource of the functional genomics of the alternative splicing events in AD, ExonSkipAD will be useful for many researchers in the fields of pathology, AD genomics and precision medicine, and pharmaceutical and therapeutic researches. Mengyuan Yang 0001, Ke Yiya, Pora Kim, Xiaobo Zhou 0005 |
Briefings Bioinform. | 3 |
| 2019 | Distinct telomere length and molecular signatures in seminoma and non-seminoma of testicular germ cell tumorabstractTesticular germ cell tumors (TGCTs) are classified into two main subtypes, seminoma (SE) and non-seminoma (NSE), but their molecular distinctions remain largely unexplored. Here, we used expression data for mRNAs and microRNAs (miRNAs) from The Cancer Genome Atlas (TCGA) to perform a systematic investigation to explain the different telomere length (TL) features between NSE (n = 48) and SE (n = 55). We found that TL elongation was dominant in NSE, whereas TL shortening prevailed in SE. We further showed that both mRNA and miRNA expression profiles could clearly distinguish these two subtypes. Notably, four telomere-related genes (TelGenes) showed significantly higher expression and positively correlated with telomere elongation in NSE than SE: three telomerase activity-related genes (TERT, WRAP53 and MYC) and an independent telomerase activity gene (ZSCAN4). We also found that the expression of genes encoding Yamanaka factors was positively correlated with telomere lengthening in NSE. Among them, SOX2 and MYC were highly expressed in NSE versus SE, while POU5F1 and KLF4 had the opposite patterns. These results suggested that enhanced expression of both TelGenes (TERT, WRAP53, MYC and ZSCAN4) and Yamanaka factors might induce telomere elongation in NSE. Conversely, the relative lack of telomerase activation and low expression of independent telomerase activity pathway during cell division may be contributed to telomere shortening in SE. Taken together, our results revealed the potential molecular profiles and regulatory roles involving the TL difference between NSE and SE, and provided a better molecular understanding of this complex disease. Hua Sun 0003, Pora Kim, Peilin Jia, Aekyung Park, Zhongming Zhao |
Briefings Bioinform. | 2 |
| 2018 | Kinase impact assessment in the landscape of fusion genes that retain kinase domains: a pan-cancer studyabstractAssessing the impact of kinase in gene fusion is essential for both identifying driver fusion genes (FGs) and developing molecular targeted therapies. Kinase domain retention is a crucial factor in kinase fusion genes (KFGs), but such a systematic investigation has not been done yet. To this end, we analyzed kinase domain retention (KDR) status in chimeric protein sequences of 914 KFGs covering 312 kinases across 13 major cancer types. Based on 171 kinase domain-retained KFGs including 101 kinases, we studied their recurrence, kinase groups, fusion partners, exon-based expression depth, short DNA motifs around the break points and networks. Our results, such as more KDR than 5'-kinase fusion genes, combinatorial effects between 3'-KDR kinases and their 5'-partners and a signal transduction-specific DNA sequence motif in the break point intronic sequences, supported positive selection on 3'-kinase fusion genes in cancer. We introduced a degree-of-frequency (DoF) score to measure the possible number of KFGs of a kinase. Interestingly, kinases with high DoF scores tended to undergo strong gene expression alteration at the break points. Furthermore, our KDR gene fusion network analysis revealed six of the seven kinases with the highest DoF scores (ALK, BRAF, MET, NTRK1, NTRK3 and RET) were all observed in thyroid carcinoma. Finally, we summarized common features of 'effective' (highly recurrent) kinases in gene fusions such as expression alteration at break point, redundant usage in multiple cancer types and 3'-location tendency. Collectively, our findings are useful for prioritizing driver kinases and FGs and provided insights into KFGs' clinical implications. Pora Kim, Peilin Jia, Zhongming Zhao |
Briefings Bioinform. | 1 |
| 2008 | Intelligent Positioning and Optimal Diversity Schemes for Mobile Agents in Ubiquitous NetworksabstractA lot of schemes have been proposed for the establishment of u-city. Especially, most of the schemes are based on ubiquitous networks. In this paper, it is assumed that the ubiquitous networks include some mobile agents. Therefore, an intelligent positioning scheme is proposed for the efficient positioning of the mobile agents in the ubiquitous networks. The approach consists of location detection and location tracking. Furthermore, an optimal diversity technique is presented in this paper to improve the positioning performance. Simulation results indicate that using the suggested algorithms, the excellent detection and tracking performance can be achieved in ubiquitous networks for u-city. Pora Kim, Sekchin Chang |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2007 | An Intelligent Positioning Scheme for Mobile Agents in Ubiquitous Networks for U-City
Pora Kim, Sekchin Chang |
KES-AMSTA | 1 |