EDBT 2026 Demo / reviewers in the wild / expert
Jijun Tang
dblp:21/234
· DBLP profile ↗
121ranked-venue papers
5as first author
64since 2021 · last 2026
0000-0002-6377-536XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 89 · 3 first-author · 49 since 2021Artificial intelligence and machine learning · 19 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 3 since 2021Theory of computation · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Closer to Biological Mechanism: Drug-Drug Interaction Prediction from the Perspective of PharmacophoreabstractDrug combinations are widely used in modern medicine but may cause severe adverse drug reactions. Therefore, making effective drug-drug interactions (DDI) prediction is crucial for pharmacovigilance. Existing DDI prediction models are typically built from a structural perspective, assuming that drugs with similar molecular structures may exhibit similar interactions. However, such approaches overlook the biological mechanisms underlying DDI in the human body. This not only weakens the generalization ability of the model, but also makes its interpretability less convincing. Inspired by this, we propose a new method called PC-DDI. Unlike structure-based models, PC-DDI utilizes pharmacophores as basic unit, and designs a complete pharmacophore feature processing framework. It further constructs a pharmacophore-based bipartite graph to model interactions between pharmacophores. This approach allows us to explore the underlying mechanisms of DDI from a functional perspective. We also design a spatial attention weight graph convolution module to optimize the message passing process by integrating pharmacophore position features with node features. Furthermore, we apply causal inference to identify key pharmacophores in pharmacophore bipartite graph, enhancing the interpretability. Compared with the SOTA, PC-DDI achieves an accuracy improvement of 1.84% under the transductive setting and consistently outperforms others in all other experiments. Mingliang Dou, Linfeng Wen 0005, Jinyang Xie, Jijun Tang, Shiqiang Ma, Fei Guo 0001 |
AAAI | 4 |
| 2026 | Enhancing Sample Discrimination: Drug-Drug Interaction Prediction Based on Bidirectional Event Semantics Guidance
Shiqiang Ma, Mingliang Dou, Fei Guo 0001, Jijun Tang |
ICIC (15) | 5 |
| 2026 | Refprogen: a reference-guided molecular generation model with protein-ligand joint representation for property-aware drug design
Chengwei Ai, Jijun Tang, Fei Guo 0001 |
Expert Syst. Appl. | 4 |
| 2026 | MultiPert: An adversarial alignment and dual attention framework for single-cell multi-omics perturbation predictionabstractPrecise prediction of perturbation responses is essential in systems biology research, as it plays a pivotal role in characterizing cellular identities and elucidating the regulatory mechanisms of biological pathways. Existing perturbation-responses prediction approaches are predominantly confined to single-modality transcriptomic data, limiting their capacity to capture cross-layer molecular effects. Here, we present MultiPert, a deep learning framework specifically designed for predicting perturbation responses in single-cell multi-omics data. MultiPert employs modality-specific encoders with dedicated pretraining, integrates perturbation through a dual-attention mechanism, and achieves cross-modal alignment via adversarial training. Benchmarking on human THP-1 and kidney multi-omics datasets demonstrates that MultiPert reliably predicts both perturbed gene expression and protein abundance profiles, achieving superior accuracy and stability compared to state-of-the-art strategies. MultiPert generalizes to unseen perturbations and uncovers regulatory mechanisms of immune checkpoint molecules based on perturbed proteomic predictions. In addition, enrichment analyses of perturbed transcriptomic predictions reveal immune-related pathways. By providing an integrated and interpretable framework, MultiPert expands the scope of perturbation modeling at the multi-omics level, thereby offering a robust methodological foundation for comprehensive research into pathogenesis and drug discovery. Xinyue Tang, Jiawei Li 0018, Cheng Liang 0001, Jijun Tang, Fei Guo 0001 |
PLoS Comput. Biol. | 5 |
| 2025 | MF-DocDDI: Drug Entity Multi-Feature Fusion for Document-Level Drug-Drug Interaction Relation ExtractionabstractDrug-drug interactions (DDIs) are crucial in clinical medicine, as they can lead to adverse events. Existing DDI extraction methods focus on sentence-level tasks, limiting their ability to identify cross-sentence DDIs. Moreover, the only document-level method available considers only internal drug features, leading to suboptimal performance. To address this, we propose MF-DocDDI, a document-level DDI extraction model using drug entity multi-feature fusion. We first construct a document-level dataset based on DDI Extraction 2013. Then, we introduce document-entity embeddings to capture internal drug features and employ a simplified U-shaped network to extract external features. Finally, we integrate these features to enhance interaction modeling. Experimental results show MF-DocDDI outperforms existing methods, improving the F1 score by 5.33 %. Case studies confirm its ability to identify cross-sentence DDIs, such as (naloxone, morphine) and (HEXALEN, cisplatin). Beyond DDI extraction, MF-DocDDI can be applied to other biomedical tasks like protein-protein interaction (PPI) extraction. Mingliang Dou, Jijun Tang, Fei Guo 0001 |
BIBM | 3 |
| 2025 | DCT-Net: Dual-Branch CT Reconstruction from Orthogonal X-Rays with Diffusion Model and Contrastive Learning
Jijun Tang, Zhijun Liao |
MICCAI (3) | 3 |
| 2025 | HyperPhS: a pharmacophore-guided multimodal representation framework for metabolic stability prediction through contrastive hypergraph learningabstractMOTIVATION: Metabolic stability is crucial in the early stage of drug discovery and development. Drug candidate screening and optimization can be streamlined through the accurate prediction of stability. Functional groups within drug molecules are known as pharmacophores, which bind directly to receptors or biological macromolecules to produce biological effects, thereby affecting metabolic stability. Therefore, determining metabolic stability via the pharmacophore groups remains a significant challenge. RESULTS: To address these issues, we propose a Pharmacophore-guided Hypergraph representation framework for predicting metabolic Stability (HyperPhS). In this study, we introduce a hypergraph-based method to extract features from metabolic pharmacophores with multi-view representation and contrastive learning. In particular, we introduce a pharmacophore-based contrastive learning encoder that captures the consistency between functional and nonfunctional structures. Our method applies ChatGPT simultaneously to metabolites and heterogeneous encoders and integrates multimodal representations by using attention-driven fusion modules coupled with fully connected neural networks. On the HLM dataset, HyperPhS achieves outstanding performance with 87.6% in AUC and 62.6% in MCC, alongside an external test AUC of 88.3%. In addition, pharmacophore groups studied by HyperPhS are validated for their interpretability through case studies. Overall, HyperPhS is an effective and interpretable tool for determining metabolic stability, identifying critical functional groups, and optimizing compounds. AVAILABILITY AND IMPLEMENTATION: The code and data are available at https://github.com/xiaoyiliu-usc/HyperPhS. Chenglong Kang, Chengwei Ai, Hongpeng Yang, Jijun Tang, Fei Guo 0001 |
Bioinform. | 6 |
| 2025 | Multi-granularity semantic relational mapping for image caption
Nan Gao 0001, Renyuan Yao, Peng Chen 0008, Ronghua Liang, Guodao Sun, Jijun Tang |
Expert Syst. Appl. | 6 |
| 2025 | RFAE: A high-robust feature selector based on fractal autoencoder
Jingfeng Ou, Jiawei Li 0018, Zhiliang Xia, Shurui Dai, Limin Jiang, Jijun Tang |
Expert Syst. Appl. | 7 |
| 2024 | ACPNet: Enhancing Small-Scale Dieases Detection in Panoramic X-raysabstractDeep learning-based disease detection can automatically identify dental diseases in panoramic X-rays and improve the accuracy and efficiency of doctors’ diagnoses. However, due to the complex data distribution of panoramic oral X-rays, significant scale differences among lesions, and the presence of many small-scale diseases, automated disease detection in panoramic oral X-rays faces considerable challenges. To alleviate the aforementioned issues, we propose ACPNet, which introduces a novel two-stage approach for detecting small-scale dental diseases in panoramic X-rays using the Contextual Attention Alignment Network (CAAN) and the Point-to-Patch Module (PTPM). To the best of our knowledge, we are the first to explore the detection of small-scale dental diseases in panoramic X-rays under limited sample conditions. Specifically, CAAN integrates deformable convolution with the global attention mechanism of transformer attention, enabling the model to more accurately extract small target foreground features in the complex background of panoramic X-rays. PTPM employs key point detection and cascade dynamic patches to adjust the bounding boxes of lesions, ensuring that small-scale diseases have sufficient high-quality proposals, thereby enhancing detector performance. Additionally, we collected a dataset containing 1157 instances of dental diseases to validate the effectiveness of our algorithm. Extensive experiments demonstrate that ACPNet achieves state-of-the-art performance, highlighting its superiority over baseline and other detection methods. Nan Gao 0001, Junchao Zhu, Peng Chen 0008, Jijun Tang, Ronghua Liang |
BIBM | 4 |
| 2024 | Multi-Task Driven Multi-Level Dynamical Fusion for Single-Cell Multi-Omics Cell Type AnnotationabstractThe emergence of single-cell multi-omics sequencing technology has enabled the simultaneous profiling of diverse omics data within individual cells. It offers a more comprehensive perspective on cellular phenotypes and heterogeneity. However, single-cell multi-omics data are inherently high-dimensional and heterogeneous. Due to technical limitations and scarce starting materials, the data are often affected by noise and dropout effects. To address these challenges, we propose a novel multitask driven multi-level dynamical fusion algorithm for single-cell multi-omics cell type annotation, named scMMDyn. Our approach incorporates reconstruction and classification auxiliary tasks to guide the training of trustworthy modules at both the feature and modality levels. It executes dynamical fusion during these stages and finally achieves cross-modality fusion via an attention mechanism. This method effectively mitigates data quality issues through reconstruction tasks and feature-level dynamical fusion while providing interpretability at both feature and modality levels. Experimental results across diverse single-cell multi-omics datasets show that our method surpasses existing approaches in cell type annotation. Jiawei Li 0018, Shizhan Chen, Zongbo Han, Jijun Tang, Fei Guo 0001 |
BIBM | 5 |
| 2024 | scCADE: A Superior Tool for Predicting Perturbation Responses in Single-Cell Gene Expression Using Contrastive Learning and Attention MechanismsabstractThe advent of single-cell transcriptomics has revolutionized our ability to analyze cellular heterogeneity and dynamics at a fine resolution, yet covering the vast array of potential perturbations remains challenging due to biological variability. To address this, we propose scCADE, a novel computational approach utilizing contrastive learning and an attention mechanism to decouple gene expression signatures and predict cellular responses to perturbations. scCADE excels in predicting responses in cells to perturbations observed in other cells but not yet seen in the target cells. Through rigorous ablation studies and validation across three datasets involving drug and gene editing perturbations, scCADE consistently outperformed existing methods, underscoring its efficacy and potential to advance genomics and personalized medicine by accurately forecasting responses to novel perturbations. Jingfeng Ou, Jiawei Li 0018, Zhiliang Xia, Shurui Dai, Yulian Ding, Limin Jiang, Jijun Tang |
BIBM | 8 |
| 2024 | SiamSegNet: A multimodal Segmentation Method Based on Cross-modal Generation for Medical Image Segmentation
Shiqiang Ma, Fei Guo 0001, Jijun Tang |
DASFAA (3) | 3 |
| 2024 | RetroCaptioner: beyond attention in end-to-end retrosynthesis transformer via contrastively captioned learnable graph representationabstractMOTIVATION: Retrosynthesis identifies available precursor molecules for various and novel compounds. With the advancements and practicality of language models, Transformer-based models have increasingly been used to automate this process. However, many existing methods struggle to efficiently capture reaction transformation information, limiting the accuracy and applicability of their predictions. RESULTS: We introduce RetroCaptioner, an advanced end-to-end, Transformer-based framework featuring a Contrastive Reaction Center Captioner. This captioner guides the training of dual-view attention models using a contrastive learning approach. It leverages learned molecular graph representations to capture chemically plausible constraints within a single-step learning process. We integrate the single-encoder, dual-encoder, and encoder-decoder paradigms to effectively fuse information from the sequence and graph representations of molecules. This involves modifying the Transformer encoder into a uni-view sequence encoder and a dual-view module. Furthermore, we enhance the captioning of atomic correspondence between SMILES and graphs. Our proposed method, RetroCaptioner, achieved outstanding performance with 67.2% in top-1 and 93.4% in top-10 exact matched accuracy on the USPTO-50k dataset, alongside an exceptional SMILES validity score of 99.4%. In addition, RetroCaptioner has demonstrated its reliability in generating synthetic routes for the drug protokylol. AVAILABILITY AND IMPLEMENTATION: The code and data are available at https://github.com/guofei-tju/RetroCaptioner. Chengwei Ai, Hongpeng Yang, Ruihan Dong, Jijun Tang, Shuangjia Zheng, Fei Guo 0001 |
Bioinform. | 5 |
| 2024 | TranSiam: Aggregating multi-modal visual features with locality for medical image segmentation
Shiqiang Ma, Junhai Xu, Jijun Tang, Shengfeng He, Fei Guo 0001 |
Expert Syst. Appl. | 4 |
| 2024 | PPRTGI: A Personalized PageRank Graph Neural Network for TF-Target Gene Interaction DetectionabstractTranscription factors (TFs) regulation is required for the vast majority of biological processes in living organisms. Some diseases may be caused by improper transcriptional regulation. Identifying the target genes of TFs is thus critical for understanding cellular processes and analyzing disease molecular mechanisms. Computational approaches can be challenging to employ when attempting to predict potential interactions between TFs and target genes. In this paper, we present a novel graph model (PPRTGI) for detecting TF-target gene interactions using DNA sequence features. Feature representations of TFs and target genes are extracted from sequence embeddings and biological associations. Then, by combining the aggregated node feature with graph structure, PPRTGI uses a graph neural network with personalized PageRank to learn interaction patterns. Finally, a bilinear decoder is applied to predict interaction scores between TF and target gene nodes. We designed experiments on six datasets from different species. The experimental results show that PPRTGI is effective in regulatory interaction inference, with our proposed model achieving an area under receiver operating characteristic score of 93.87% and an area under precision-recall curves score of 88.79% on the human dataset. This paper proposes a new method for predicting TF-target gene interactions, which provides new insights into modeling molecular networks and can thus be used to gain a better understanding of complex biological systems. Jiawei Li 0018, Ibrahim Zamit, Fei Guo 0001, Jijun Tang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2024 | DMAMP: A Deep-Learning Model for Detecting Antimicrobial Peptides and Their Multi-ActivitiesabstractDue to the broad-spectrum and high-efficiency antibacterial activity, antimicrobial peptides (AMPs) and their functions have been studied in the field of drug discovery. Using biological experiments to detect the AMPs and corresponding activities require a high cost, whereas computational technologies do so for much less. Currently, most computational methods solve the identification of AMPs and their activities as two independent tasks, which ignore the relationship between them. Therefore, the combination and sharing of patterns for two tasks is a crucial problem that needs to be addressed. In this study, we propose a deep learning model, called DMAMP, for detecting AMPs and activities simultaneously, which is benefited from multi-task learning. The first stage is to utilize convolutional neural network models and residual blocks to extract the sharing hidden features from two related tasks. The next stage is to use two fully connected layers to learn the distinct information of two tasks. Meanwhile, the original evolutionary features from the peptide sequence are also fed to the predictor of the second task to complement the forgotten information. The experiments on the independent test dataset demonstrate that our method performs better than the single-task model with 4.28% of Matthews Correlation Coefficient (MCC) on the first task, and achieves 0.2627 of an average MCC which is higher than the single-task model and two existing methods for five activities on the second task. To understand whether features derived from the convolutional layers of models capture the differences between target classes, we visualize these high-dimensional features by projecting into 3D space. In addition, we show that our predictor has the ability to identify peptides that achieve activity against Severe Acute Respiratory Syndrome Coronavirus-2 (SARS-CoV-2). We hope that our proposed method can give new insights into the discovery of novel antiviral peptide drugs. Qiaozhen Meng, Genlang Chen, Shixin Zheng, Yulai Lin, Jijun Tang, Fei Guo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2024 | Prediction of LncRNA-Protein Interactions Based on Kernel Combinations and Graph Convolutional NetworksabstractThe complexes of long non-coding RNAs bound to proteins can be involved in regulating life activities at various stages of organisms. However, in the face of the growing number of lncRNAs and proteins, verifying LncRNA-Protein Interactions (LPI) based on traditional biological experiments is time-consuming and laborious. Therefore, with the improvement of computing power, predicting LPI has met new development opportunity. In virtue of the state-of-the-art works, a framework called LncRNA-Protein Interactions based on Kernel Combinations and Graph Convolutional Networks (LPI-KCGCN) has been proposed in this article. We first construct kernel matrices by taking advantage of extracting both the lncRNAs and protein concerning the sequence features, sequence similarity features, expression features, and gene ontology. Then reconstruct the existent kernel matrices as the input of the next step. Combined with known LPI interactions, the reconstructed similarity matrices, which can be used as features of the topology map of the LPI network, are exploited in extracting potential representations in the lncRNA and protein space using a two-layer Graph Convolutional Network. The predicted matrix can be finally obtained by training the network to produce scoring matrices w.r.t. lncRNAs and proteins. Different LPI-KCGCN variants are ensemble to derive the final prediction results and testify on balanced and unbalanced datasets. The 5-fold cross-validation shows that the optimal feature information combination on a dataset with 15.5% positive samples has an AUC value of 0.9714 and an AUPR value of 0.9216. On another highly unbalanced dataset with only 5% positive samples, LPI-KCGCN also has outperformed the state-of-the-art works, which achieved an AUC value of 0.9907 and an AUPR value of 0.9267. Dongdong Mao, Jijun Tang, Zhijun Liao, Shengyong Chen |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | BTCN: Bridging the Gap Between Pre-trained and Downstream Models for Endoscopic Caries DetectionabstractAlthough deep learning has been widely applied in the field of dental caries detection, there are still certain challenges that need to be addressed. The limitations of sharing the same backbone between the pre-trained model and the downstream model hinder the feature alignment capability of self-supervised learning (SSL) during the fine-tuning stage, leading to incomplete transfer from the pre-trained model to the downstream model. To address this challenge, we introduce an SSL pre-trained model called Bi-branches Transformer CNN Network (BTCN). BTCN adopts a parallel structure combining the CNN and Transformer branches. This parallel structure allows the pre-trained model to capture additional global representations, which helps alleviate feature differences during fine-tuning and better adapt to downstream detection models. Additionally, to further enhance the fusion quality of the bi-branches encoder, we introduced the Multi-layer Supervision Strategy (MSS) to increase the supervision on features at different layers. To validate the effectiveness of our approach, we collected a dedicated dataset for caries detection, comprising 1039 endoscopic images of dental caries. Through extensive experimental research, our results demonstrate the effectiveness of the proposed BTCN and MSS, showing significant improvements compared to the current state-of-the-art methods. Nan Gao 0001, Peng Chen 0008, Yukai Li, Jijun Tang, Ronghua Liang, Tianshuang Liu |
BIBM | 4 |
| 2023 | MixUNet: Mix the 2D and 3D Models for Robust Medical Image SegmentationabstractBrain tumor segmentation is pivotal in the diagnosis and treatment of brain tumors. As functional imaging technologies like CT and MR advance, analyzing 3D medical image data becomes more time-consuming. Several challenges exist in 3D medical image segmentation: 1) 2D networks, when applied to 3D segmentation tasks, suffer from a lack of 3D structural information. 2) Pure 3D networks, due to their vast parameter count and smaller training sample, are susceptible to overfitting. 3) Current 2.5D networks do not fully leverage the available 3D structural information. In this study, we introduce the Mix-UNet, a multi-branch network that synergizes 2D and 3D networks. This design preserves essential 3D structural details for precise segmentation while ensuring computational efficiency. Our model comprises two main branches and a fusion module: a 2D branch for coarse segmentation without 3D structural information, a 3D branch to capture comprehensive 3D structural details, and a fusion module for pixel-level integration to produce the final segmentation. Experimental results demonstrate the model’s ability to reduce parameter count, increase robustness, and maintain high precision. When tested on the BraTS 2020 validation dataset, our model achieved mean dice coefficients of 90.4%, 80.7%, and 71.2% for the whole tumor, tumor core, and enhancing tumor, respectively, with only 2.2M parameters. Jiawei Li 0018, Shizhan Chen, Shiqiang Ma, Fei Guo 0001, Jijun Tang |
BIBM | 5 |
| 2023 | Prediction of LncRNA-Protein Interactions Based on Multi-kernel Fusion and Graph Auto-Encoders
Dongdong Mao, Ruilin Wu, Yankai Wu, Jinxuan Wang, Jijun Tang, Zhijun Liao |
ICIC (3) | 7 |
| 2023 | DETA-Net: A Dual Encoder Network with Text-Guided Attention Mechanism for Skin-Lesions Segmentation
Jijun Tang, Zhijun Liao |
ICIC (3) | 3 |
| 2023 | IK-DDI: a novel framework based on instance position embedding and key external text for DDI extractionabstractDetermining drug-drug interactions (DDIs) is an important part of pharmacovigilance and has a vital impact on public health. Compared with drug trials, obtaining DDI information from scientific articles is a faster and lower cost but still a highly credible approach. However, current DDI text extraction methods consider the instances generated from articles to be independent and ignore the potential connections between different instances in the same article or sentence. Effective use of external text data could improve prediction accuracy, but existing methods cannot extract key information from external data accurately and reasonably, resulting in low utilization of external data. In this study, we propose a DDI extraction framework, instance position embedding and key external text for DDI (IK-DDI), which adopts instance position embedding and key external text to extract DDI information. The proposed framework integrates the article-level and sentence-level position information of the instances into the model to strengthen the connections between instances generated from the same article or sentence. Moreover, we introduce a comprehensive similarity-matching method that uses string and word sense similarity to improve the matching accuracy between the target drug and external text. Furthermore, the key sentence search method is used to obtain key information from external data. Therefore, IK-DDI can make full use of the connection between instances and the information contained in external text data to improve the efficiency of DDI extraction. Experimental results show that IK-DDI outperforms existing methods on both macro-averaged and micro-averaged metrics, which suggests our method provides complete framework that can be used to extract relationships between biomedical entities and process external text data. Mingliang Dou, Jiaqi Ding, Genlang Chen, Junwen Duan, Fei Guo 0001, Jijun Tang |
Briefings Bioinform. | 6 |
| 2023 | MVML-MPI: Multi-View Multi-Label Learning for Metabolic Pathway InferenceabstractDevelopment of robust and effective strategies for synthesizing new compounds, drug targeting and constructing GEnome-scale Metabolic models (GEMs) requires a deep understanding of the underlying biological processes. A critical step in achieving this goal is accurately identifying the categories of pathways in which a compound participated. However, current machine learning-based methods often overlook the multifaceted nature of compounds, resulting in inaccurate pathway predictions. Therefore, we present a novel framework on Multi-View Multi-Label Learning for Metabolic Pathway Inference, hereby named MVML-MPI. First, MVML-MPI learns the distinct compound representations in parallel with corresponding compound encoders to fully extract features. Subsequently, we propose an attention-based mechanism that offers a fusion module to complement these multi-view representations. As a result, MVML-MPI accurately represents and effectively captures the complex relationship between compounds and metabolic pathways and distinguishes itself from current machine learning-based methods. In experiments conducted on the Kyoto Encyclopedia of Genes and Genomes pathways dataset, MVML-MPI outperformed state-of-the-art methods, demonstrating the superiority of MVML-MPI and its potential to utilize the field of metabolic pathway design, which can aid in optimizing drug-like compounds and facilitating the development of GEMs. The code and data underlying this article are freely available at https://github.com/guofei-tju/MVML-MPI. Contact: [email protected], [email protected] or [email protected]. Hongpeng Yang, Chengwei Ai, Yijie Ding, Fei Guo 0001, Jijun Tang |
Briefings Bioinform. | 6 |
| 2023 | Improved structure-related prediction for insufficient homologous proteins using MSA enhancement and pre-trained language modelabstractIn recent years, protein structure problems have become a hotspot for understanding protein folding and function mechanisms. It has been observed that most of the protein structure works rely on and benefit from co-evolutionary information obtained by multiple sequence alignment (MSA). As an example, AlphaFold2 (AF2) is a typical MSA-based protein structure tool which is famous for its high accuracy. As a consequence, these MSA-based methods are limited by the quality of the MSAs. Especially for orphan proteins that have no homologous sequence, AlphaFold2 performs unsatisfactorily as MSA depth decreases, which may pose a barrier to its widespread application in protein mutation and design problems in which there are no rich homologous sequences and rapid prediction is needed. In this paper, we constructed two standard datasets for orphan and de novo proteins which have insufficient/none homology information, called Orphan62 and Design204, respectively, to fairly evaluate the performance of the various methods in this case. Then, depending on whether or not utilizing scarce MSA information, we summarized two approaches, MSA-enhanced and MSA-free methods, to effectively solve the issue without sufficient MSAs. MSA-enhanced model aims to improve poor MSA quality from the data source by knowledge distillation and generation models. MSA-free model directly learns the relationship between residues on enormous protein sequences from pre-trained models, bypassing the step of extracting the residue pair representation from MSA. Next, we evaluated the performance of four MSA-free methods (trRosettaX-Single, TRFold, ESMFold and ProtT5) and MSA-enhanced (Bagging MSA) method compared with a traditional MSA-based method AlphaFold2, in two protein structure-related prediction tasks, respectively. Comparison analyses show that trRosettaX-Single and ESMFold which belong to MSA-free method can achieve fast prediction ($\sim\! 40$s) and comparable performance compared with AF2 in tertiary structure prediction, especially for short peptides, $\alpha $-helical segments and targets with few homologous sequences. Bagging MSA utilizing MSA enhancement improves the accuracy of our trained base model which is an MSA-based method when poor homology information exists in secondary structure prediction. Our study provides biologists an insight of how to select rapid and appropriate prediction tools for enzyme engineering and peptide drug development. CONTACT: [email protected], [email protected]. Qiaozhen Meng, Fei Guo 0001, Jijun Tang |
Briefings Bioinform. | 3 |
| 2023 | CoMutDB: the landscape of somatic mutation co-occurrence in cancersabstractMOTIVATION: Somatic mutation co-occurrence has been proven to have a profound effect on tumorigenesis. While some studies have been conducted on co-mutations, a centralized resource dedicated to co-mutations in cancer is still lacking. RESULTS: Using multi-omics data from over 30 000 subjects and 1747 cancer cell lines, we present the Cancer co-mutation database (CoMutDB), the most comprehensive resource devoted to describing cancer co-mutations and their characteristics. AVAILABILITY AND IMPLEMENTATION: The data underlying this article are available in the online database CoMutDB: http://www.innovebioinfo.com/Database/CoMutDB/Home.php. Limin Jiang, Jijun Tang |
Bioinform. | 3 |
| 2023 | A multi-scale multi-model deep neural network via ensemble strategy on high-throughput microscopy image for protein subcellular localization
Jiaqi Ding, Junhai Xu, Jianguo Wei, Jijun Tang, Fei Guo 0001 |
Expert Syst. Appl. | 4 |
| 2023 | Boosting Short Text Classification by Solving the OOV ProblemabstractIn the field of natural language processing, text classification has received a lot of attention. Compared with long texts, short texts have fewer words and lack contextual semantic information. Existing approaches enrich short text information by linking the external knowledge graph, but they ignore the out-of-vocabulary (OOV) problem during entity linking, especially when dealing with domain-oriented data, which has some rare words or domain-specific nouns. In this paper, to alleviate the OOV problem caused by linking the external knowledge graph(KG), we propose a domain knowledge graph and entity complementation strategy to improve the performance of short text classification. Specifically, the external knowledge graph is used to enrich the information of short texts. The self-build domain knowledge graph is used to solve the problem of entities failing to link to the external knowledge graph. Finally, we conduct experiments on various datasets: 1. a labeled Chinese electronic domain dataset; 2. an open-source dataset to test the performance of our algorithm in different data distribution scenarios. The results demonstrate our dual knowledge graph model outperforms the state-of-the-art short text classification methods, especially when the OOV problem is severe. Nan Gao 0001, Peng Chen 0008, Jijun Tang |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Low Rank Matrix Factorization Algorithm Based on Multi-Graph Regularization for Detecting Drug-Disease AssociationabstractDetecting potential associations between drugs and diseases plays an indispensable role in drug development, which has also become a research hotspot in recent years. Compared with traditional methods, some computational approaches have the advantages of fast speed and low cost, which greatly accelerate the progress of predicting the drug-disease association. In this study, we propose a novel similarity-based method of low-rank matrix decomposition based on multi-graph regularization. On the basis of low-rank matrix factorization with$L_{2}$regularization, the multi-graph regularization constraint is constructed by combining a variety of similarity matrices from drugs and diseases respectively. In the experiments, we analyze the difference in the combination of different similarities, resulting that combining all the similarity information on drug space is unnecessary, and only a part of the similarity information can achieve the desired performance. Then our method is compared with other existing models on three data sets (Fdataset, Cdataset and LRSSLdataset) and have a good advantage in the evaluation measurement of AUPR. Besides, a case study experiment is conducted and showing that the superior ability for predicting the potential disease-related drugs of our model. Finally, we compare our model with some methods on six real world datasets, and our model has a good performance in detecting real world data. Chengwei Ai, Hongpeng Yang, Yijie Ding, Jijun Tang, Fei Guo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2023 | Laplacian Regularized Sparse Representation Based Classifier for Identifying DNA N4-Methylcytosine Sites via $L_{2,1/2}$L2,1/2-Matrix NormabstractN4-methylcytosine (4mC) is one of important epigenetic modifications in DNA sequences. Detecting 4mC sites is time-consuming. The computational method based on machine learning has provided effective help for identifying 4mC. To further improve the performance of prediction, we propose a Laplacian Regularized Sparse Representation based Classifier with L2,1/2-matrix norm (LapRSRC). We also utilize kernal trick to derive the kernel LapRSRC for nonlinear modeling. Matrix factorization technology is employed to solve the sparse representation coefficients of all test samples in the training set. And an efficient iterative algorithm is proposed to solve the objective function. We implement our model on six benchmark datasets of 4mC and eight UCI datasets to test evaluate performance. The results show that the performance of our method is better or comparable. Yijie Ding, Wenying He, Jijun Tang, Quan Zou 0001, Fei Guo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | BP-DDI: Drug-drug interaction prediction based on biological information and pharmacological textabstractIn the treatment of many diseases, combination drug therapy has been widely used and achieved good clinical efficacy. However, drug-drug interaction (DDI) may occur between multiple drugs and pose a huge threat to the health of patients. Therefore, predicting the presence or absence of DDI among multiple drugs is an important part of pharmacovigilance. Currently, various computational methods for DDI prediction usually use biological information such as molecular structures, targets and enzymes of drugs, or construct heterogeneous networks about drugs, diseases, and genes, so as to obtain abundant information related to drugs. In addition to biological data, pharmacology texts also contain a wealth of information about drug properties, but these texts have not yet been applied to DDI predictions. In this study, we first collect six types of pharmacology texts from DrugBank that can reflect properties of drugs, and propose a novel method named BP-DDI which can combine biological information and pharmacological text to realize DDI event prediction. BP-DDI first extracts biological features (chemical substructure features and target features) from biological data, and then extracts specific types of text features from the collected pharmacology text data. Finally, the biological features are fused with different types of pharmacological text features in order to predict DDI events. Our experiments demonstrate that BP-DDI outperforms existing methods on all three types of prediction tasks. BP-DDI achieves 0.9052 on ACC, and achieves 0.9612 on AUPR. Mingliang Dou, Genlang Chen, Fei Guo 0001, Jijun Tang |
BIBM | 5 |
| 2022 | Integrating Prior Knowledge with Graph Encoder for Gene Regulatory Inference from Single-cell RNA-Seq DataabstractInferring gene regulatory networks based on single-cell transcriptomes is critical for systematically understanding cell-specific regulatory networks and discovering drug targets in tumor cells. Here we show that existing methods mainly perform co-expression analysis and apply the image-based model to deal with the non-euclidean scRNA-seq data, which may not reasonably handle the dropout problem and not fully take advantage of the validated gene regulatory topology. We propose a graph-based end-to-end deep learning model for GRN inference (GRNInfer) with the help of known regulatory relations through transductive learning. The robustness and superiority of the model are demonstrated by comparative experiments. Jiawei Li 0018, Fan Yang 0081, Fang Wang 0028, Yu Rong 0001, Peilin Zhao, Shizhan Chen, Jianhua Yao 0001, Jijun Tang, Fei Guo 0001 |
BIBM | 8 |
| 2022 | Multi-scale Neighborhood Attention Transformer on U-Net for Medical Image SegmentationabstractU-shaped network structures with skip connections played an irreplaceable role in medical image analysis, but the limitation of convolution makes it unable to learn long-distance semantic information well. The recent success of Transformer in natural language processing and image classification shows that it can benefit from global information modeling by using self-attention mechanisms. However, both local and global features are equally important for dense prediction tasks. Transformer ignores local semantic information to a certain extent. In this study, we propose a Unet-like Transformer for medical image segmentation, named MN-Unet, which can simultaneously extract local and global features. MN-Unet consists of encoder, decoder, and skip connections. Specially, we design an encoder based on the Neighborhood Attention Transformer, which fuse three neighborhood sizes of different dimensions to simultaneously extract local and global features. In the decoder, we use bilinear interpolation to restore the image to its original size. Skip connection is added to alleviate the distortion of low resolution to high resolution. MN-Unet can achieve accurate segmentation of medical images without any pre-training. Extensive experimental results on two medical image datasets (LiTS 2017 and BraTS 2020) show that we achieve relatively better performance than state-of-the-art methods. The codes and trained models will be publicly available a https://github.com/hutchinsonian/MN_Unet Nanxing Zhang, Shiqiang Ma, Jijun Tang, Fei Guo 0001 |
BIBM | 5 |
| 2022 | Identification of protein-nucleotide binding residues via graph regularized k-local hyperplane distance nearest neighbor model
Yijie Ding, Jijun Tang, Fei Guo 0001 |
Appl. Intell. | 3 |
| 2022 | Identification of drug-target interactions via multiple kernel-based triple collaborative matrix factorizationabstractTargeted drugs have been applied to the treatment of cancer on a large scale, and some patients have certain therapeutic effects. It is a time-consuming task to detect drug-target interactions (DTIs) through biochemical experiments. At present, machine learning (ML) has been widely applied in large-scale drug screening. However, there are few methods for multiple information fusion. We propose a multiple kernel-based triple collaborative matrix factorization (MK-TCMF) method to predict DTIs. The multiple kernel matrices (contain chemical, biological and clinical information) are integrated via multi-kernel learning (MKL) algorithm. And the original adjacency matrix of DTIs could be decomposed into three matrices, including the latent feature matrix of the drug space, latent feature matrix of the target space and the bi-projection matrix (used to join the two feature spaces). To obtain better prediction performance, MKL algorithm can regulate the weight of each kernel matrix according to the prediction error. The weights of drug side-effects and target sequence are the highest. Compared with other computational methods, our model has better performance on four test data sets. Yijie Ding, Jijun Tang, Fei Guo 0001, Quan Zou 0001 |
Briefings Bioinform. | 2 |
| 2022 | Two-stage-vote ensemble framework based on integration of mutation data and gene interaction network for uncovering driver genesabstractIdentifying driver genes, exactly from massive genes with mutations, promotes accurate diagnosis and treatment of cancer. In recent years, a lot of works about uncovering driver genes based on integration of mutation data and gene interaction networks is gaining more attention. However, it is in suspense if it is more effective for prioritizing driver genes when integrating various types of mutation information (frequency and functional impact) and gene networks. Hence, we build a two-stage-vote ensemble framework based on somatic mutations and mutual interactions. Specifically, we first represent and combine various kinds of mutation information, which are propagated through networks by an improved iterative framework. The first vote is conducted on iteration results by voting methods, and the second vote is performed to get ensemble results of the first poll for the final driver gene list. Compared with four excellent previous approaches, our method has better performance in identifying driver genes on $33$ types of cancer from The Cancer Genome Atlas. Meanwhile, we also conduct a comparative analysis about two kinds of mutation information, five gene interaction networks and four voting strategies. Our framework offers a new view for data integration and promotes more latent cancer genes to be admitted. Yingxin Kan, Limin Jiang, Jijun Tang, Fei Guo 0001 |
Briefings Bioinform. | 4 |
| 2022 | A hybrid deep learning framework for gene regulatory network inference from single-cell transcriptomic dataabstractInferring gene regulatory networks (GRNs) based on gene expression profiles is able to provide an insight into a number of cellular phenotypes from the genomic level and reveal the essential laws underlying various life phenomena. Different from the bulk expression data, single-cell transcriptomic data embody cell-to-cell variance and diverse biological information, such as tissue characteristics, transformation of cell types, etc. Inferring GRNs based on such data offers unprecedented advantages for making a profound study of cell phenotypes, revealing gene functions and exploring potential interactions. However, the high sparsity, noise and dropout events of single-cell transcriptomic data pose new challenges for regulation identification. We develop a hybrid deep learning framework for GRN inference from single-cell transcriptomic data, DGRNS, which encodes the raw data and fuses recurrent neural network and convolutional neural network (CNN) to train a model capable of distinguishing related gene pairs from unrelated gene pairs. To overcome the limitations of such datasets, it applies sliding windows to extract valuable features while preserving the direction of regulation. DGRNS is constructed as a deep learning model containing gated recurrent unit network for exploring time-dependent information and CNN for learning spatially related information. Our comprehensive and detailed comparative analysis on the dataset of mouse hematopoietic stem cells illustrates that DGRNS outperforms state-of-the-art methods. The networks inferred by DGRNS are about 16% higher than the area under the receiver operating characteristic curve of other unsupervised methods and 10% higher than the area under the precision recall curve of other supervised methods. Experiments on human datasets show the strong robustness and excellent generalization of DGRNS. By comparing the predictions with standard network, we discover a series of novel interactions which are proved to be true in some specific cell types. Importantly, DGRNS identifies a series of regulatory relationships with high confidence and functional consistency, which have not yet been experimentally confirmed and merit further research. Wenying He, Jijun Tang, Quan Zou 0001, Fei Guo 0001 |
Briefings Bioinform. | 3 |
| 2022 | Inferring gene regulatory network via fusing gene expression image and RNA-seq dataabstractMOTIVATION: Recently, with the development of high-throughput experimental technology, reconstruction of gene regulatory network (GRN) has ushered in new opportunities and challenges. Some previous methods mainly extract gene expression information based on RNA-seq data, but the associated information is very limited. With the establishment of gene expression image database, it is possible to infer GRN from image data with rich spatial information. RESULTS: First, we propose a new convolutional neural network (called SDINet), which can extract gene expression information from images and identify the interaction between genes. SDINet can obtain the detailed information and high-level semantic information from the images well. And it can achieve satisfying performance on image data (Acc: 0.7196, F1: 0.7374). Second, we apply the idea of our SDINet to build an RNA-model, which also achieves good results on RNA-seq data (Acc: 0.8962, F1: 0.8950). Finally, we combine image data and RNA-seq data, and design a new fusion network to explore the potential relationship between them. Experiments show that our proposed network fusing two modalities can obtain satisfying performance (Acc: 0.9116, F1: 0.9118) than any single data. AVAILABILITY AND IMPLEMENTATION: Data and code are available from https://github.com/guofei-tju/Combine-Gene-Expression-images-and-RNA-seq-data-For-infering-GRN. Shiqiang Ma, Jin Liu 0012, Jijun Tang, Fei Guo 0001 |
Bioinform. | 4 |
| 2022 | A multi-layer multi-kernel neural network for determining associations between non-coding RNAs and diseases
Chengwei Ai, Hongpeng Yang, Yijie Ding, Jijun Tang, Fei Guo 0001 |
Neurocomputing | 4 |
| 2022 | Inferring human microbe-drug associations via multiple kernel fusion on graph neural network
Hongpeng Yang, Yijie Ding, Jijun Tang, Fei Guo 0001 |
Knowl. Based Syst. | 3 |
| 2022 | Res2Unet: A multi-scale channel attention network for retinal vessel segmentation
Jiaqi Ding, Jijun Tang, Fei Guo 0001 |
Neural Comput. Appl. | 3 |
| 2022 | DeepFusionDTA: Drug-Target Binding Affinity Prediction With Information Fusion and Hybrid Deep-Learning Ensemble ModelabstractIdentification of drug-target interaction (DTI) is the most important issue in the broad field of drug discovery. Using purely biological experiments to verify drug-target binding profiles takes lots of time and effort, so computational technologies for this task obviously have great benefits in reducing the drug search space. Most of computational methods to predict DTI are proposed to solve a binary classification problem, which ignore the influence of binding strength. Therefore, drug-target binding affinity prediction is still a challenging issue. Currently, lots of studies only extract sequence information that lacks feature-rich representation, but we consider more spatial features in order to merge various data in drug and target spaces. In this study, we propose a two-stage deep neural network ensemble model for detecting drug-target binding affinity, called DeepFusionDTA, via various information analysis modules. First stage is to utilize sequence and structure information to generate fusion feature map of candidate protein and drug pair through various analysis modules based deep learning. Second stage is to apply bagging-based ensemble learning strategy for regression prediction, and we obtain outstanding results by combining the advantages of various algorithms in efficient feature abstraction and regression calculation. Importantly, we evaluate our novel method, DeepFusionDTA, which delivers 1.5 percent CI increase on KIBA dataset and 1.0 percent increase on Davis dataset, by comparing with existing prediction tools, DeepDTA. Furthermore, the ideas we have offered can be applied to in-silico screening of the interaction space, to provide novel DTIs which can be experimentally pursued. The codes and data are available from https://github.com/guofei-tju/DeepFusionDTA. Yuqian Pu, Jiawei Li 0018, Jijun Tang, Fei Guo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | Identify ncRNA Subcellular Localization via Graph Regularized $k$k-Local Hyperplane Distance Nearest Neighbor Model on Multi-Kernel LearningabstractNon-coding RNAs (ncRNAs) are a type of RNAs which are not used to encode protein sequences. Emerging evidence shows that lots of ncRNAs may participate in many biological processes and must be widely involved in many types of cancers. Therefore, understanding their functionality is of great importance. Similar to proteins, various functions of ncRNAs relies on their subcellular localizations. Traditional high-throughput methods in wet-lab to identify subcellular localization is time-consuming and costly. In this paper, we propose a novel computational method based on multi-kernel learning to identify multi-label ncRNA subcellular localizations, via graph regularized k-local hyperplane distance nearest neighbor algorithm. First, we construct six types of sequence-based feature descriptors and select important feature vectors. Then, we build a multi-kernel learning model with Hilbert-Schmidt independence criterion (HSIC) to obtain optimal weights for vairous features. Furthermore, we propose the graph regularized k-local hyperplane distance nearest neighbor algorithm (GHKNN) as a binary classification model for detecting one kind of non-coding RNA subcellular localization. Finally, we apply One-vs-Rest strategy to decompose multi-label problem of non-coding RNA subcellular localizations. Our method achieves excellent performance on three ncRNA datasets and three human ncRNA datasets, and out-performs other outstanding machine learning methods. Comparing to existing method, our model also performs well especially on small datasets. We expect that this model will be useful for the prediction of subcellular localization and the study of important functional mechanisms of ncRNAs. Furthermore, we establish user-friendly web server (http://ncrna.lbci.net/) with the implementation of our method, which can be easily used by most experimental scientists. Haohao Zhou, Jijun Tang, Yijie Ding, Fei Guo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2021 | Document-level DDI relation extraction with document-entity embeddingabstractDDI is an important part of drug-related research and pharmacovigilance. Extracting DDI information from scientific literature has become a low-cost and highly reliable way. Currently, existing works are all sentence-level DDI relation extraction. In fact, the entity relationship is often expressed by multiple sentences. Moreover, the sentence-level DDI relation extraction also causes a large amount of redundancy in the whole dataset with increasing in negative instance data. In this study, we propose a document-level DDI relation extraction method based on document-entity embedding. Our method performs special processing on the DDI Extraction 2013 for the first time, in order to calculate document-level relation extraction. For obtaining document-level entity information, we propose a document-entity embedding method to integrate the information of all same drugs in the same article. The experimental results show that the processing of DDI Extraction 2013 dataset is reasonable. In addition, the proposed method has achieved good performance on document-level DDI dataset, and the best F1 score is 62.51%. This is the first time that DDI Extraction 2013 has been processed into a document-level dataset, and document-level DDI relation extraction has been realized. Mingliang Dou, Jijun Tang, Fei Guo 0001 |
BIBM | 2 |
| 2021 | MIASNet: A medical image segmentation method predicting future based on past and current casesabstractFast and accurate segmentation of medical images is essential for the diagnosis and treatment of diseases. The automatic segmentation technology based on deep learning has achieved encouraging performance in segmentation accuracy. However, the improvement of segmentation accuracy usually requires a larger network structure, which also leads to a decrease in segmentation speed. In this study, we propose a medical image anticipation segmentation net (MIASNet), in order to further improve the segmentation speed under the premise of excellent segmentation accuracy. For 3D medical images, we use the spatial association of the previous frame and the current frame as input data to predict the segmentation results of the next frame. Our approach consists of three lightweight sub-networks, which are used to learn the mapping relationship of the spatial domain. In order to make full use of the generating ability of the deep learning network, we use group convolution to obtain diversified prediction results. On the multimodal brain tumor image segmentation (BraTS) 2020 dataset, MIASNet achieves excellent segmentation accuracy without using the target frame that need to be segmented as the network input. Therefore, our proposed segmentation network can be used in a wider range of real-time medical applications. Shiqiang Ma, Jijun Tang, Fei Guo 0001 |
BIBM | 3 |
| 2021 | GEU-Net: Rethinking the information transmission in the skip connection of U-Net architectureabstractWith the wide application of deep learning technology in medical image processing, the performance of medical image segmentation has been improved in a breakthrough. U-Net architecture has excellent performance in medical image segmentation tasks. In order to solve the problem of image signal loss caused by the autoencoder structure, U-Net has added skip connections to its network to transfer the low-level features of the encoder path to the decoder path. Although this method can roughly solve the problem of image information loss, while it introduces a new problem, that is, the simple feature fusion method causes the high-level semantic information to be diluted. In order to solve the problem that the simple fusion of low-level edge information and high-level semantic information creates the semantic gap and dilutes high-level semantic information, we propose a novel U-shaped architecture, namely GEU-Net. GEU-Net utilizes ensemble learning methods to obtain better segmentation performance with a small computational cost. In addition, We propose a multi-scale group convolution block namely Group Residual (GR) module to reduce the semantic gap between encoder and decoder. We have evaluated our model on the BraTS 2020 Challenge, and have achieved competitive segmentation results. Shiqiang Ma, Jijun Tang, Fei Guo 0001 |
BIBM | 4 |
| 2021 | Multi-AMP: detecting the antimicrobial peptides and their activities using the multi-task learningabstractRecently due to the broad-spectrum and high-efficiency antibacterial activity, antimicrobial peptides (AMPs) have become the best alternative to antibiotics. With the rapid increase of the antibacterial peptides, many computational methods have been developed to identify the AMPs and their specific antibacterial activities. However, most existing methods regard these two problems as independent sub-problems and ignore the correlation between tasks. In this paper, we propose a method, Multi-AMP, which utilizes multi-task learning and solves two tasks simultaneously: 1) whether a given peptide is AMP, 2) which activities it performs. The two tasks share the parameters at the bottom layers of the model and learn the specific information at the top layers. Experiments indicate that our multi-task model performs better than single-task models and two existing predictors, which can give insights to the drug discovery process. Qiaozhen Meng, Jijun Tang, Fei Guo 0001 |
BIBM | 2 |
| 2021 | Membrane Protein Identification via Multi-view Graph Regularized k-Local Hyperplane Distance Nearest Neighbor ModelabstractX-ray diffraction and nuclear magnetic resonance spectroscopy are the main methods for measuring membrane proteins. The traditional methods are time-consuming and labor-intensive. To large-scale prediction and screening of membrane proteins, a graph regularized k-local hyperplane distance nearest neighbor model (GHKNN) is proposed to identify of membrane protein types. For effectively integrating features, multi-view learning (MVL) is employed to estimate the weight of each graph. We test GHKNN on 2 data sets of membrane protein. Compared with other methods, the accuracy of GHKNN is better or comparable. Mengwei Sun, Yuqing Qian, Yijie Ding, Jijun Tang, Quan Zou 0001 |
BIBM | 4 |
| 2021 | RCGA-Net: An Improved Multi-hybrid Attention Mechanism Network in Biomedical Image SegmentationabstractDrawing support from an effective Medical Image Segmentation (MIS) is conducive to a substantial diagnostic basis for the physicians to identify the focus lesion in the patient body and give the subsequent clinical assessment of the patient status. Although various works have tried the challenging quantitative analysis problem, it is still difficult to conduct precise automatic segmentation, especially the soft tissue organs. In this decade, with the increased amount of available datasets, deep learning-based networks have achieved remarkable performance in image processing. Inspired by the state-of-the-art deep learning works, in this paper, we propose an end-to-end multi-layer network named RCGA-Net. It consists of an encoder-decoder backbone that integrates a coordinate attention mechanism based on space and channel and a global context extraction module to highlight more valuable information. To evaluate the performance of RCGA-Net, we apply it to different kinds of clinical and experimental MIS tasks to testify its generalization ability. Extensive experiments represent that our schema has taken the outperform or compatible results among the comparison methods group. Specifically, the numeric result of RCGA-Net on the pulmonary dataset has achieved a 99.12% optimum F1-score. Feng Xiao 0005, Shengyong Chen, Zhijun Liao, Jijun Tang |
BIBM | 7 |
| 2021 | A Zero-Shot Method for 3D Medical Image SegmentationabstractAccurate automatic medical image segmentation technology plays an important role for the diagnosis and treatment of brain tumor. However, existing methods based on outstanding 2.5D and 3D segmentation strategies are time-consumption and hardware-consumption while ensuring high accuracy. In order to reduce the high demand for automatic segmentation of tumor images and avoid the noise interference in a single input image, we propose an end-to-end zero-shot CNN segmentation method. Our method only utilizes two adjacent images, instead of the target image, as the input data of deep neural network to predict the brain tumor area in the target image. Avoiding noise interference in the target image, this method makes full use of the spatial context feature between adjacent slices in order to obtain accurate zero-shot segmentation results. We compare with the state-of-the-art segmentation frameworks on the same benchmark and notice that our method has strong competitiveness. Shiqiang Ma, Jijun Tang, Fei Guo 0001 |
ICME | 3 |
| 2021 | MMFGRN: a multi-source multi-model fusion method for gene regulatory network reconstructionabstractLots of biological processes are controlled by gene regulatory networks (GRNs), such as growth and differentiation of cells, occurrence and development of the diseases. Therefore, it is important to persistently concentrate on the research of GRN. The determination of the gene-gene relationships from gene expression data is a complex issue. Since it is difficult to efficiently obtain the regularity behind the gene-gene relationship by only relying on biochemical experimental methods, thus various computational methods have been used to construct GRNs, and some achievements have been made. In this paper, we propose a novel method MMFGRN (for "Multi-source Multi-model Fusion for Gene Regulatory Network reconstruction") to reconstruct the GRN. In order to make full use of the limited datasets and explore the potential regulatory relationships contained in different data types, we construct the MMFGRN model from three perspectives: single time series data model, single steady-data model and time series and steady-data joint model. And, we utilize the weighted fusion strategy to get the final global regulatory link ranking. Finally, MMFGRN model yields the best performance on the DREAM4 InSilico_Size10 data, outperforming other popular inference algorithms, with an overall area under receiver operating characteristic score of 0.909 and area under precision-recall (AUPR) curves score of 0.770 on the 10-gene network. Additionally, as the network scale increases, our method also has certain advantages with an overall AUPR score of 0.335 on the DREAM4 InSilico_Size100 data. These results demonstrate the good robustness of MMFGRN on different scales of networks. At the same time, the integration strategy proposed in this paper provides a new idea for the reconstruction of the biological network model without prior knowledge, which can help researchers to decipher the elusive mechanism of life. Wenying He, Jijun Tang, Quan Zou 0001, Fei Guo 0001 |
Briefings Bioinform. | 2 |
| 2021 | Predicting MHC class I binder: existing approaches and a novel recurrent neural network solutionabstractMajor histocompatibility complex (MHC) possesses important research value in the treatment of complex human diseases. A plethora of computational tools has been developed to predict MHC class I binders. Here, we comprehensively reviewed 27 up-to-date MHC I binding prediction tools developed over the last decade, thoroughly evaluating feature representation methods, prediction algorithms and model training strategies on a benchmark dataset from Immune Epitope Database. A common limitation was identified during the review that all existing tools can only handle a fixed peptide sequence length. To overcome this limitation, we developed a bilateral and variable long short-term memory (BVLSTM)-based approach, named BVLSTM-MHC. It is the first variable-length MHC class I binding predictor. In comparison to the 10 mainstream prediction tools on an independent validation dataset, BVLSTM-MHC achieved the best performance in six out of eight evaluated metrics. A web server based on the BVLSTM-MHC model was developed to enable accurate and efficient MHC class I binder prediction in human, mouse, macaque and chimpanzee. Limin Jiang, Jiawei Li 0018, Jijun Tang, Fei Guo 0001 |
Briefings Bioinform. | 4 |
| 2021 | DeepATT: a hybrid category attention neural network for identifying functional effects of DNA sequencesabstractQuantifying DNA properties is a challenging task in the broad field of human genomics. Since the vast majority of non-coding DNA is still poorly understood in terms of function, this task is particularly important to have enormous benefit for biology research. Various DNA sequences should have a great variety of representations, and specific functions may focus on corresponding features in the front part of learning model. Currently, however, for multi-class prediction of non-coding DNA regulatory functions, most powerful predictive models do not have appropriate feature extraction and selection approaches for specific functional effects, so that it is difficult to gain a better insight into their internal correlations. Hence, we design a category attention layer and category dense layer in order to select efficient features and distinguish different DNA functions. In this study, we propose a hybrid deep neural network method, called DeepATT, for identifying $919$ regulatory functions on nearly $5$ million DNA sequences. Our model has four built-in neural network constructions: convolution layer captures regulatory motifs, recurrent layer captures a regulatory grammar, category attention layer selects corresponding valid features for different functions and category dense layer classifies predictive labels with selected features of regulatory functions. Importantly, we compare our novel method, DeepATT, with existing outstanding prediction tools, DeepSEA and DanQ. DeepATT performs significantly better than other existing tools for identifying DNA functions, at least increasing $1.6\%$ area under precision recall. Furthermore, we can mine the important correlation among different DNA functions according to the category attention module. Moreover, our novel model can greatly reduce the number of parameters by the mechanism of attention and locally connected, on the basis of ensuring accuracy. Jiawei Li 0018, Yuqian Pu, Jijun Tang, Quan Zou 0001, Fei Guo 0001 |
Briefings Bioinform. | 3 |
| 2021 | Exploring associations of non-coding RNAs in human diseases via three-matrix factorization with hypergraph-regular terms on center kernel alignmentabstractRelationship of accurate associations between non-coding RNAs and diseases could be of great help in the treatment of human biomedical research. However, the traditional technology is only applied on one type of non-coding RNA or a specific disease, and the experimental method is time-consuming and expensive. More computational tools have been proposed to detect new associations based on known ncRNA and disease information. Due to the ncRNAs (circRNAs, miRNAs and lncRNAs) having a close relationship with the progression of various human diseases, it is critical for developing effective computational predictors for ncRNA-disease association prediction. In this paper, we propose a new computational method of three-matrix factorization with hypergraph regularization terms (HGRTMF) based on central kernel alignment (CKA), for identifying general ncRNA-disease associations. In the process of constructing the similarity matrix, various types of similarity matrices are applicable to circRNAs, miRNAs and lncRNAs. Our method achieves excellent performance on five datasets, involving three types of ncRNAs. In the test, we obtain best area under the curve scores of $0.9832$, $0.9775$, $0.9023$, $0.8809$ and $0.9185$ via 5-fold cross-validation and $0.9832$, $0.9836$, $0.9198$, $0.9459$ and $0.9275$ via leave-one-out cross-validation on five datasets. Furthermore, our novel method (CKA-HGRTMF) is also able to discover new associations between ncRNAs and diseases accurately. Availability: Codes and data are available: https://github.com/hzwh6910/ncRNA2Disease.git. Contact:[email protected]. Jijun Tang, Yijie Ding, Fei Guo 0001 |
Briefings Bioinform. | 2 |
| 2021 | Exploring effectiveness of ab-initio protein-protein docking methods on a novel antibacterial protein complex datasetabstractDiseases caused by bacterial infections become a critical problem in public heath. Antibiotic, the traditional treatment, gradually loses their effectiveness due to the resistance. Meanwhile, antibacterial proteins attract more attention because of broad spectrum and little harm to host cells. Therefore, exploring new effective antibacterial proteins is urgent and necessary. In this paper, we are committed to evaluating the effectiveness of ab-initio docking methods in antibacterial protein-protein docking. For this purpose, we constructed a three-dimensional (3D) structure dataset of antibacterial protein complex, called APCset, which contained $19$ protein complexes whose receptors or ligands are homologous to antibacterial peptides from Antimicrobial Peptide Database. Then we selected five representative ab-initio protein-protein docking tools including ZDOCK3.0.2, FRODOCK3.0, ATTRACT, PatchDock and Rosetta to identify these complexes' structure, whose performance differences were obtained by analyzing from five aspects, including top/best pose, first hit, success rate, average hit count and running time. Finally, according to different requirements, we assessed and recommended relatively efficient protein-protein docking tools. In terms of computational efficiency and performance, ZDOCK was more suitable as preferred computational tool, with average running time of $6.144$ minutes, average Fnat of best pose of $0.953$ and average rank of best pose of $4.158$. Meanwhile, ZDOCK still yielded better performance on Benchmark 5.0, which proved ZDOCK was effective in performing docking on large-scale dataset. Our survey can offer insights into the research on the treatment of bacterial infections by utilizing the appropriate docking methods. Qiaozhen Meng, Jijun Tang, Fei Guo 0001 |
Briefings Bioinform. | 3 |
| 2021 | A comprehensive overview and critical evaluation of gene regulatory network inference technologiesabstractGene regulatory network (GRN) is the important mechanism of maintaining life process, controlling biochemical reaction and regulating compound level, which plays an important role in various organisms and systems. Reconstructing GRN can help us to understand the molecular mechanism of organisms and to reveal the essential rules of a large number of biological processes and reactions in organisms. Various outstanding network reconstruction algorithms use specific assumptions that affect prediction accuracy, in order to deal with the uncertainty of processing. In order to study why a certain method is more suitable for specific research problem or experimental data, we conduct research from model-based, information-based and machine learning-based method classifications. There are obviously different types of computational tools that can be generated to distinguish GRNs. Furthermore, we discuss several classical, representative and latest methods in each category to analyze core ideas, general steps, characteristics, etc. We compare the performance of state-of-the-art GRN reconstruction technologies on simulated networks and real networks under different scaling conditions. Through standardized performance metrics and common benchmarks, we quantitatively evaluate the stability of various methods and the sensitivity of the same algorithm applying to different scaling networks. The aim of this study is to explore the most appropriate method for a specific GRN, which helps biologists and medical scientists in discovering potential drug targets and identifying cancer biomarkers. Wenying He, Jijun Tang, Quan Zou 0001, Fei Guo 0001 |
Briefings Bioinform. | 3 |
| 2021 | Predicting subcellular location of protein with evolution information and sequence-based deep learningabstractBACKGROUND: Protein subcellular localization prediction plays an important role in biology research. Since traditional methods are laborious and time-consuming, many machine learning-based prediction methods have been proposed. However, most of the proposed methods ignore the evolution information of proteins. In order to improve the prediction accuracy, we present a deep learning-based method to predict protein subcellular locations. RESULTS: Our method utilizes not only amino acid compositions sequence but also evolution matrices of proteins. Our method uses a bidirectional long short-term memory network that processes the entire protein sequence and a convolutional neural network that extracts features from protein sequences. The position specific scoring matrix is used as a supplement to protein sequences. Our method was trained and tested on two benchmark datasets. The experiment results show that our method yields accurate results on the two datasets with an average precision of 0.7901, ranking loss of 0.0758 and coverage of 1.2848. CONCLUSION: The experiment results show that our method outperforms five methods currently available. According to those experiments, we can see that our method is an acceptable alternative to predict protein subcellular location. Zhijun Liao, Gaofeng Pan, Jijun Tang |
BMC Bioinform. | 4 |
| 2021 | A sequence-based multiple kernel model for identifying DNA-binding proteinsabstractBACKGROUND: DNA-Binding Proteins (DBP) plays a pivotal role in biological system. A mounting number of researchers are studying the mechanism and detection methods. To detect DBP, the tradition experimental method is time-consuming and resource-consuming. In recent years, Machine Learning methods have been used to detect DBP. However, it is difficult to adequately describe the information of proteins in predicting DNA-binding proteins. In this study, we extract six features from protein sequence and use Multiple Kernel Learning-based on Centered Kernel Alignment to integrate these features. The integrated feature is fed into Support Vector Machine to build predictive model and detect new DBP. RESULTS: In our work, date sets of PDB1075 and PDB186 are employed to test our method. From the results, our model obtains better results (accuracy) than other existing methods on PDB1075 ([Formula: see text]) and PDB186 ([Formula: see text]), respectively. CONCLUSION: Multiple kernel learning could fuse the complementary information between different features. Compared with existing methods, our method achieves comparable and best results on benchmark data sets. Yuqing Qian, Limin Jiang, Yijie Ding, Jijun Tang, Fei Guo 0001 |
BMC Bioinform. | 4 |
| 2021 | Identification of drug-target interactions via multi-view graph regularized link propagation model
Yijie Ding, Jijun Tang, Fei Guo 0001 |
Neurocomputing | 2 |
| 2021 | Granular multiple kernel learning for identifying RNA-binding protein residues via integrating sequence and structure information
Yijie Ding, Qiaozhen Meng, Jijun Tang, Fei Guo 0001 |
Neural Comput. Appl. | 4 |
| 2021 | Protein Crystallization Identification via Fuzzy Model on Linear Neighborhood RepresentationabstractX-ray crystallography is the most popular approach for analyzing protein 3D structure. However, the success rate of protein crystallization is very low (2-10 percent). To reduce the cost of time and resources, lots of computation-based methods are developed to detect the protein crystallization. Improving the accuracy of predicting protein crystallization is very important for the determination of protein structure by X-ray crystallography. At present, many machine learning methods are used to predict protein crystallization. In this article, we propose a Fuzzy Support Vector Machine based on Linear Neighborhood Representation (FSVM-LNR) to predict the crystallization propensity of proteins. Proteins are represented by three types of features (PsePSSM, PSSM-DWT, MMI-PS), and these features are serially combined and fed into FSVM-LNR. FSVM-LNR can filter outliers by membership score, which is calculated via reconstruction residuals of k nearest samples. To evaluate the performance of our predictive model, we test FSVM-LNR on the datasets of TRAIN3587, TEST3585 and TEST500. Our method achieves better Mathew's correlation coefficient (MCC) on TRAIN3587 (MCC: 0.56) and TEST3585 (MCC: 0.58). Although the performance of independent test is not the best on TEST500, FSVM-LNR also has a certain predictability (MCC: 0.70) in the identification of protein crystallization. The good performance on the datasets proves the effectiveness of our method and the better performance on large datasets further demonstrates the stability and superiority of our method. Yijie Ding, Jijun Tang, Fei Guo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | CrystalM: A Multi-View Fusion Approach for Protein Crystallization PredictionabstractImproving the accuracy of predicting protein crystallization is very important for protein crystallization projects, which is a critical step for the determination of protein structure by X-ray crystallography. At present, many machine learning methods are used to predict protein crystallization. Here, we use a novel feature combination to construct a SVM model in the prediction of protein crystallization, called as CrystalM. In this work, we extract six features to represent protein sequences, namely Average Block-Position specific scoring matrix (AVBlock-PSSM), Average Block-Secondary Structure (AVBlock-SS), Global Encoding (GE), Pseudo-Position specific scoring matrix (PsePSSM), Protscale, and Discrete Wavelet Transform-Position specific scoring matrix (DWT-PSSM). Moreover, we employ two training datasets (TRAIN3587 and TRAIN1500) and their corresponding independent test datasets (TEST3585 and TEST500) to evaluate CrystalM by feeding multi-view features into Support Vector Machine (SVM) classifier. Two training datasets are employed for five-fold cross validation, and two test datasets are separately used to test the corresponding datasets. Finally, we compare CrystalM with other existing methods in the performance. For the datasets of TRAIN3587 and TEST3585, CrystalM achieves best Accuracy (ACC), best Specificity (SP), and the same Mathew's correlation coefficient (MCC) as the previous outperforming methods in the five-fold cross validation. In particular, ACC, SP, and MCC have surpassed the existing methods in independent test, which proves the effectiveness of CrystalM. Meanwhile, ACC, SP, and MCC are higher than existing methods in the five-fold cross validation for TRAIN1500. Although the performance of independent test for TEST500 is not the best, CrystalM also has a certain predictability in the prediction of protein crystallization. In addition, we find that only choosing the first four features can improve the performance of prediction for TRAIN1500 and TEST500, not only in independent tests but also in five-fold cross validation. This phenomenon indicates that the latter two features can not effectively represent proteins of TRAIN1500 and TEST500. CrystalM is a sequence-based protein crystallization prediction method. The good performance on the datasets proves the effectiveness of CrystalM and the better performance on large datasets further demonstrates the stability and superiority of CrystalM. Yijie Ding, Jijun Tang, Yu Dai 0005, Fei Guo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2021 | AIEpred: An Ensemble Predictive Model of Classifier Chain to Identify Anti-Inflammatory PeptidesabstractAnti-inflammatory peptides (AIEs) have recently emerged as promising therapeutic agent for treatment of various inflammatory diseases, such as rheumatoid arthritis and Alzheimer's disease. Therefore, detecting the correlation between amino acid sequence and its anti-inflammatory property is of great importance for the discovery of new AIEs. To address this issue, we propose a novel prediction tool for accurate identification of peptides as anti-inflammatory epitopes or non anti-inflammatory epitopes. Most of all, we encode the original peptide sequence for better mining and exploring the information and patterns, based on the three feature representations as amino acid contact, position specific scoring matrix, physicochemical property. At the same time, we exploit several feature extraction models and utilize one feature selection model, in order to construct many base classifiers from various feature representations. More specifically, we develop an effective classification model, with which we can extract and learn a set of informative features from the ensemble classifier chain model with different group of base classifiers. Furthermore, in order to test the predictive power of our model, we conduct the comparative experiments on the leave-one-out cross-validation and the independent test. It shows that our novel predictor performs great accurate for identification of AIEs as well as existing outstanding prediction tools. Source codes are available at https://github.com/guofei-tju/Ensemble-classifier-chain-model. Lianrong Pu, Jijun Tang, Fei Guo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2021 | Multi-Scale Time-Series Kernel-Based Learning Method for Brain Disease DiagnosisabstractThe functional magnetic resonance imaging (fMRI) is a noninvasive technique for studying brain activity, such as brain network analysis, neural disease automated diagnosis and so on. However, many existing methods have some drawbacks, such as limitations of graph theory, lack of global topology characteristic, local sensitivity of functional connectivity, and absence of temporal or context information. In addition to many numerical features, fMRI time series data also cover specific contextual knowledge and global fluctuation information. Here, we propose multi-scale time-series kernel-based learning model for brain disease diagnosis, based on Jensen-Shannon divergence. First, we calculate correlation value within and between brain regions over time. In addition, we extract multi-scale synergy expression probability distribution (interactional relation) between brain regions. Also, we produce state transition probability distribution (sequential relation) on single brain regions. Then, we build time-series kernel-based learning model based on Jensen-Shannon divergence to measure similarity of brain functional connectivity. Finally, we provide an efficient system to deal with brain network analysis and neural disease automated diagnosis. On Alzheimer's Disease Neuroimaging Initiative (ADNI) dataset, our proposed method achieves accuracy of 0.8994 and AUC of 0.8623. On Major Depressive Disorder (MDD) dataset, our proposed method achieves accuracy of 0.9166 and AUC of 0.9263. Experiments show that our proposed method outperforms other existing excellent neural disease automated diagnosis approaches. It shows that our novel prediction method performs great accurate for identification of brain diseases as well as existing outstanding prediction tools. Jiaqi Ding, Junhai Xu, Jijun Tang, Fei Guo 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2020 | Melanoma Classification in Dermoscopy Images via Ensemble Learning on Deep Neural NetworkabstractAuotmatic melanoma classification in dermoscopy images is a very important task, which can help improve diagnostic accuracy and reduce mortality. Deep convolutional neural network (DCNN) has developed rapidly in recent years, but it is still a challenging task due to the intra-class variation and inter-class similarity of melanoma. We proposed a novel neural network integration model, which is composed of three parts: First, we use U-net segmentation network to generate masks and use the masks to crop original images; Second, we use five state-of-the-art DCNNs to extract features of cropped images, and add the squeeze-excitation block (SE block) to emphasize useful features; Finally, we construct a new neural network with local connection to integrate the classification results, extract features of different class of results, and integrate the results of each class separately. Local connection can integrate each class separately, maximizing the advantages of different networks in various classes. We evaluate our model on ISIC 2017 challenge dataset, and the result shows that our method has better performance compared with the existing methods. Jiawei Li 0018, Shiqiang Ma, Jijun Tang, Fei Guo 0001 |
BIBM | 4 |
| 2020 | An two-layer predictive model of ensemble classifier chain for detecting antimicrobial peptidesabstractAntimicrobial peptides (AMPs) are innate immune molecules that exhibit activities against a range of microbes. According to their special functions, AMPs are generally classified into several categories. Over the last decade, a number of AMP prediction tools have been designed and made freely available online, which show potential to discriminate AMPs from non-AMPs. However, the relative quality of existing AMP predictions produced by various tools is difficult to quantify. In fact, a comprehensive benchmark dataset used to train the prediction model is one of key points to solving the problem. Also, how to address the multi-label character of new synthetic instance is obviously very important to both basic research and drug development. In view of this, AMPs prediction should be a task of two-level multi-label classification, in which the first step is to identify whether a query peptide is AMP, and the second step is to identify which functional type(s) the peptide belongs to. To establish a really useful prediction method, we construct a valid benchmark dataset to train the predictor, and develop a powerful algorithm to operate the prediction. In this paper, we propose a novel two-layer prediction model for identifying AMP and its functional types, using ADASYN oversampling technology to solve imbalance multi-label classification problem. First, we construct a novel benchmark AMPs dataset with seven different AMP functional types. Then, we encode AMPs within three different feature representations, and use various feature extraction models to convert the variable length coding matrix into some equidimensional features. Furthermore, we use modelbased feature selection method for filtering effective and sparse features. Finally, we apply ensemble classifier chain model to identify whether a query peptide is an AMPs or non-AMPs. In the second layer prediction, we use ADASYN to oversample different functional types of AMPs, and build a multi-label multi-class prediction model to identify which functional type(s) it belongs to. To be specific, our novel method outperforms outstanding rather than other tools in most respects on our novel benchmark datasets. Our novel benchmark dataset and source codes are available at https://github.com/guofei-tju/Two_Level_Ensemble-classifier-chain. Yijie Ding, Jijun Tang, Fei Guo 0001 |
BIBM | 4 |
| 2020 | Critical evaluation of web-based prediction tools for human protein subcellular localizationabstractHuman protein subcellular localization has an important research value in biological processes, also in elucidating protein functions and identifying drug targets. Over the past decade, a number of protein subcellular localization prediction tools have been designed and made freely available online. The purpose of this paper is to summarize the progress of research on the subcellular localization of human proteins in recent years, including commonly used data sets proposed by the predecessors and the performance of all selected prediction tools against the same benchmark data set. We carry out a systematic evaluation of several publicly available subcellular localization prediction methods on various benchmark data sets. Among them, we find that mLASSO-Hum and pLoc-mHum provide a statistically significant improvement in performance, as measured by the value of accuracy, relative to the other methods. Meanwhile, we build a new data set using the latest version of Uniprot database and construct a new GO-based prediction method HumLoc-LBCI in this paper. Then, we test all selected prediction tools on the new data set. Finally, we discuss the possible development directions of human protein subcellular localization. Availability: The codes and data are available from http://www.lbci.cn/syn/. Yinan Shen, Yijie Ding, Jijun Tang, Quan Zou 0001, Fei Guo 0001 |
Briefings Bioinform. | 3 |
| 2020 | Achieving large and distant ancestral genome inference by using an improved discrete quantum-behaved particle swarm optimization algorithmabstractBACKGROUND: Reconstructing ancestral genomes is one of the central problems presented in genome rearrangement analysis since finding the most likely true ancestor is of significant importance in phylogenetic reconstruction. Large scale genome rearrangements can provide essential insights into evolutionary processes. However, when the genomes are large and distant, classical median solvers have failed to adequately address these challenges due to the exponential increase of the search space. Consequently, solving ancestral genome inference problems constitutes a task of paramount importance that continues to challenge the current methods used in this area, whose difficulty is further increased by the ongoing rapid accumulation of whole-genome data. RESULTS: In response to these challenges, we provide two contributions for ancestral genome inference. First, an improved discrete quantum-behaved particle swarm optimization algorithm (IDQPSO) by averaging two of the fitness values is proposed to address the discrete search space. Second, we incorporate DCJ sorting into the IDQPSO (IDQPSO-Median). In comparison with the other methods, when the genomes are large and distant, IDQPSO-Median has the lowest median score, the highest adjacency accuracy, and the closest distance to the true ancestor. In addition, we have integrated our IDQPSO-Median approach with the GRAPPA framework. Our experiments show that this new phylogenetic method is very accurate and effective by using IDQPSO-Median. CONCLUSIONS: Our experimental results demonstrate the advantages of IDQPSO-Median approach over the other methods when the genomes are large and distant. When our experimental results are evaluated in a comprehensive manner, it is clear that the IDQPSO-Median approach we propose achieves better scalability compared to existing algorithms. Moreover, our experimental results by using simulated and real datasets confirm that the IDQPSO-Median, when integrated with the GRAPPA framework, outperforms other heuristics in terms of accuracy, while also continuing to infer phylogenies that were equivalent or close to the true trees within 5 days of computation, which is far beyond the difficulty level that can be handled by GRAPPA. Zhaojuan Zhang, Wanliang Wang, Ruofan Xia, Gaofeng Pan, Jijun Tang |
BMC Bioinform. | 6 |
| 2020 | Identification of membrane protein types via multivariate information fusion with Hilbert-Schmidt Independence Criterion
Yijie Ding, Jijun Tang, Fei Guo 0001 |
Neurocomputing | 3 |
| 2020 | Identification of Drug-Target Interactions via Dual Laplacian Regularized Least Squares with Multiple Kernel Fusion
Yijie Ding, Jijun Tang, Fei Guo 0001 |
Knowl. Based Syst. | 2 |
| 2020 | Identification of drug-target interactions via fuzzy bipartite local model
Yijie Ding, Jijun Tang, Fei Guo 0001 |
Neural Comput. Appl. | 2 |
| 2020 | DeepAVP: A Dual-Channel Deep Neural Network for Identifying Variable-Length Antiviral PeptidesabstractAntiviral peptides (AVPs) have been experimentally verified to block virus into host cells, which have antiviral activity with decapeptide amide. Therefore, utilization of experimentally validated antiviral peptides is a potential alternative strategy for targeting medically important viruses. In this article, we propose a dual-channel deep neural network ensemble method for analyzing variable-length antiviral peptides. The LSTM channel can capture long-term dependencies for effectively studying original variable-length sequence data. The CONV channel can build dynamic neural network for analyzing the local evolution information. Also, our model can fine-tune the substitution matrix for specifically functional peptides. Applying it to a novel experimentally verified dataset, our AVPs predictor, DeepAVP, demonstrates state-of-the-art performance of [Formula: see text] accuracy and 0.85 MCC, which is far better than existing prediction methods for identifying antiviral peptides. Therefore, DeepAVP, web server for predicting the effective AVPs, would make significantly contributions to peptide-based antiviral research. Jiawei Li 0018, Yuqian Pu, Jijun Tang, Quan Zou 0001, Fei Guo 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2019 | Identification of DNA-Binding Proteins via Fuzzy Multiple Kernel Model and Sequence Information
Yijie Ding, Jijun Tang, Fei Guo 0001 |
ICIC (2) | 2 |
| 2019 | Identifying protein-protein interface via a novel multi-scale local sequence and structural representationabstractAbstract Background Protein-protein interaction plays a key role in a multitude of biological processes, such as signal transduction, de novo drug design, immune responses, and enzymatic activities. Gaining insights of various binding abilities can deepen our understanding of the interaction. It is of great interest to understand how proteins in a complex interact with each other. Many efficient methods have been developed for identifying protein-protein interface. Results In this paper, we obtain the local information on protein-protein interface, through multi-scale local average block and hexagon structure construction. Given a pair of proteins, we use a trained support vector regression (SVR) model to select best configurations. On Benchmark v4.0, our method achieves average Irmsd value of 3.28Å and overall Fnat value of 63%, which improves upon Irmsd of 3.89Å and Fnat of 49% for ZRANK, and Irmsd of 3.99Å and Fnat of 46% for ClusPro. On CAPRI targets, our method achieves average Irmsd value of 3.45Å and overall Fnat value of 46%, which improves upon Irmsd of 4.18Å and Fnat of 40% for ZRANK, and Irmsd of 5.12Å and Fnat of 32% for ClusPro. The success rates by our method, FRODOCK 2.0, InterEvDock and SnapDock on Benchmark v4.0 are 41.5%, 29.0%, 29.4% and 37.0%, respectively. Conclusion Experiments show that our method performs better than some state-of-the-art methods, based on the prediction quality improved in terms of CAPRI evaluation criteria. All these results demonstrate that our method is a valuable technological tool for identifying protein-protein interface. Fei Guo 0001, Quan Zou 0001, Jijun Tang, Junhai Xu |
BMC Bioinform. | 5 |
| 2019 | Identification of drug-side effect association via multiple information integration with centered kernel alignment
Yijie Ding, Jijun Tang, Fei Guo 0001 |
Neurocomputing | 2 |
| 2019 | Phylogenetic Reconstruction for Copy-Number Evolution ProblemsabstractCancer is known for its heterogeneity and is regarded as an evolutionary process driven by somatic mutations and clonal expansions. This evolutionary process can be modeled by a phylogenetic tree and phylogenetic analysis of multiple subclones of cancer cells can facilitate the study of the tumor variants progression. Copy-number aberration occurs frequently in many types of tumors in terms of segmental amplifications and deletions. In this paper, we developed a distance-based method for reconstructing phylogenies from copy-number profiles of cancer cells. We demonstrate the importance of distance correction from the edit (minimum) distance to the estimated actual number of events. Experimental results show that our approaches provide accurate and scalable results in estimating the actual number of evolutionary events between copy number profiles and in reconstructing phylogenies. Ruofan Xia, Yu Lin 0001, Tieming Geng, Bing Feng, Jijun Tang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2019 | Identification of Drug-Side Effect Association via Semisupervised Model and Multiple Kernel LearningabstractDrug-side effect association contains the information on marketed medicines and their recorded adverse drug reactions. Traditional experimental method is time consuming and expensive. All associations of drugs and side-effects are seen as a bipartite network. Therefore, many computational approaches have been developed to deal with this problem, which are used to predict new potential associations. However, lots of methods did not consider multiple kernel learning (MKL) algorithm, which can integrate multiple sources of information and further improve prediction performance. In this study, we develop a novel predictor of drug-side effect association. First, we build multiple kernels from drug space and side-effect space. What is more, these corresponding kernels are linear weighted by MKL algorithm in drug space and side-effect space, respectively. Finally, a graph-based semisupervised learning is employed to construct drug-side effect predictor. Compared with existing methods, our method achieves better results on three benchmark data sets. The values of area under the precision recall curve are 0.668, 0.673, and 0.670 on three benchmark data sets, respectively. Our method is a useful tool for the side-effects prediction of drugs. Yijie Ding, Jijun Tang, Fei Guo 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2018 | Prediction of human protein subcellular localization using deep learning
Leyi Wei, Yijie Ding, Ran Su, Jijun Tang, Quan Zou 0001 |
J. Parallel Distributed Comput. | 4 |
| 2018 | Guest Editorial for the 14th Asia Pacific Bioinformatics ConferenceabstractThe eight papers in this special section were presented at the 14th Asia Pacific Bioinformatics Conference (APBC2016), which was held in San Francisco, USA, 11-13 January 2016. Jijun Tang, Yi-Ping Phoebe Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2017 | Down syndrome prediction/screening model based on deep learning and illumina genotyping arrayabstractDown syndrome (DS) is a genetic disorder with genome dosage imbalances and micro-duplications of human chromosome 21. It is usually associated with a group of serious diseases, including intellectual disabilities, cardiac diseases, physical abnormalities, and other abnormalities. Currently, since there is no cure for human DS, screening and early detection have become the most efficient way for DS prevention. In this study, we used deep learning techniques to build accurate DS prediction/screening models based on the analysis of newly introduced Illumina genotyping array. Specifically, we built chromosome SNP maps based on clinical genotyping data collected by Vanderbilt University Medical Center. Then we proposed a convolutional neural network (CNN) architecture with ten layers and two merged CNN models, which took two input chromosome SNP maps in combination. Our CNN DS prediction/screening model achieved over 99.3% average accuracy, as well as very low false positive and false negative rate, which are critical to disease prediction and screening in medical practice. It also had better performances in terms of all evaluating metrics when compared with three conventional machine-learning algorithms. Finally, we visualized the feature maps and the trained filter weights from intermediate layers of our trained CNN model. We further discussed the advantages of our method and the underlying reasons for its robust performance. Bing Feng, David C. Samuels, William Hoskins, Jijun Tang, Zibo Meng |
BIBM | 6 |
| 2017 | HiCComp: Multiple-level comparative analysis of Hi-C data by triplet networkabstractHi-C technique is an important tool for the study of 3D genome organization. In the past few years, we have seen an explosion of Hi-C data in a variety of cell/tissue types. While these publicly available data presents an unprecedented opportunity to interrogate chromosomal architecture, how to quantitatively compare Hi-C data from different tissues and identify tissue-specific chromatin interactions remains challenging. Here, we present HiCComp, a comprehensive framework for comparing Hi-C data. HiCComp utilizes convolutional neural networks to extract key features in Hi-C interaction matrices in a fully automatic way. The core component of HiCComp is a triplet network, which contains three identical convolutional neural networks with shared parameters. The inputs to our network are three Hi-C matrices: two of them are biological replicates from the same cell type and the third one is from another cell type. The HiCComp network takes advantages of the two biological replicates to estimate the natural variation in the experiments and further use it to identify significant variations between Hi-C matrices from different cell types. Furthermore, we incorporate systematic occluding method into our framework so that we can identify the dynamic interaction regions from Hi-C maps. Finally, we show that the dynamic regions between two cell types are enriched for transcription factor binding sites and histone modifications that are associated with cis-regulatory functions, suggesting these variations in 3D genome structure are potentially gene regulatory events. W. Jim Zheng, Jijun Tang |
BIBM | 4 |
| 2017 | A Median Solver and Phylogenetic Inference Based on DCJ Sorting
Ruofan Xia, Lingxi Zhou, Bing Feng, Jijun Tang |
ISBRA | 5 |
| 2017 | Identification of drug-target interactions via multiple information integration
Yijie Ding, Jijun Tang, Fei Guo 0001 |
Inf. Sci. | 2 |
| 2017 | Local-DPP: An improved DNA-binding protein prediction method by exploring local evolutionary information
Leyi Wei, Jijun Tang, Quan Zou 0001 |
Inf. Sci. | 2 |
| 2016 | Predicting protein-protein interactions via multivariate mutual information of protein sequencesabstractBACKGROUND: Protein-protein interactions (PPIs) are central to a lot of biological processes. Many algorithms and methods have been developed to predict PPIs and protein interaction networks. However, the application of most existing methods is limited since they are difficult to compute and rely on a large number of homologous proteins and interaction marks of protein partners. In this paper, we propose a novel sequence-based approach with multivariate mutual information (MMI) of protein feature representation, for predicting PPIs via Random Forest (RF). METHODS: Our method constructs a 638-dimentional vector to represent each pair of proteins. First, we cluster twenty standard amino acids into seven function groups and transform protein sequences into encoding sequences. Then, we use a novel multivariate mutual information feature representation scheme, combined with normalized Moreau-Broto Autocorrelation, to extract features from protein sequence information. Finally, we feed the feature vectors into a Random Forest model to distinguish interaction pairs from non-interaction pairs. RESULTS: To evaluate the performance of our new method, we conduct several comprehensive tests for predicting PPIs. Experiments show that our method achieves better results than other outstanding methods for sequence-based PPIs prediction. Our method is applied to the S.cerevisiae PPIs dataset, and achieves 95.01 % accuracy and 92.67 % sensitivity repectively. For the H.pylori PPIs dataset, our method achieves 87.59 % accuracy and 86.81 % sensitivity respectively. In addition, we test our method on other three important PPIs networks: the one-core network, the multiple-core network, and the crossover network. CONCLUSIONS: Compared to the Conjoint Triad method, accuracies of our method are increased by 6.25,2.06 and 18.75 %, respectively. Our proposed method is a useful tool for future proteomics studies. Yijie Ding, Jijun Tang, Fei Guo 0001 |
BMC Bioinform. | 2 |
| 2016 | Learning from real imbalanced data of 14-3-3 proteins binding specificity
Jijun Tang, Fei Guo 0001 |
Neurocomputing | 2 |
| 2015 | Ancestral reconstruction under weighted maximum matchingabstractAncestral genome reconstruction has attracted increasing interests from both biologists and computer scientists. It has been conducted using various evolutionary models ever since comparative genomics moved from sequence data to gene order data. We propose a Flexible Ancestral Reconstruction Model, FARM, based on the maximum likelihood and weighted maximum matching algorithms, to infer ancestral gene orders. This will accommodate various evolutionary scenarios, including not only genomic rearrangements, but also insertion/deletions (indels), segment duplications, and whole genome duplications. We evaluate this work by using various simulated evolution experiments while comparing FARM to existing methods, like InferCarsPro, GASTS and PMAG++. FARM shows significant improvement in running time and the final assembling process and, therefore, can be used in large-scale real biological data ancestral inference. Lingxi Zhou, William Hoskins, Jieyi Zhao, Jijun Tang |
BIBM | 4 |
| 2015 | An Iterative Approach for Phylogenetic Analysis of Tumor Progression Using FISH Copy Number
Yu Lin 0001, William Hoskins, Jijun Tang |
ISBRA | 4 |
| 2015 | Maximum Parsimony Analysis of Gene Copy Number Changes
Yu Lin 0001, Vaibhav Rajan, William Hoskins, Jijun Tang |
WABI | 5 |
| 2015 | A Cooperative Co-Evolutionary Genetic Algorithm for Tree Scoring and Ancestral Genome InferenceabstractRecent advances of technology have made it easy to obtain and compare whole genomes. Rearrangements of genomes through operations such as reversals and transpositions are rare events that enable researchers to reconstruct deep evolutionary history among species. Some of the popular methods need to search a large tree space for the best scored tree, thus it is desirable to have a fast and accurate method that can score a given tree efficiently. During the tree scoring procedure, the genomic structures of internal tree nodes are also provided, which provide important information for inferring ancestral genomes and for modeling the evolutionary processes. However, computing tree scores and ancestral genomes are very difficult and a lot of researchers have to rely on heuristic methods which have various disadvantages. In this paper, we describe the first genetic algorithm for tree scoring and ancestor inference, which uses a fitness function considering co-evolution, adopts different initial seeding methods to initialize the first population pool, and utilizes a sorting-based approach to realize evolution. Our extensive experiments show that compared with other existing algorithms, this new method is more accurate and can infer ancestral genomes that are much closer to the true ancestors. Bing Feng, Jijun Tang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2014 | Assessing ancestral genome reconstruction methods by resamplingabstractInference of ancestral or extinct genomes is important in evolutionary biology, cancer research and many other research areas. Whole-genome data has become readily available due to advances in large-scale sequencing technology. During the past decades, a number of evolutionary models and related algorithms have been designed to infer ancestral genome sequence or gene order. Since it is so hard or even impossible to know the true scenario of the ancestral genomes, there must be some tools used to test the robustness of the adjacencies obtained from existing methods. However, till now, there is still no systematic work being conducted to tackle this problem. On the other hand, it is a common practice in phylogenetic analysis to assess the confidence rate of the inferred branches, using techniques such as bootstrapping and jackknifing, which have been evaluated as good resampling tools for phylogenetic reconstruction methods. Some of these resampling methods can be potentially used as a robustness test for ancestral genomes. In this paper, we conduct large-scale experiments by using three existing best ancestral genome inference methods and four different resampling techniques. Experimental results show that resampling techniques are useful for assessing the quality of ancestral genomes, while different methods have different preferred resampling techniques. It is suggested the cut-off threshold of 75% should be used to filter bad gene adjacencies. William Hoskins, Jijun Tang |
BIBM | 4 |
| 2014 | A Lin-Kernighan Heuristic for the DCJ Median Problem of Genomes with Unequal Contents
Zhaoming Yin, Jijun Tang, Stephen W. Schaeffer, David A. Bader |
COCOON | 2 |
| 2014 | MLGO: phylogeny reconstruction and ancestral inference from gene-order dataabstractBACKGROUND: The rapid accumulation of whole-genome data has renewed interest in the study of using gene-order data for phylogenetic analyses and ancestral reconstruction. Current software and web servers typically do not support duplication and loss events along with rearrangements. RESULTS: MLGO (Maximum Likelihood for Gene-Order Analysis) is a web tool for the reconstruction of phylogeny and/or ancestral genomes from gene-order data. MLGO is based on likelihood computation and shows advantages over existing methods in terms of accuracy, scalability and flexibility. CONCLUSIONS: To the best of our knowledge, it is the first web tool for analysis of large-scale genomic changes including not only rearrangements but also gene insertions, deletions and duplications. The web tool is available from http://www.geneorder.org/server.php . Yu Lin 0001, Jijun Tang |
BMC Bioinform. | 3 |
| 2014 | Probabilistic Reconstruction of Ancestral Gene Orders with Insertions and DeletionsabstractChanges of gene orderings have been extensively used as a signal to reconstruct phylogenies and ancestral genomes. Inferring the gene order of an extinct species has a wide range of applications, including the potential to reveal more detailed evolutionary histories, to determine gene content and ordering, and to understand the consequences of structural changes for organismal function and species divergence. In this study, we propose a new adjacency-based method, PMAG(+) , to infer ancestral genomes under a more general model of gene evolution involving gene insertions and deletions (indels), in addition to gene rearrangements. PMAG(+) improves on our previous method PMAG by developing a new approach to infer ancestral gene contents and reducing the adjacency assembly problem to an instance of TSP. We designed a series of experiments to extensively validate PMAG(+) and compared the results with the most recent and comparable method GapAdj. According to the results, ancestral gene contents predicted by PMAG(+) coincides highly with the actual contents with error rates less than 1 percent. Under various degrees of indels, PMAG(+) consistently achieves more accurate prediction of ancestral gene orders and at the same time, produces contigs very close to the actual chromosomes. Lingxi Zhou, Jijun Tang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2013 | Exploring genomes with a game engineabstractStudying genomes continues to be beneficial for evolutionary discover, and prognosticating/diagnosing many genetic disorders and diseases. Most of the these studies have used systems that view the DNA in a linear structure, but having this information is only a small part of fully understanding what they can reveal. Visualizing genomes in real time 3D can give researchers more insight, but this is fraught with hardware limitations. Each element contains vast amounts of information that cannot be processed at once. However, by using a game engine and sophisticated video game visualization techniques, we were able to construct a multi-platform real-time 3D genome viewer. Jeremiah J. Shepherd, Lingxi Zhou, W. Jim Zheng, Jijun Tang |
BIBM | 5 |
| 2013 | Exploring genomes with a game engine
Jeremiah J. Shepherd, Bill Arndt, W. Jim Zheng, Jijun Tang |
FDG | 5 |
| 2013 | Reconstructing Ancestral Genomic Orders Using Binary Encoding and Probabilistic Models
Lingxi Zhou, Jijun Tang |
ISBRA | 3 |
| 2011 | Emulating Insertion and Deletion Events in Genome Rearrangement AnalysisabstractGreat advancements have been achieved in phylogenetic reconstruction from genome rearrangement events but difficult problems still remain. One challenge is to deal with more complex events such as gene insertions and deletions such that we can analyze both gene order and gene content changes in tandem. We propose the concept of prosthetic chromosomes to incorporate these events into the standard double-cut-and-join (DCJ) distance metric widely used for genome scale rearrangement analysis. In this paper, we also introduce our new software package Egchel (Extended Gene Content HEuristic Layer) which implements this prosthetic chromosome model. Egchel can modify unequal content gene order data sets such that they are able to be analyzed by existing phylogenetic tree builders specifically designed to work with equal content gene order data sets. When compared to existing pairwise analysis of unequal content data sets, Egchel uses a global approach, produces substantially more accurate phylogenetic trees, and is significantly more likely to generate the best available tree. William Arndt, Jijun Tang |
BIBM | 2 |
| 2011 | Maximum likelihood phylogenetic reconstruction using gene order encodingsabstractGene order changes under rearrangement events such as inversions and transpositions have attracted increasing attention as a new type of data for phylogenetic analysis. Since these events are rare, they allow the reconstruction of evolutionary history far back in time. Many software have been developed for the inference of gene order phylogenies, including widely used maximum parsimony methods such as GRAPPA and MGR. However, these methods confronted great difficulties in dealing with emerging large nuclear genomes. In this study, we proposed three simple yet powerful maximum likelihood(ML) based methods for phylogenetic reconstruction by first encoding the gene orders into binary or multistate strings based on gene adjacency information presented in the given genomes and further converting these strings into molecular sequences. RAxML is at last used to compute the maximum likelihood phylogeny. We conducted extensive experiments using simulated datasets and found that although the multistate encoding is more complex and more time-consuming, it did not improve accuracy over the methods using simpler binary encodings. Among all methods tested in our experiments, MLBE is of the most accuracy in most cases and often returns phylogenies without errors. ML methods is also fast and in the most difficult case only takes up to three days to compute datasets with 40 genomes, making it very suitable for large scale analysis. We give three simple and robust phylogenetic reconstruction methods using different encodings based on maximum likelihood which has not been successfully applied for gene orderings before. Our development of these ML methods showed great potential in gene order analysis with respect to the high accuracy and stability, although formal mathematical and statistical analysis of these methods are much desired. Meng Zhang 0006, Jijun Tang |
CIBCB | 4 |
| 2011 | Isolating - a new resampling method for gene order dataabstractThe purpose of using resampling methods on phylogenetic data is to estimate the confidence value of branches. In recent years, bootstrapping and jackknifing are the two most popular resampling schemes which are widely used in biological reserach. However, for gene order data, traditional bootstrap procedures can not be applied because gene order data is viewed as one character with various states. Experience in the biological community has shown that jackknifing is a useful means of determining the confidence value of a gene order phylogeny. When genomes are distant, however, applying jackknifing tends to give low confidence values to many valid branches, causing them to be mistakenly removed. In this paper, we propose a new method that overcomes this disadvantage of jackknifing and achieves better accuracy and confidence values for gene order data. Compared to jackknifing, our experimental results show that the proposed method can produce phylogenies with lower error rates and much stronger support for good branches. We also establish a theoretic lower bound regarding how many genes should be isolated, which is confirmed empirically. William Arndt, Jijun Tang |
CIBCB | 4 |
| 2011 | A different approach to teaching Chinese through serious gamesabstractFrom the days of computer assisted language learning (CALL) [10], using computers as a means to acquire a new language has been a long standing research field. Chinese is a notoriousloy difficult language to learn, and teaching methods that are used take a long time and can be tedious[7], but these techniques have been shown to be effective[11]. We believe that by creating a serious game that uses newer second language acauisition techniques, the learning process can be expidited and enjoyable. Our game exposes the player to an abundance of simple, comprehensible target language input, which provides an interesting, motivating, and low stress setting. Also Lost in the Middle Kingdom utilizes total immersion, focusing on the language's culture to create a holistic experience. Jeremiah J. Shepherd, Renaldo J. Doe, Matthew Arnold, Jijun Tang |
FDG | 5 |
| 2010 | Phylogenetic reconstruction with gene rearrangements and gene lossesabstractReconstructing phylogenies from gene-order data has become very attractive in the research of evolution these years. So far, most methods can only treat genomes with equal gene contents with each gene appearing exactly once in each genome. In this paper, we propose a new distance measurement for genomes with inversions and insertions/deletions that comply with triangle inequality. Based on this distance, we develop a new method to solve the median problem of unequal gene content, which are used to reconstruct both phylogenies and ancestral genomes. We test our method on simulated datasets under various conditions and the experimental results show that our distance measurement can produce more accurate phylogenetic trees compared with other popular methods for unequal genomes. Also our median algorithm produces remarkably more accurate ancestral genomes than the only unequal genome median solver that is currently available. Jijun Tang |
BIBM | 3 |
| 2010 | Genome3D: A viewer-model framework for integrating and visualizing multi-scale epigenomic information within a three-dimensional genomeabstractBACKGROUND: New technologies are enabling the measurement of many types of genomic and epigenomic information at scales ranging from the atomic to nuclear. Much of this new data is increasingly structural in nature, and is often difficult to coordinate with other data sets. There is a legitimate need for integrating and visualizing these disparate data sets to reveal structural relationships not apparent when looking at these data in isolation. RESULTS: We have applied object-oriented technology to develop a downloadable visualization tool, Genome3D, for integrating and displaying epigenomic data within a prescribed three-dimensional physical model of the human genome. In order to integrate and visualize large volume of data, novel statistical and mathematical approaches have been developed to reduce the size of the data. To our knowledge, this is the first such tool developed that can visualize human genome in three-dimension. We describe here the major features of Genome3D and discuss our multi-scale data framework using a representative basic physical model. We then demonstrate many of the issues and benefits of multi-resolution data integration. CONCLUSIONS: Genome3D is a software visualization tool that explores a wide range of structural genomic and epigenetic data. Data from various sources of differing scales can be integrated within a hierarchical framework that is easily adapted to new developments concerning the structure of the physical genome. In addition, our tool has a simple annotation mechanism to incorporate non-structural information. Genome3D is unique is its ability to manipulate large amounts of multi-resolution data from diverse sources to uncover complex and new structural relationships within the genome. Thomas M. Asbury, Matt Mitman, Jijun Tang, W. Jim Zheng |
BMC Bioinform. | 3 |
| 2010 | Using jackknife to assess the quality of gene order phylogeniesabstractBACKGROUND: In recent years, gene order data has attracted increasing attention from both biologists and computer scientists as a new type of data for phylogenetic analysis. If gene orders are viewed as one character with a large number of states, traditional bootstrap procedures cannot be applied. Researchers began to use a jackknife resampling method to assess the quality of gene order phylogenies. RESULTS: In this paper, we design and conduct a set of experiments to validate the performance of this jackknife procedure and provide discussions on how to conduct it properly. Our results show that jackknife is very useful to determine the confidence level of a phylogeny obtained from gene orders and a jackknife rate of 40% should be used. However, although a branch with support value of 85% can be trusted, low support branches require careful investigation before being discarded. CONCLUSIONS: Our experiments show that jackknife is indeed necessary and useful for gene order data, yet some caution should be taken when the results are interpreted. Haiwei Luo, Jijun Tang |
BMC Bioinform. | 4 |
| 2009 | Maximum independent sets of commuting and noninterfering inversionsabstractBACKGROUND: Given three signed permutations, an inversion median is a fourth permutation that minimizes the sum of the pairwise inversion distances between it and the three others. This problem is NP-hard as well as hard to approximate. Yet median-based approaches to phylogenetic reconstruction have been shown to be among the most accurate, especially in the presence of long branches. Most existing approaches have used heuristics that attempt to find a longest sequence of inversions from one of the three permutations that, at each step in the sequence, moves closer to the other two permutations; yet very little is known about the quality of solutions returned by such approaches. RESULTS: Recently, Arndt and Tang took a step towards finding longer such sequences by using sets of commuting inversions. In this paper, we formalize the problem of finding such sequences of inversions with what we call signatures and provide algorithms to find maximum cardinality sets of commuting and noninterfering inversions. CONCLUSION: Our results offer a framework in which to study the inversion median problem, faster algorithms to obtain good medians, and an approach to study characteristic events along an evolutionary path. Krister M. Swenson, Yokuki To, Jijun Tang, Bernard M. E. Moret |
BMC Bioinform. | 3 |
| 2009 | Simultaneous phylogeny reconstruction and multiple sequence alignmentabstractBACKGROUND: A phylogeny is the evolutionary history of a group of organisms. To date, sequence data is still the most used data type for phylogenetic reconstruction. Before any sequences can be used for phylogeny reconstruction, they must be aligned, and the quality of the multiple sequence alignment has been shown to affect the quality of the inferred phylogeny. At the same time, all the current multiple sequence alignment programs use a guide tree to produce the alignment and experiments showed that good guide trees can significantly improve the multiple alignment quality. RESULTS: We devise a new algorithm to simultaneously align multiple sequences and search for the phylogenetic tree that leads to the best alignment. We also implemented the algorithm as a C program package, which can handle both DNA and protein data and can take simple cost model as well as complex substitution matrices, such as PAM250 or BLOSUM62. The performance of the new method are compared with those from other popular multiple sequence alignment tools, including the widely used programs such as ClustalW and T-Coffee. Experimental results suggest that this method has good performance in terms of both phylogeny accuracy and alignment quality. CONCLUSION: We present an algorithm to align multiple sequences and reconstruct the phylogenies that minimize the alignment score, which is based on an efficient algorithm to solve the median problems for three sequences. Our extensive experiments suggest that this method is very promising and can produce high quality phylogenies and alignments. Jijun Tang |
BMC Bioinform. | 3 |
| 2008 | Phylogenetic Reconstruction from Complete Gene Orders of Whole Genomes
Krister M. Swenson, William Arndt, Jijun Tang, Bernard M. E. Moret |
APBC | 3 |
| 2008 | Phylogenetic reconstruction with disk-covering and Bayesian approachesabstractThe DCM approach is commonly used to divide the dataset into smaller subproblems, analyze each subproblem using a base method to obtain subtrees, then recombine these subtrees to build the final phylogeny over the whole dataset. In recent years, the new and improved method MrBayes, a Bayesian Markov Chain Monte Carlo (MCMC) approach is widely used for phylogeny analysis. In this paper, a new method for large scale Bayesian phylogeny analysis is proposed. This new method (DCM3-MrBayes) is an improved version of Rec-I-DCM3 (recursive iterative disk-covering method), which uses a divide-and-conquer approach and is designed for large dataset analysis. To integrate MrBayes with Rec-I-DCM3, we have to deal with some unique problems and proposed several methods to tackle these problems. Our improvements include a cache system that can avoid unnecessary computations and a method to eliminate weak branches indicated by the Bayesian analysis to filter out potential bad branches. Our experiments on simulated datasets shows promising improvement over the original DCM. One of the most important advantages of using Bayesian method for phylogeny reconstruction is being able to calculate the posterior probabilities. A divide-and-conquer Bayesian method looses its ability to calculate the posterior probabilities due to the fact that each subproblem generates its own posterior probabilities, which posts some difficulties for obtaining the posterior probability for the whole problem. In order to preserve the advantage of Bayesian approach, we also introduce an algorithm that calculates the posterior probabilities of the whole phylogeny from the subproblemspsila posterior probabilities. Jijun Tang |
BIBE | 3 |
| 2008 | A Branch-and-Bound Method for the Multichromosomal Reversal Median Problem
Meng Zhang 0006, William Arndt, Jijun Tang |
WABI | 3 |
| 2007 | FPGA Acceleration of Phylogeny Reconstruction for Whole Genome DataabstractIn this paper we describe our design and characterization of a co-processor architecture to accelerate median-based phylogenetic reconstruction for gene-rearrangement data. Our current design performs a parallelized version of the breakpoint median computation and achieves an average speedup of 876 for simulated input data having a high evolution rate. After integrating our hardware-based median computation into the GRAPPA toolset, we have achieved an average speedup of 189 over the entire phylogenetic reconstruction procedure. The results in this paper suggest that FPGA-based acceleration is a promising approach for computationally expensive phylogenetic problems that are based on combinatorial optimization. Jason D. Bakos, Panormitis E. Elenis, Jijun Tang |
BIBE | 3 |
| 2007 | A Heuristic for Phylogenetic Reconstruction Using TranspositionabstractBecause of the advent of high-throughput sequencing and the consequent reduction in cost of sequencing, many organisms have been completely sequenced and most of their genes identified; homologies among these genes are also getting established. It thus has become possible to represent whole genomes as ordered lists of gene identifiers and to study the evolution of these entities through computational means, in systematics as well as in comparative genomics. As a result, gene order data (also known as genome rearrangement data) has attracted increasing attention from both biologists and computer scientists as a new type of data for phylogenetic analysis. Methods for reconstructing phylogeny from genome rearrangements include distance-based methods, MCMC methods and direct optimization methods. The latter, pioneered by Sankoff and extended in the software packages of GRAPPA and MGR, is the most accurate approach for inversion phylogeny. However, due to the difficulty of computing the transposition distance, this type of methods has not been applied to datasets where transposition is the only or dominant event. In this paper, we present a heuristic transposition median solver and extend GRAPPA to handle transpositions. Our extensive testing using simulated datasets shows that this method (GRAPPA-TP) is very accurate in terms of ancestor genome inference and phylogenetic reconstruction. It also suggests that model match is critical in phylogenetic analysis, and a fast and accurate method for transposition distance computation is still very important. The new GRAPPA-TP is available from phylo.cse.sc.edu. Meng Zhang 0006, Jijun Tang |
BIBE | 3 |
| 2007 | A Divide-and-Conquer Implementation of Three Sequence Alignment and Ancestor InferenceabstractIn this paper, we present an algorithm to simultaneously align three biological sequences with affine gap model and infer their common ancestral sequence. Our algorithm can be further extended to perform tree alignment for more se- quences, and eventually unify the two procedures of phylo- genetic reconstruction and sequence alignment. The nov- elty of our algorithm is: it applies the divide-and-conquer strategy so that the memory usage is reduced from O (n3) to O (n2), while at the same time, it is based on dynamic programming and optimal alignment is guaranteed. Tra- ditionally, three sequence alignment is limited by the huge demand of memory space and can only handle sequences less than two hundred characters long. With the new im- proved algorithm, we can produce the optimal alignment of sequences of several thousand characters long. We implemented our algorithm as a C program package MSAM . It has been extensively tested with BAliBASE, a real manually refined multiple sequence alignment database, as well as simulated datasets generated by Rose (Ran- dom Model of Sequence Evolution). We compared our re- sults with those of other popular multiple sequence align- ment tools, including the widely used programs such as ClustalW and T-Coffee. The experiment shows that MSAM produces not only better alignment, but also better ancestral sequence. The software can be downloaded for free at http://www.cse.sc.edu/phylo/MSAM.html Jijun Tang |
BIBM | 2 |
| 2007 | A Fast 3D Correspondence Method for Statistical Shape ModelingabstractAccurately identifying corresponded landmarks from a population of shape instances is the major challenge in constructing statistical shape models. In this paper, we address this landmark-based shape-correspondence problem for 3D cases by developing a highly efficient landmark-sliding algorithm. This algorithm is able to quickly refine all the landmarks in a parallel fashion by sliding them on the 3D shape surfaces. We use 3D thin-plate splines to model the shape-correspondence error so that the proposed algorithm is invariant to affine transformations and more accurately reflects the nonrigid biological shape deformations between different shape instances. In addition, the proposed algorithm can handle both open-and closed-surface shape, while most of the current 3D shape-correspondence methods can only handle genus-0 closed surfaces. We conduct experiments on 3D hippocampus data and compare the performance of the proposed algorithm to the state-of-the-art MDL and SPHARM methods. We find that, while the proposed algorithm produces a shape correspondence with a better or comparable quality to the other two, it takes substantially less CPU time. We also apply the proposed algorithm to correspond 3D diaphragm data which have an open-surface shape. Pahal Dalal, Brent C. Munsell, Song Wang 0002, Jijun Tang, Kenton Oliver, Hiroaki Ninomiya, Xiangrong Zhou, Hiroshi Fujita 0001 |
CVPR | 4 |
| 2006 | Succinct Text Indexes on Large Alphabet
Meng Zhang 0006, Jijun Tang, Dong Guo 0002, Liang Hu 0001, Qiang Li 0008 |
TAMC | 2 |
| 2005 | Improving Genome Rearrangement Phylogeny Using Sequence-Style ParsimonyabstractThe study of genome rearrangements, the evolutionary events that change the order and strandedness of genes within genomes, presents new opportunities for discoveries about deep evolutionary events. The best software so far, GRAPPA, solves breakpoint and inversion phylogenies by scoring each tree topology through iterative improvements of internal node gene orders. We find that the greedy hill-climbing approach means the accuracy is limited because of multiple local optima. To address this problem, we propose integration GRAPPA with MPME, a string encoding of gene adjacency relationships whose optimal internal node assignments can be determined globally in polynomial time, to provide better initializations for GRAPPA. In simulation studies, the new algorithm yields shorter tree lengths and better accuracy in phylogeny reconstruction. Jijun Tang, Li-San Wang |
BIBE | 1 |
| 2005 | Quartet-Based Phylogeny Reconstruction from Gene Orders
Jijun Tang, Bernard M. E. Moret |
COCOON | 2 |
| 2005 | Linear Programming for Phylogenetic Reconstruction Based on Gene Rearrangements
Jijun Tang, Bernard M. E. Moret |
CPM | 1 |
| 2004 | Phylogenetic Reconstruction from Arbitrary Gene-Order DataabstractPhylogenetic reconstruction from gene-order data has attracted attention from both biologists and computer scientists over the last few years. So far, our software suite GRAPPA is the most accurate approach, but it requires that all genomes have identical gene content, with each gene appearing exactly once in each genome. Some progress has been made in handling genomes with unequal gene content, both in terms of computing pair-wise genomic distances and in terms of reconstruction. In this paper, we present a new approach for computing the median of three arbitrary genomes and apply it to the reconstruction of phylogenies from arbitrary gene-order data. We implemented these methods within GRAPPA and tested them on simulated datasets under various conditions as well as on a real dataset of chloroplast genomes; we report the results of our simulations and our analysis of the real dataset and compare them to reconstructions made by using neighbor-joining and using the original GRAPPA on the same genomes with equalized gene contents. Our new approach is remarkably accurate both in simulations and on the real dataset, in contrast to the distance-based approaches and to reconstructions using the original GRAPPA applied to equalized gene contents. Jijun Tang, Bernard M. E. Moret, Liying Cui, Claude W. dePamphilis |
BIBE | 1 |
| 2003 | Phylogenetic Reconstruction from Gene-Rearrangement Data with Unequal Gene Content
Jijun Tang, Bernard M. E. Moret |
WADS | 1 |
| 2002 | Inversion Medians Outperform Breakpoint Medians in Phylogeny Reconstruction from Gene-Order Data
Bernard M. E. Moret, Adam C. Siepel, Jijun Tang |
WABI | 3 |
| 2002 | Steps toward accurate reconstructions of phylogenies from gene-order data
Bernard M. E. Moret, Jijun Tang, Li-San Wang, Tandy J. Warnow |
J. Comput. Syst. Sci. | 2 |