Jinbo Xu

dblp:45/2923 · DBLP profile ↗
← Back
84ranked-venue papers
15as first author
20since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 51 · 7 first-author · 12 since 2021Artificial intelligence and machine learning · 13 · 4 since 2021Systems, architecture and hardware · 11 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Theory of computation · 3Computer networks · 2 · 2 first-author
YearPublicationVenuePosition
2025 PIAR: Path-Improved Adaptive Routing for Dragonfly Networks
abstract
For the next-generation exascale supercomputing communication systems, Dragonfly topology offers strong scalability, low latency, and cost efficiency. Dragonfly networks have already been implemented in current supercomputers and will continue to expand in future systems. Adaptive routing in Dragonfly topologies is critical for network performance. The traditional UGAL routing algorithm, which uses the valiant mechanism to select non-minimal paths, does not adequately consider the impact of high hops in non-minimal paths, often unnecessarily increasing the average path length, thereby increasing network latency and load. Furthermore, UGAL inaccurately estimates the congestion of the entire routing path based on local information, leading to suboptimal routing decisions that limit the algorithm's performance. In this paper, we propose PIAR, a novel pathimproved adaptive routing algorithm. PIAR dynamically selects paths based on the status of local and global channels, prioritizing non-minimal paths with fewer hops to reduce network latency and load, thereby improving network performance. Additionally, we present the microarchitecture of the routing computation unit. Our evaluation results demonstrate that, compared with advanced algorithms such as PAR$_{\text {PH }}$, TPR, and UGAL LE, PIAR achieves an average throughput improvement of 19.2 % and reduces latency by up to$\mathbf{1 3. 4 \%}$under the single synthetic traffic. Under mixed traffic, PIAR achieves an average throughput improvement of$\mathbf{2 3. 6 \%}$and reduces the latency by up to$\mathbf{3 3. 8 \%}$. For application workloads, PIAR achieves an average reduction of 24.0 % in packet latency.
Qiang Wang 0006, Jinbo Xu, Guo Chen 0001
CLUSTER5
2025 Retrieval Augmented Zero-Shot Enzyme Generation for Specified Substrate
abstract
Generating novel enzymes for target molecules in zero-shot scenarios is a fundamental challenge in biomaterial synthesis and chemical production. Without known enzymes for a target molecule, training generative models becomes difficult due to the lack of direct supervision. To address this, we propose a retrieval-augmented generation method that uses existing enzyme-substrate data to guide enzyme design. Our method retrieves enzymes with substrates that share structural similarities with the target molecule, leveraging functional similarities in catalytic activity. Since none of the retrieved enzymes directly catalyze the target molecule, we use a conditioned discrete diffusion model to generate new enzymes based on the retrieved examples. An enzyme-substrate relationship classifier guides the generation process to ensure optimal protein sequence distributions. We evaluate our model on enzyme design tasks with diverse real-world substrates and show that it outperforms existing protein generation methods in catalytic capability, foldability, and docking accuracy. Additionally, we define the zero-shot substrate-specified enzyme generation task and introduce a dataset with evaluation benchmarks.
Jiahe Du, Kaixiong Zhou, Xinyu Hong, Zhaozhuo Xu, Jinbo Xu, Xiao Huang 0001
ICML5
2025 Surface-based Molecular Design with Multi-modal Flow Matching
abstract
Therapeutic peptides show promise in targeting previously undruggable binding sites, with recent advancements in deep generative models enabling full-atom peptide co-design for specific protein receptors.However, the critical role of molecular surfaces in proteinprotein interactions (PPIs) has been underexplored.To bridge this gap, we propose an omni-design peptides generation paradigm, called SurfFlow, a novel surface-based generative algorithm that enables comprehensive co-design of sequence, structure, and surface for peptides.SurfFlow employs a multi-modality conditional flow matching (CFM) architecture to learn distributions of surface geometries and biochemical properties, enhancing peptide binding accuracy.Evaluated on the comprehensive PepMerge benchmark, SurfFlow consistently outperforms full-atom baselines across all metrics.These results highlight the advantages of considering molecular surfaces in de novo peptide discovery and demonstrate the potential of integrating multiple protein modalities for more effective therapeutic peptide discovery.
Fang Wu 0002, Zhengyuan Zhou, Shuting Jin, Xiangxiang Zeng, Jure Leskovec, Jinbo Xu
KDD (2)6
2024 Prediction of single-stranded DNA binding proteins with protein language model
abstract
Single-stranded DNA binding proteins (SSBs) represent one important type of nucleic acid binding proteins and are crucial for many cellular functions, such as DNA replication, transcriptional regulation, and maintaining genome stability and integrity. Compared to other types of nucleic acid binding proteins, SSBs are underexplored due to a relatively small number of annotated SSBs and experimentally solved SSB structures. Only a few computational approaches have been developed so far to classify SSBs from double-stranded DNA binding proteins (DSBs), which typically assume the inputs are known DNA binding proteins (DBPs). However, our recent study reveals that a large number of SSBs are predicted as RNA binding proteins (RBPs) for any given nucleic acid binding proteins, suggesting that SSB/DSB classification alone is not enough and SSB/RBP classification is equally important in accurately predicting novel SSBs from proteins with unknown functions. In this study, in addition to a much improved SSB/DSB classifier, we developed the first SSB/RBP classifier with machine learning models. Besides the widely used PSSM (position specific scoring matrix), we tested a structural feature and a feature extracted from a protein language model ESM2 for training and testing. Our results show that ESM2 improves prediction accuracy dramatically, up to 95% for the SSB/DSB classifier and 92% for the SSB/RBP classifier respectively with a support vector machine (SVM) approach.
Siwen Wu, Jinbo Xu, Jun-tao Guo
BIBM2
2024 The Self-adaptive and Topology-aware MPI_Bcast leveraging Collective offload on Tianhe Express Interconnect
abstract
Large parallel applications have heavily used MPI (Massage Passing Interface) collectives that support portable and efficient group communication operations. MPI_Bcast is one of the most commonly used MPI collectives that broadcast data to all processes of the communication domain. However, traditional software-based broadcast algorithms fail to fully utilize modern interconnection networks’ advanced features such as offloading collectives to the network hardware for efficient group communications. Besides, the semantic gap between MPI_Bcast and hardware multicast of underlying interconnects presents challenges for offload-based algorithms to accelerate MPI_Bcast for a wide range of message sizes.In this paper, we propose a hardware-software co-design MPI_Bcast by efficiently leveraging the NIC-based collective offload provided by Tianhe-express interconnect, which completely precludes the involvement of CPU to accelerate message broadcast. We detail this broadcast mechanism that can be adaptively tuned to offload MPI_Bcast operations from the CPU to the NIC for various message and system sizes. In addition, we further propose a topology-aware broadcast design in conjunction with this offload method to significantly reduce the broadcast latency by constructing the optimal global inter-node communication tree. We implement and evaluate the proposed Tianhe-Express Offload-based Broadcast (TOB) design on Tianhe-2A and Tianhe-EP supercomputers. Extensive experiments have been conducted to evaluate TOB performance at both microbenchmark and application levels. Our solution offers up to 4.94x significant performance speedup at the microbenchmark level over state-of-the-art MPI libraries. For the application-level evaluation, our technique accelerates scientific applications by a maximum speedup of 1.34x.
Chongshan Liang, Jinbo Xu, Jintao Peng, Weixia Xu 0001, Jie Liu 0002, Zhiquan Lai, Sheng Ma
IPDPS4
2024 Collective Communication Acceleration Architecture for Reduce Operation with Long Operands Based on Efficient Data Access and Reconfigurable On-Chip Buffer
abstract
The Reduce operation with long operands is widely used in high-performance computing and AI(Artificial Intelligence) computing. A hardware acceleration architecture for Reduce operation with long operands is designed and implemented in FPGA. Flow control of the accelerator is achieved by using a dedicated hardware trigger mechanism. Based on DMA(Direct Memory Access), an efficient data access method is proposed between host memory and the accelerator. A reconfigurable on-chip buffer structure is designed to achieve flexible and efficient buffering of a large number of operands. Calculations are directly performed on the accelerator with on-chip ALU(Arithmetic Logic Unit) arrays to reduce the communications between the host and NIC(Network Interface Card). Experimental results show that this work has a significant acceleration effect compared to the non-offloading method and the current offloading method in "Tianhe" interconnection.
Jinbo Xu, Zhengbin Pang
ISPA1
2024 ABAG-docking benchmark: a non-redundant structure benchmark dataset for antibody-antigen computational docking
abstract
Accurate prediction of antibody-antigen complex structures is pivotal in drug discovery, vaccine design and disease treatment and can facilitate the development of more effective therapies and diagnostics. In this work, we first review the antibody-antigen docking (ABAG-docking) datasets. Then, we present the creation and characterization of a comprehensive benchmark dataset of antibody-antigen complexes. We categorize the dataset based on docking difficulty, interface properties and structural characteristics, to provide a diverse set of cases for rigorous evaluation. Compared with Docking Benchmark 5.5, we have added 112 cases, including 14 single-domain antibody (sdAb) cases and 98 monoclonal antibody (mAb) cases, and also increased the proportion of Difficult cases. Our dataset contains diverse cases, including human/humanized antibodies, sdAbs, rodent antibodies and other types, opening the door to better algorithm development. Furthermore, we provide details on the process of building the benchmark dataset and introduce a pipeline for periodic updates to keep it up to date. We also utilize multiple complex prediction methods including ZDOCK, ClusPro, HDOCK and AlphaFold-Multimer for testing and analyzing this dataset. This benchmark serves as a valuable resource for evaluating and advancing docking computational methods in the analysis of antibody-antigen interaction, enabling researchers to develop more accurate and effective tools for predicting and designing antibody-antigen complexes. The non-redundant ABAG-docking structure benchmark dataset is available at https://github.com/Zhaonan99/Antibody-antigen-complex-structure-benchmark-dataset.
Bingqing Han, Cuicui Zhao, Jinbo Xu, Xinqi Gong
Briefings Bioinform.4
2024 Support-Query Mutual Promotion and Classification Correction Network for Few-Shot Object Detection
abstract
Recently, finetuning-based methods have shown great performance in few-shot object detection. These methods employ pure convolutional structures for feature extraction, and then finetune the last layers of the base detector on novel classes. However, most of them have the following two issues: 1) the feature extraction processes of support and query sets lack mutual promotion, and 2) the features extracted by pure convolutional structures have low discriminability for confused classes, which could lead to false positive results. To overcome these two issues, we propose a support-query mutual promotion and classification correction network (SMPCCNet). First, we build a support-query mutual promotion module. In this module, we perform class attention operations with query features to produce enhanced support features, and use the dense kernels generated by the enhanced support features to obtain support-relevant query features, which achieves the mutual promotion of support and query features. In addition, we design a classification correction branch with a hybrid attention module. The hybrid attention module uses spatial and channel attentions to generate features that focus on the positions and types of detected objects, respectively, which can better discriminate confused classes and correct false positive results. Extensive experiments demonstrate the superiority of SMPCCNet over state-of-the-art methods.
Jinbo Xu, Yong Wang 0002, Yiqun Zou
IEEE Signal Process. Lett.1
2023 Improved the heterodimer protein complex prediction with protein language models
abstract
AlphaFold-Multimer has greatly improved the protein complex structure prediction, but its accuracy also depends on the quality of the multiple sequence alignment (MSA) formed by the interacting homologs (i.e. interologs) of the complex under prediction. Here we propose a novel method, ESMPair, that can identify interologs of a complex using protein language models. We show that ESMPair can generate better interologs than the default MSA generation method in AlphaFold-Multimer. Our method results in better complex structure prediction than AlphaFold-Multimer by a large margin (+10.7% in terms of the Top-5 best DockQ), especially when the predicted complex structures have low confidence. We further show that by combining several MSA generation methods, we may yield even better complex structure prediction accuracy than Alphafold-Multimer (+22% in terms of the Top-5 best DockQ). By systematically analyzing the impact factors of our algorithm we find that the diversity of MSA of interologs significantly affects the prediction accuracy. Moreover, we show that ESMPair performs particularly well on complexes in eucaryotes.
Bo Chen 0026, Ziwei Xie, Jiezhong Qiu, Zhaofeng Ye, Jinbo Xu, Jie Tang 0001
Briefings Bioinform.5
2023 Improving protein structure prediction using templates and sequence embedding
abstract
MOTIVATION: Protein structure prediction has been greatly improved by deep learning, but the contribution of different information is yet to be fully understood. This article studies the impacts of two kinds of information for structure prediction: template and multiple sequence alignment (MSA) embedding. Templates have been used by some methods before, such as AlphaFold2, RoseTTAFold and RaptorX. AlphaFold2 and RosetTTAFold only used templates detected by HHsearch, which may not perform very well on some targets. In addition, sequence embedding generated by pre-trained protein language models has not been fully explored for structure prediction. In this article, we study the impact of templates (including the number of templates, the template quality and how the templates are generated) on protein structure prediction accuracy, especially when the templates are detected by methods other than HHsearch. We also study the impact of sequence embedding (generated by MSATransformer and ESM-1b) on structure prediction. RESULTS: We have implemented a deep learning method for protein structure prediction that may take templates and MSA embedding as extra inputs. We study the contribution of templates and MSA embedding to structure prediction accuracy. Our experimental results show that templates can improve structure prediction on 71 of 110 CASP13 (13th Critical Assessment of Structure Prediction) targets and 47 of 91 CASP14 targets, and templates are particularly useful for targets with similar templates. MSA embedding can improve structure prediction on 63 of 91 CASP14 (14th Critical Assessment of Structure Prediction) targets and 87 of 183 CAMEO targets and is particularly useful for proteins with shallow MSAs. When both templates and MSA embedding are used, our method can predict correct folds (TMscore > 0.5) for 16 of 23 CASP14 FM targets and 14 of 18 Continuous Automated Model Evaluation (CAMEO) targets, outperforming RoseTTAFold by 5% and 7%, respectively. AVAILABILITY AND IMPLEMENTATION: Available at https://github.com/xluo233/RaptorXFold. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Fandi Wu, Xiaoyang Jing, Jinbo Xu
Bioinform.4
2022 Accurate protein function prediction via graph attention networks with predicted structure information
abstract
Experimental protein function annotation does not scale with the fast-growing sequence databases. Only a tiny fraction (<0.1%) of protein sequences has experimentally determined functional annotations. Computational methods may predict protein function very quickly, but their accuracy is not very satisfactory. Based upon recent breakthroughs in protein structure prediction and protein language models, we develop GAT-GO, a graph attention network (GAT) method that may substantially improve protein function prediction by leveraging predicted structure information and protein sequence embedding. Our experimental results show that GAT-GO greatly outperforms the latest sequence- and structure-based deep learning methods. On the PDB-mmseqs testset where the train and test proteins share <15% sequence identity, our GAT-GO yields Fmax (maximum F-score) 0.508, 0.416, 0.501, and area under the precision-recall curve (AUPRC) 0.427, 0.253, 0.411 for the MFO, BPO, CCO ontology domains, respectively, much better than the homology-based method BLAST (Fmax 0.117, 0.121, 0.207 and AUPRC 0.120, 0.120, 0.163) that does not use any structure information. On the PDB-cdhit testset where the training and test proteins are more similar, although using predicted structure information, our GAT-GO obtains Fmax 0.637, 0.501, 0.542 for the MFO, BPO, CCO ontology domains, respectively, and AUPRC 0.662, 0.384, 0.481, significantly exceeding the just-published method DeepFRI that uses experimental structures, which has Fmax 0.542, 0.425, 0.424 and AUPRC only 0.313, 0.159, 0.193.
Boqiao Lai, Jinbo Xu
Briefings Bioinform.2
2022 A tale of solving two computational challenges in protein science: neoantigen prediction and protein structure prediction
abstract
In this article, we review two challenging computational questions in protein science: neoantigen prediction and protein structure prediction. Both topics have seen significant leaps forward by deep learning within the past five years, which immediately unlocked new developments of drugs and immunotherapies. We show that deep learning models offer unique advantages, such as representation learning and multi-layer architecture, which make them an ideal choice to leverage a huge amount of protein sequence and structure data to address those two problems. We also discuss the impact and future possibilities enabled by those two applications, especially how the data-driven approach by deep learning shall accelerate the progress towards personalized biomedicine.
Ngoc Hieu Tran, Jinbo Xu, Ming Li 0001
Briefings Bioinform.2
2022 Deep graph learning of inter-protein contacts
abstract
MOTIVATION: Inter-protein (interfacial) contact prediction is very useful for in silico structural characterization of protein-protein interactions. Although deep learning has been applied to this problem, its accuracy is not as good as intra-protein contact prediction. RESULTS: We propose a new deep learning method GLINTER (Graph Learning of INTER-protein contacts) for interfacial contact prediction of dimers, leveraging a rotational invariant representation of protein tertiary structures and a pretrained language model of multiple sequence alignments. Tested on the 13th and 14th CASP-CAPRI datasets, the average top L/10 precision achieved by GLINTER is 54% on the homodimers and 52% on all the dimers, much higher than 30% obtained by the latest deep learning method DeepHomo on the homodimers and 15% obtained by BIPSPI on all the dimers. Our experiments show that GLINTER-predicted contacts help improve selection of docking decoys. AVAILABILITY AND IMPLEMENTATION: The software is available at https://github.com/zw2x/glinter. The datasets are available at https://github.com/zw2x/glinter/data. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ziwei Xie, Jinbo Xu
Bioinform.2
2022 Annotating functional effects of non-coding variants in neuropsychiatric cell types by deep transfer learning
abstract
Genomewide association studies (GWAS) have identified a large number of loci associated with neuropsychiatric traits, however, understanding the molecular mechanisms underlying these loci remains difficult. To help prioritize causal variants and interpret their functions, computational methods have been developed to predict regulatory effects of non-coding variants. An emerging approach to variant annotation is deep learning models that predict regulatory functions from DNA sequences alone. While such models have been trained on large publicly available dataset such as ENCODE, neuropsychiatric trait-related cell types are under-represented in these datasets, thus there is an urgent need of better tools and resources to annotate variant functions in such cellular contexts. To fill this gap, we collected a large collection of neurodevelopment-related cell/tissue types, and trained deep Convolutional Neural Networks (ResNet) using such data. Furthermore, our model, called MetaChrom, borrows information from public epigenomic consortium to improve the accuracy via transfer learning. We show that MetaChrom is substantially better in predicting experimentally determined chromatin accessibility variants than popular variant annotation tools such as CADD and delta-SVM. By combining GWAS data with MetaChrom predictions, we prioritized 31 SNPs for Schizophrenia, suggesting potential risk genes and the biological contexts where they act. In summary, MetaChrom provides functional annotations of any DNA variants in the neuro-development context and the general method of MetaChrom can also be extended to other disease-related cell or tissue types.
Boqiao Lai, Sheng Qian, Hanwei Zhang 0004, Anna A. Kozlova, Jubao Duan, Jinbo Xu
PLoS Comput. Biol.7
2021 XOR-CD: Linearly Convergent Constrained Structure Generation
abstract
We propose XOR-Contrastive Divergence learning (XOR-CD), a provable approach for constrained structure generation, which remains difficult for state-of-the-art neural network and constraint reasoning approaches. XOR-CD harnesses XOR-Sampling to generate samples from the model distribution in CD learning and is guaranteed to generate valid structures. In addition, XOR-CD has a linear convergence rate towards the global maximum of the likelihood function within a vanishing constant in learning exponential family models. Constraint satisfaction enabled by XOR-CD also boosts its learning performance. Our real-world experiments on data-driven experimental design, dispatching route generation, and sequence-based protein homology detection demonstrate the superior performance of XOR-CD compared to baseline approaches in generating valid structures as well as capturing the inductive bias in the training set.
Jianzhu Ma, Jinbo Xu, Yexiang Xue
ICML3
2021 PALM: Probabilistic area loss Minimization for Protein Sequence Alignment
abstract
Protein sequence alignment is a fundamental problem in computational structure biology and popular for protein 3D structural prediction and protein homology detection. Most of the developed programs for detecting protein sequence alignments are based upon the likelihood information of amino acids and are sensitive to alignment noises. We present a novel method PALM for modeling pairwise protein structure alignments, using the area distance to reduce the biological measurement noise. PALM generatively learn the alignment of two protein sequences with probabilistic area distance objective, which can denoise the measurement errors contained in the ground-truth alignments. During learning, we show that the optimization is computationally efficient by estimating the gradients via dynamically sampling alignments. Empirically, we show that PALM can generate sequence alignments with higher precision and recall, as well as smaller area distance than the competing methods especially for long protein sequences and remote homologies. This study implies for learning over large-scale protein sequence alignment problems, one could potentially give PALM a try.
Nan Jiang 0012, Jianzhu Ma, Jian Peng 0001, Jinbo Xu, Yexiang Xue
UAI5
2021 ORFLine: a bioinformatic pipeline to prioritize small open reading frames identifies candidate secreted small proteins from lymphocytes
abstract
MOTIVATION: The annotation of small open reading frames (smORFs) of <100 codons (<300 nucleotides) is challenging due to the large number of such sequences in the genome. RESULTS: In this study, we developed a computational pipeline, which we have named ORFLine, that stringently identifies smORFs and classifies them according to their position within transcripts. We identified a total of 5744 unique smORFs in datasets from mouse B and T lymphocytes and systematically characterized them using ORFLine. We further searched smORFs for the presence of a signal peptide, which predicted known secreted chemokines as well as novel micropeptides. Four novel micropeptides show evidence of secretion and are therefore candidate mediators of immunoregulatory functions. AVAILABILITY AND IMPLEMENTATION: Freely available on the web at https://github.com/boboppie/ORFLine. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Fengyuan Hu, Louise S. Matheson, Manuel D. Díaz-Muñoz, Alexander Saveliev, Jinbo Xu, Martin Turner
Bioinform.6
2021 Improved protein model quality assessment by integrating sequential and pairwise features using deep learning
abstract
MOTIVATION: Accurately estimating protein model quality in the absence of experimental structure is not only important for model evaluation and selection but also useful for model refinement. Progress has been steadily made by introducing new features and algorithms (especially deep neural networks), but the accuracy of quality assessment (QA) is still not very satisfactory, especially local QA on hard protein targets. RESULTS: We propose a new single-model-based QA method ResNetQA for both local and global quality assessment. Our method predicts model quality by integrating sequential and pairwise features using a deep neural network composed of both 1D and 2D convolutional residual neural networks (ResNet). The 2D ResNet module extracts useful information from pairwise features such as model-derived distance maps, co-evolution information, and predicted distance potential from sequences. The 1D ResNet is used to predict local (global) model quality from sequential features and pooled pairwise information generated by 2D ResNet. Tested on the CASP12 and CASP13 datasets, our experimental results show that our method greatly outperforms existing state-of-the-art methods. Our ablation studies indicate that the 2D ResNet module and pairwise features play an important role in improving model quality assessment. AVAILABILITY AND IMPLEMENTATION: https://github.com/AndersJing/ResNetQA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiaoyang Jing, Jinbo Xu
Bioinform.2
2021 Study of real-valued distance prediction for protein structure prediction with deep learning
abstract
MOTIVATION: Inter-residue distance prediction by convolutional residual neural network (deep ResNet) has greatly advanced protein structure prediction. Currently, the most successful structure prediction methods predict distance by discretizing it into dozens of bins. Here, we study how well real-valued distance can be predicted and how useful it is for 3D structure modeling by comparing it with discrete-valued prediction based upon the same deep ResNet. RESULTS: Different from the recent methods that predict only a single real value for the distance of an atom pair, we predict both the mean and standard deviation of a distance and then fold a protein by the predicted mean and deviation. Our findings include: (i) tested on the CASP13 FM (free-modeling) targets, our real-valued distance prediction obtains 81% precision on top L/5 long-range contact prediction, much better than the best CASP13 results (70%); (ii) our real-valued prediction can predict correct folds for the same number of CASP13 FM targets as the best CASP13 group, despite generating only 20 decoys for each target; (iii) our method greatly outperforms a very new real-valued prediction method DeepDist in both contact prediction and 3D structure modeling and (iv) when the same deep ResNet is used, our real-valued distance prediction has 1-6% higher contact and distance accuracy than our own discrete-valued prediction, but less accurate 3D structure models. AVAILABILITY AND IMPLEMENTATION: https://github.com/j3xugit/RaptorX-3DModeling. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jinbo Xu
Bioinform.2
2021 Deep template-based protein structure prediction
abstract
MOTIVATION: Protein structure prediction has been greatly improved by deep learning, but most efforts are devoted to template-free modeling. But very few deep learning methods are developed for TBM (template-based modeling), a popular technique for protein structure prediction. TBM has been studied extensively in the past, but its accuracy is not satisfactory when highly similar templates are not available. RESULTS: This paper presents a new method NDThreader (New Deep-learning Threader) to address the challenges of TBM. NDThreader first employs DRNF (deep convolutional residual neural fields), which is an integration of deep ResNet (convolutional residue neural networks) and CRF (conditional random fields), to align a query protein to templates without using any distance information. Then NDThreader uses ADMM (alternating direction method of multipliers) and DRNF to further improve sequence-template alignments by making use of predicted distance potential. Finally, NDThreader builds 3D models from a sequence-template alignment by feeding it and sequence coevolution information into a deep ResNet to predict inter-atom distance distribution, which is then fed into PyRosetta for 3D model construction. Our experimental results show that NDThreader greatly outperforms existing methods such as CNFpred, HHpred, DeepThreader and CEthreader. NDThreader was blindly tested in CASP14 as a part of RaptorX server, which obtained the best average GDT score among all CASP14 servers on the 58 TBM targets.
Fandi Wu, Jinbo Xu
PLoS Comput. Biol.2
2020 Automated inference of Boolean models from molecular interaction maps using CaSQ
abstract
MOTIVATION: Molecular interaction maps have emerged as a meaningful way of representing biological mechanisms in a comprehensive and systematic manner. However, their static nature provides limited insights to the emerging behaviour of the described biological system under different conditions. Computational modelling provides the means to study dynamic properties through in silico simulations and perturbations. We aim to bridge the gap between static and dynamic representations of biological systems with CaSQ, a software tool that infers Boolean rules based on the topology and semantics of molecular interaction maps built with CellDesigner. RESULTS: We developed CaSQ by defining conversion rules and logical formulas for inferred Boolean models according to the topology and the annotations of the starting molecular interaction maps. We used CaSQ to produce executable files of existing molecular maps that differ in size, complexity and the use of Systems Biology Graphical Notation (SBGN) standards. We also compared, where possible, the manually built logical models corresponding to a molecular map to the ones inferred by CaSQ. The tool is able to process large and complex maps built with CellDesigner (either following SBGN standards or not) and produce Boolean models in a standard output format, Systems Biology Marked Up Language-qualitative (SBML-qual), that can be further analyzed using popular modelling tools. References, annotations and layout of the CellDesigner molecular map are retained in the obtained model, facilitating interoperability and model reusability. AVAILABILITY AND IMPLEMENTATION: The present tool is available online: https://lifeware.inria.fr/∼soliman/post/casq/ and distributed as a Python package under the GNU GPLv3 license. The code can be accessed here: https://gitlab.inria.fr/soliman/casq. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sara Sadat Aghamiri, Vidisha Singh, Aurélien Naldi, Tomás Helikar, Sylvain Soliman, Anna Niarakis, Jinbo Xu
Bioinform.7
2019 stance-Based Protein Folding Powered by Deep Learning
Jinbo Xu
RECOMB1
2019 PredMP: a web server for de novo prediction and visualization of membrane proteins
abstract
MOTIVATION: PredMP is the first web service, to our knowledge, that aims at de novo prediction of the membrane protein (MP) 3D structure followed by the embedding of the MP into the lipid bilayer for visualization. Our approach is based on a high-throughput Deep Transfer Learning (DTL) method that first predicts MP contacts by learning from non-MPs and then predicts the 3D model of the MP using the predicted contacts as distance restraints. This algorithm is derived from our previous Deep Learning (DL) method originally developed for soluble protein contact prediction, which has been officially ranked No. 1 in CASP12. The DTL framework in our approach overcomes the challenge that there are only a limited number of solved MP structures for training the deep learning model. There are three modules in the PredMP server: (i) The DTL framework followed by the contact-assisted folding protocol has already been implemented in RaptorX-Contact, which serves as the key module for 3D model generation; (ii) The 1D annotation module, implemented in RaptorX-Property, is used to predict the secondary structure and disordered regions; and (iii) the visualization module to display the predicted MPs embedded in the lipid bilayer guided by the predicted transmembrane topology. RESULTS: Tested on 510 non-redundant MPs, our server predicts correct folds for ∼290 MPs, which significantly outperforms existing methods. Tested on a blind and live benchmark CAMEO from September 2016 to January 2018, PredMP can successfully model all 10 MPs belonging to the hard category. AVAILABILITY AND IMPLEMENTATION: PredMP is freely accessed on the web at http://www.predmp.com. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sheng Wang 0001, Shiyang Fei, Zongan Wang, Yu Li 0006, Jinbo Xu, Feng Zhao 0004, Xin Gao 0001
Bioinform.5
2018 Deep Learning Reveals Many More Inter-protein Residue-Residue Contacts than Direct Coupling Analysis
Tianming Zhou, Sheng Wang 0001, Jinbo Xu
RECOMB3
2018 Predicting protein-protein interactions through sequence-based deep learning
abstract
Motivation: High-throughput experimental techniques have produced a large amount of protein-protein interaction (PPI) data, but their coverage is still low and the PPI data is also very noisy. Computational prediction of PPIs can be used to discover new PPIs and identify errors in the experimental PPI data. Results: We present a novel deep learning framework, DPPI, to model and predict PPIs from sequence information alone. Our model efficiently applies a deep, Siamese-like convolutional neural network combined with random projection and data augmentation to predict PPIs, leveraging existing high-quality experimental PPI data and evolutionary information of a protein pair under prediction. Our experimental results show that DPPI outperforms the state-of-the-art methods on several benchmarks in terms of area under precision-recall curve (auPR), and computationally is more efficient. We also show that DPPI is able to predict homodimeric interactions where other methods fail to work accurately, and the effectiveness of DPPI in specific applications such as predicting cytokine-receptor binding affinities. Availability and implementation: Predicting protein-protein interactions through sequence-based deep learning): https://github.com/hashemifar/DPPI/. Supplementary information: Supplementary data are available at Bioinformatics online.
Somaye Hashemifar, Behnam Neyshabur, Aly Azeem Khan, Jinbo Xu
Bioinform.4
2018 Protein threading using residue co-variation and deep learning
abstract
Motivation: Template-based modeling, including homology modeling and protein threading, is a popular method for protein 3D structure prediction. However, alignment generation and template selection for protein sequences without close templates remain very challenging. Results: We present a new method called DeepThreader to improve protein threading, including both alignment generation and template selection, by making use of deep learning (DL) and residue co-variation information. Our method first employs DL to predict inter-residue distance distribution from residue co-variation and sequential information (e.g. sequence profile and predicted secondary structure), and then builds sequence-template alignment by integrating predicted distance information and sequential features through an ADMM algorithm. Experimental results suggest that predicted inter-residue distance is helpful to both protein alignment and template selection especially for protein sequences without very close templates, and that our method outperforms currently popular homology modeling method HHpred and threading method CNFpred by a large margin and greatly outperforms the latest contact-assisted protein threading method EigenTHREADER. Availability and implementation: http://raptorx.uchicago.edu/. Supplementary information: Supplementary data are available at Bioinformatics online.
Jianwei Zhu, Sheng Wang 0001, Dongbo Bu, Jinbo Xu
Bioinform.4
2018 RaptorX-Angle: real-value prediction of protein backbone dihedral angles through a hybrid method of clustering and deep learning
abstract
BACKGROUND: Protein dihedral angles provide a detailed description of protein local conformation. Predicted dihedral angles can be used to narrow down the conformational space of the whole polypeptide chain significantly, thus aiding protein tertiary structure prediction. However, direct angle prediction from sequence alone is challenging. RESULTS: In this article, we present a novel method (named RaptorX-Angle) to predict real-valued angles by combining clustering and deep learning. Tested on a subset of PDB25 and the targets in the latest two Critical Assessment of protein Structure Prediction (CASP), our method outperforms the existing state-of-art method SPIDER2 in terms of Pearson Correlation Coefficient (PCC) and Mean Absolute Error (MAE). Our result also shows approximately linear relationship between the real prediction errors and our estimated bounds. That is, the real prediction error can be well approximated by our estimated bounds. CONCLUSIONS: Our study provides an alternative and more accurate prediction of dihedral angles, which may facilitate protein structure prediction and functional study.
Yujuan Gao, Sheng Wang 0001, Minghua Deng, Jinbo Xu
BMC Bioinform.4
2017 Folding Membrane Proteins by Deep Transfer Learning
Zhen Li 0026, Sheng Wang 0001, Yizhou Yu, Jinbo Xu
RECOMB4
2017 Accurate De Novo Prediction of Protein Contact Map by Ultra-Deep Learning Model
abstract
MOTIVATION: Protein contacts contain key information for the understanding of protein structure and function and thus, contact prediction from sequence is an important problem. Recently exciting progress has been made on this problem, but the predicted contacts for proteins without many sequence homologs is still of low quality and not very useful for de novo structure prediction. METHOD: This paper presents a new deep learning method that predicts contacts by integrating both evolutionary coupling (EC) and sequence conservation information through an ultra-deep neural network formed by two deep residual neural networks. The first residual network conducts a series of 1-dimensional convolutional transformation of sequential features; the second residual network conducts a series of 2-dimensional convolutional transformation of pairwise information including output of the first residual network, EC information and pairwise potential. By using very deep residual networks, we can accurately model contact occurrence patterns and complex sequence-structure relationship and thus, obtain higher-quality contact prediction regardless of how many sequence homologs are available for proteins in question. RESULTS: Our method greatly outperforms existing methods and leads to much more accurate contact-assisted folding. Tested on 105 CASP11 targets, 76 past CAMEO hard targets, and 398 membrane proteins, the average top L long-range prediction accuracy obtained by our method, one representative EC method CCMpred and the CASP11 winner MetaPSICOV is 0.47, 0.21 and 0.30, respectively; the average top L/10 long-range accuracy of our method, CCMpred and MetaPSICOV is 0.77, 0.47 and 0.59, respectively. Ab initio folding using our predicted contacts as restraints but without any force fields can yield correct folds (i.e., TMscore>0.6) for 203 of the 579 test proteins, while that using MetaPSICOV- and CCMpred-predicted contacts can do so for only 79 and 62 of them, respectively. Our contact-assisted models also have much better quality than template-based models especially for membrane proteins. The 3D models built from our contact prediction have TMscore>0.5 for 208 of the 398 membrane proteins, while those from homology modeling have TMscore>0.5 for only 10 of them. Further, even if trained mostly by soluble proteins, our deep learning method works very well on membrane proteins. In the recent blind CAMEO benchmark, our fully-automated web server implementing this method successfully folded 6 targets with a new fold and only 0.3L-2.3L effective sequence homologs, including one β protein of 182 residues, one α+β protein of 125 residues, one α protein of 140 residues, one α protein of 217 residues, one α/β of 260 residues and one α protein of 462 residues. Our method also achieved the highest F1 score on free-modeling targets in the latest CASP (Critical Assessment of Structure Prediction), although it was not fully implemented back then. AVAILABILITY: http://raptorx.uchicago.edu/ContactMap/.
Sheng Wang 0001, Zhen Li 0026, Jinbo Xu
PLoS Comput. Biol.5
2017 A Sparse Learning Framework for Joint Effect Analysis of Copy Number Variants
abstract
Copy number variants (CNVs), including large deletions and duplications, represent an unbalanced change of DNA segments. Abundant in human genomes, CNVs contribute to a large proportion of human genetic diversity, with impact on many human phenotypes. Although recent advances in genetic studies have shed light on the impact of individual CNVs on different traits, the analysis of joint effect of multiple interactive CNVs lags behind from many perspectives. A primary reason is that the large number of CNV combinations and interactions in the human genome make it computationally challenging to perform such joint analysis. To address this challenge, we developed a novel sparse learning framework that combines sparse learning with biological networks to identify interacting CNVs with joint effect on particular traits. We showed that our approach performs well in identifying CNVs with joint phenotypic effect using simulated data. Applied to a real human genomic dataset from the 1,000 Genomes Project, our approach identified multiple CNVs that collectively contribute to population differentiation. We found a set of multiple CNVs that have joint effect in different populations, and affect gene expression differently in distinct populations. These results provided a collection of CNVs that likely have downstream biomedical implications in individuals from diverse population backgrounds.
Zhiyong Wang 0005, Benika Hall, Jinbo Xu, Xinghua Shi
IEEE ACM Trans. Comput. Biol. Bioinform.3
2016 AUC-Maximized Deep Convolutional Neural Fields for Protein Sequence Labeling
Sheng Wang 0001, Jinbo Xu
ECML/PKDD (2)3
2016 Joint Alignment of Multiple Protein-Protein Interaction Networks via Convex Optimization
Somaye Hashemifar, Qixing Huang, Jinbo Xu
RECOMB3
2016 ModuleAlign: module-based global alignment of protein-protein interaction networks
abstract
MOTIVATION: As an increasing amount of protein-protein interaction (PPI) data becomes available, their computational interpretation has become an important problem in bioinformatics. The alignment of PPI networks from different species provides valuable information about conserved subnetworks, evolutionary pathways and functional orthologs. Although several methods have been proposed for global network alignment, there is a pressing need for methods that produce more accurate alignments in terms of both topological and functional consistency. RESULTS: In this work, we present a novel global network alignment algorithm, named ModuleAlign, which makes use of local topology information to define a module-based homology score. Based on a hierarchical clustering of functionally coherent proteins involved in the same module, ModuleAlign employs a novel iterative scheme to find the alignment between two networks. Evaluated on a diverse set of benchmarks, ModuleAlign outperforms state-of-the-art methods in producing functionally consistent alignments. By aligning Pathogen-Human PPI networks, ModuleAlign also detects a novel set of conserved human genes that pathogens preferentially target to cause pathogenesis. AVAILABILITY: http://ttic.uchicago.edu/∼hashemifar/ModuleAlign.html CONTACT: [email protected] or j3xu.ttic.edu SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Somaye Hashemifar, Jianzhu Ma, Hammad Naveed, Stefan Canzar, Jinbo Xu
Bioinform.5
2016 AUCpreD: proteome-level protein disorder prediction by AUC-maximized deep convolutional neural fields
abstract
MOTIVATION: Protein intrinsically disordered regions (IDRs) play an important role in many biological processes. Two key properties of IDRs are (i) the occurrence is proteome-wide and (ii) the ratio of disordered residues is about 6%, which makes it challenging to accurately predict IDRs. Most IDR prediction methods use sequence profile to improve accuracy, which prevents its application to proteome-wide prediction since it is time-consuming to generate sequence profiles. On the other hand, the methods without using sequence profile fare much worse than using sequence profile. METHOD: This article formulates IDR prediction as a sequence labeling problem and employs a new machine learning method called Deep Convolutional Neural Fields (DeepCNF) to solve it. DeepCNF is an integration of deep convolutional neural networks (DCNN) and conditional random fields (CRF); it can model not only complex sequence-structure relationship in a hierarchical manner, but also correlation among adjacent residues. To deal with highly imbalanced order/disorder ratio, instead of training DeepCNF by widely used maximum-likelihood, we develop a novel approach to train it by maximizing area under the ROC curve (AUC), which is an unbiased measure for class-imbalanced data. RESULTS: Our experimental results show that our IDR prediction method AUCpreD outperforms existing popular disorder predictors. More importantly, AUCpreD works very well even without sequence profile, comparing favorably to or even outperforming many methods using sequence profile. Therefore, our method works for proteome-wide disorder prediction while yielding similar or better accuracy than the others. AVAILABILITY AND IMPLEMENTATION: http://raptorx2.uchicago.edu/StructurePropertyPred/predict/ CONTACT: [email protected], [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sheng Wang 0001, Jianzhu Ma, Jinbo Xu
Bioinform.3
2015 Inferring Block Structure of Graphical Models in Exponential Families
abstract
Learning the structure of a graphical model is a fundamental problem and it is used extensively to infer the relationship between random variables. In many real world applications, we usually have some prior knowledge about the underlying graph structure, such as degree distribution and block structure. In this paper, we propose a novel generative model for describing the block structure in general exponential families, and optimize it by an Expectation-Maximization(EM) algorithm with variational Bayes. Experimental results show that our method performs well on both synthetic and real data. Further, our method can predict overlapped block structure of a graphical model in general exponential families.
Jinbo Xu
AISTATS3
2015 Joint inference of tissue-specific networks with a scale free topology
abstract
High-throughput experimental techniques have produced an enormous number of gene expression profiles for various tissues of the human body. Tissue-specificity is a key component in reflecting the potentially different roles of proteins in diverse cell lineages. One way of understanding the tissue specificity is by reconstructing the tissue-specific co-expression networks (CENs) to analyze the correlation between genes. A few methods have been developed for estimating CENs, but it still remains challenging in terms of both accuracy and efficiency. In this paper we propose a new method, JointNet, for predicting tissue-specific co-expression networks. JointNet is exploiting the observation that, functionally related tissues have similar expression patterns and thus, similar networks. It uses different node penalties for hubs and non-hub nodes to accurately estimate the scale-free networks. Our experimental results show that the resulting tissue-specific CENs are accurate and that our method outperforms the current state of the art.
Somaye Hashemifar, Behnam Neyshabur, Jinbo Xu
BIBM3
2015 Predicting diverse M-best protein contact maps
abstract
Protein contacts contain important information for protein structure and functional study, but contact prediction from sequence information remains very challenging. Recently evolutionary coupling (EC) analysis, which predicts contacts by detecting co-evolved residues (or columns) in a multiple sequence alignment (MSA), has made good progress due to better statistical assessment techniques and high-throughput sequencing. Existing EC analysis methods predict only a single contact map for a given protein, which may have low accuracy especially when the protein under prediction does not have a large number of sequence homologs. Analogous to ab initio folding that usually predicts a few possible 3D models for a given protein sequence, this paper presents a novel structure learning method that can predict a set of diverse contact maps for a given protein sequence, in which the best solution usually has much better accuracy than the first one. Our experimental tests show that for many test proteins, the best out of 5 solutions generated by our method has accuracy at least 0.1 better than the first one when the top L/5 or L/10 (L is the sequence length) predicted long-range contacts are evaluated, especially for protein families with a small number of sequence homologs. Our best solutions also have better quality than those generated by the two popular EC methods Evfold and PSICOV.
Jianzhu Ma, Sheng Wang 0001, Jinbo Xu
BIBM4
2015 Learning Scale-Free Networks by Dynamic Node Specific Degree Prior
abstract
Learning network structure underlying data is an important problem in machine learning. This paper presents a novel degree prior to study the inference of scale-free networks, which are widely used to model social and biological networks. In particular, this paper formulates scale-free network inference using Gaussian Graphical model (GGM) regularized by a node degree prior. Our degree prior not only promotes a desirable global degree distribution, but also exploits the estimated degree of an individual node and the relative strength of all the edges of a single node. To fulfill this, this paper proposes a ranking-based method to dynamically estimate the degree of a node, which makes the resultant optimization problem challenging to solve. To deal with this, this paper presents a novel ADMM (alternating direction method of multipliers) procedure. Our experimental results on both synthetic and real data show that our prior not only yields a scale-free network, but also produces many more correctly predicted edges than existing scale-free inducing prior, hub-inducing prior and the l_1 norm.
Qingming Tang, Jinbo Xu
ICML3
2015 A low-latency fine-grained dynamic shared cache management scheme for chip multi-processor
abstract
In order to utilize the shared last-level cache (LLC) in chip multi-processors (CMP) more efficiently, the partitioning of LLC resources among all cores should have the characteristics of low-latency for access, fine granularity for migration and simple hardware complexity for implementation. This paper proposes a dynamic LLC management scheme to achieve these goals. The proposed scheme migrates cache resources among different cores at the granularity of cache blocks, instead of ways. The quantity of victim cache blocks that each victim core can migrate to other target cores are related to an eviction probability, which are calculated according to the performance goal. Then the victim cache blocks for a target core is chosen from the nearest victim core who has non-zero eviction probability by introducing innovate E-Table structure in CMP. The eviction probabilities are updated periodically. With the help of E-Tables, the proposal achieves low-latency accesses by always keeping the required cache blocks near to the target cores. And fine granularity is guaranteed by maintaining an eviction probability for each core. In addition, only little additional hardware changes to traditional cache structure is required. Simulation results suggest significant performance improvements from 6.8% to 22.7% over related works.
Jinbo Xu, Zhengbin Pang
IPCCC1
2015 Learning structured densities via infinite dimensional exponential families
abstract
Learning the structure of a probabilistic graphical models is a well studied problem in the machine learning community due to its importance in many applications. Current approaches are mainly focused on learning the structure under restrictive parametric assumptions, which limits the applicability of these methods. In this paper, we study the problem of estimating the structure of a probabilistic graphical model without assuming a particular parametric model. We consider probabilities that are members of an infinite dimensional exponential family, which is parametrized by a reproducing kernel Hilbert space (RKHS) H and its kernel $k$. One difficulty in learning nonparametric densities is evaluation of the normalizing constant. In order to avoid this issue, our procedure minimizes the penalized score matching objective. We show how to efficiently minimize the proposed objective using existing group lasso solvers. Furthermore, we prove that our procedure recovers the graph structure with high-probability under mild conditions. Simulation studies illustrate ability of our procedure to recover the true graph structure without the knowledge of the data generating process.
Mladen Kolar, Jinbo Xu
NIPS3
2015 Exact Hybrid Covariance Thresholding for Joint Graphical Lasso
Qingming Tang, Jian Peng 0001, Jinbo Xu
ECML/PKDD (2)4
2015 Protein Contact Prediction by Integrating Joint Evolutionary Coupling Analysis and Supervised Learning
Jianzhu Ma, Sheng Wang 0001, Zhiyong Wang 0005, Jinbo Xu
RECOMB4
2015 Structure Learning Constrained by Node-Specific Degree Distribution
Jianzhu Ma, Feng Zhao 0004, Jinbo Xu
UAI3
2015 Protein contact prediction by integrating joint evolutionary coupling analysis and supervised learning
abstract
MOTIVATION: Protein contact prediction is important for protein structure and functional study. Both evolutionary coupling (EC) analysis and supervised machine learning methods have been developed, making use of different information sources. However, contact prediction is still challenging especially for proteins without a large number of sequence homologs. RESULTS: This article presents a group graphical lasso (GGL) method for contact prediction that integrates joint multi-family EC analysis and supervised learning to improve accuracy on proteins without many sequence homologs. Different from existing single-family EC analysis that uses residue coevolution information in only the target protein family, our joint EC analysis uses residue coevolution in both the target family and its related families, which may have divergent sequences but similar folds. To implement this, we model a set of related protein families using Gaussian graphical models and then coestimate their parameters by maximum-likelihood, subject to the constraint that these parameters shall be similar to some degree. Our GGL method can also integrate supervised learning methods to further improve accuracy. Experiments show that our method outperforms existing methods on proteins without thousands of sequence homologs, and that our method performs better on both conserved and family-specific contacts. AVAILABILITY AND IMPLEMENTATION: See http://raptorx.uchicago.edu/ContactMap/ for a web server implementing the method. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jianzhu Ma, Sheng Wang 0001, Zhiyong Wang 0005, Jinbo Xu
Bioinform.4
2014 Adaptive Variable Clustering in Gaussian Graphical Models
abstract
Gaussian graphical models (GGMs) are widely-used to describe the relationship between random variables. In many real-world applications, GGMs have a block structure in the sense that the variables can be clustered into groups so that inter-group correlation is much weaker than intra-group correlation. We present a novel nonparametric Bayesian generative model for such a block-structured GGM and an efficient inference algorithm to find the clustering of variables in this GGM by combining a Gibbs sampler and a split-merge Metropolis-Hastings algorithm. Experimental results show that our method performs well on both synthetic and real data. In particular, our method outperforms generic clustering algorithms and can automatically identify the true number of clusters.
Yuancheng Zhu, Jinbo Xu
AISTATS3
2014 Low-latency last-level cache structure based on grouped cores in Chip Multi-Processor
abstract
Last-Level Cache (LLC) plays an important role in Chip Multi-Processor (CMP). The objective of this work is to optimize the structure and management strategy of LLC. Based on 8-core CMP, a LLC structure based on grouped cores is proposed, where 8 cores are divided into 4 groups. All LLC resources are classified into three types, which are fixed private cache, dynamic private cache and dynamic shared cache. The layout of the LLC structure and the corresponding dynamic partitioning strategy are designed to achieve low access latency and high efficiency. Experimental results on full-system simulator suggest that the proposed structure and method are able to reduce the access latency by 2% to 12% compared with previous works, such as tiled structure, cache-centered structure and core-centered structure. Consequently, performance measured by IPC is improved up to 7%. The contribution of this paper is useful for CMP performance, and applies to not only 8-core CMP but also all small-scale CMPs.
Jinbo Xu, Kefei Wang, Zhengbin Pang
IPCCC1
2014 MRFalign: Protein Homology Detection through Alignment of Markov Random Fields
Jianzhu Ma, Sheng Wang 0001, Zhiyong Wang 0005, Jinbo Xu
RECOMB4
2014 HubAlign: an accurate and efficient method for global alignment of protein-protein interaction networks
abstract
MOTIVATION: High-throughput experimental techniques have produced a large amount of protein-protein interaction (PPI) data. The study of PPI networks, such as comparative analysis, shall benefit the understanding of life process and diseases at the molecular level. One way of comparative analysis is to align PPI networks to identify conserved or species-specific subnetwork motifs. A few methods have been developed for global PPI network alignment, but it still remains challenging in terms of both accuracy and efficiency. RESULTS: This paper presents a novel global network alignment algorithm, denoted as HubAlign, that makes use of both network topology and sequence homology information, based upon the observation that topologically important proteins in a PPI network usually are much more conserved and thus, more likely to be aligned. HubAlign uses a minimum-degree heuristic algorithm to estimate the topological and functional importance of a protein from the global network topology information. Then HubAlign aligns topologically important proteins first and gradually extends the alignment to the whole network. Extensive tests indicate that HubAlign greatly outperforms several popular methods in terms of both accuracy and efficiency, especially in detecting functionally similar proteins. AVAILABILITY: HubAlign is available freely for non-commercial purposes at http://ttic.uchicago.edu/∼hashemifar/software/HubAlign.zip. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Somaye Hashemifar, Jinbo Xu
Bioinform.2
2014 MRFalign: Protein Homology Detection through Alignment of Markov Random Fields
abstract
Sequence-based protein homology detection has been extensively studied and so far the most sensitive method is based upon comparison of protein sequence profiles, which are derived from multiple sequence alignment (MSA) of sequence homologs in a protein family. A sequence profile is usually represented as a position-specific scoring matrix (PSSM) or an HMM (Hidden Markov Model) and accordingly PSSM-PSSM or HMM-HMM comparison is used for homolog detection. This paper presents a new homology detection method MRFalign, consisting of three key components: 1) a Markov Random Fields (MRF) representation of a protein family; 2) a scoring function measuring similarity of two MRFs; and 3) an efficient ADMM (Alternating Direction Method of Multipliers) algorithm aligning two MRFs. Compared to HMM that can only model very short-range residue correlation, MRFs can model long-range residue interaction pattern and thus, encode information for the global 3D structure of a protein family. Consequently, MRF-MRF comparison for remote homology detection shall be much more sensitive than HMM-HMM or PSSM-PSSM comparison. Experiments confirm that MRFalign outperforms several popular HMM or PSSM-based methods in terms of both alignment accuracy and remote homology detection and that MRFalign works particularly well for mainly beta proteins. For example, tested on the benchmark SCOP40 (8353 proteins) for homology detection, PSSM-PSSM and HMM-HMM succeed on 48% and 52% of proteins, respectively, at superfamily level, and on 15% and 27% of proteins, respectively, at fold level. In contrast, MRFalign succeeds on 57.3% and 42.5% of proteins at superfamily and fold level, respectively. This study implies that long-range residue interaction patterns are very helpful for sequence-based homology detection. The software is available for download at http://raptorx.uchicago.edu/download/. A summary of this paper appears in the proceedings of the RECOMB 2014 conference, April 2-5.
Jianzhu Ma, Sheng Wang 0001, Zhiyong Wang 0005, Jinbo Xu
PLoS Comput. Biol.4
2014 FPGA Implementation of a Special-Purpose VLIW Structure for Double-Precision Elementary Function
abstract
In the current article, the capability and flexibility of field programmable gate-arrays (FPGAs) to implement IEEE-754 double-precision floating-point elementary functions are explored. To perform various elementary functions on the unified hardware efficiently, we propose a special-purpose very long instruction word (VLIW) processor, called DP_VELP. This processor is equipped with multiple basic units, and its performance is improved through an explicitly parallel technique. Pipelined evaluation of polynomial approximation with Estrin's scheme is proposed, by scheduling basic components in an optimal order to avoid data hazard stalls and achieve minimal latency. The custom VLIW processor can achieve high scalability. Under the control of specific VLIW instructions, the basic units are combined into special-purpose hardware for elementary functions. Common elementary functions are presented as examples to illustrate the design of elementary function in DP_VELP in detail. Minimax approximation scheme is used to reduce degree of polynomial. Compromise between the size of lookup table and the latency is discussed, and the internal precision is carefully planned to guarantee accuracy of the result. Finally, we create a prototype of the DP_VELP unit and an FPGA accelerator based on the DP_VELP unit on a Xilinx XC6VLX760 FPGA chip to implement the SGP4/SDP4 application. Compared with previous researches, the proposed design can achieve low latency with a reasonable amount of resources and evaluate a variety of elementary functions with the unified hardware to satisfy the demands in scientific applications. Experimental results show that the proposed design guarantees more than 99% of correct rounding. Moreover, the SGP4/SDP4 accelerator, which is equipped with 39 DP_VELP units and runs at 200 MHz, outperforms the parallel software approach with hyper-thread technology on an Intel Xeon Quad E5620 CPU at 2.40 GHz by a factor of 7X.
Yuanwu Lei, Lei Guo 0029, Yong Dou, Sheng Ma, Jinbo Xu
ACM Trans. Reconfigurable Technol. Syst.5
2013 Estimating the Partition Function of Graphical Models Using Langevin Importance Sampling
abstract
Graphical models are powerful in modeling a variety of applications. Computing the partition function of a graphical model is a typical inference problem and known as an NP-hard problem for general graphs. A few sampling algorithms like MCMC, Simulated Annealing Sampling (SAS), Annealed Importance Sampling (AIS) are developed to address this challenging problem. This paper describes a Langevin Importance Sampling (LIS) algorithm to compute the partition function of a graphical model. LIS first performs a random walk in the configuration-temperature space guided by the Langevin equation and then estimates the partition function using all the samples generated during the random walk, as opposed to the other configuration-temperature sampling methods, which uses only the samples at a specific temperature. Experimental results show that LIS can obtain much more accurate partition function than the others tested on several different types of graphical models. LIS performs especially well on relatively large graph models or those with a large number of local optima.
Jianzhu Ma, Jian Peng 0001, Sheng Wang 0001, Jinbo Xu
AISTATS4
2013 Protein threading using context-specific alignment potential
abstract
MOTIVATION: Template-based modeling, including homology modeling and protein threading, is the most reliable method for protein 3D structure prediction. However, alignment errors and template selection are still the main bottleneck for current template-base modeling methods, especially when proteins under consideration are distantly related. RESULTS: We present a novel context-specific alignment potential for protein threading, including alignment and template selection. Our alignment potential measures the log-odds ratio of one alignment being generated from two related proteins to being generated from two unrelated proteins, by integrating both local and global context-specific information. The local alignment potential quantifies how well one sequence residue can be aligned to one template residue based on context-specific information of the residues. The global alignment potential quantifies how well two sequence residues can be placed into two template positions at a given distance, again based on context-specific information. By accounting for correlation among a variety of protein features and making use of context-specific information, our alignment potential is much more sensitive than the widely used context-independent or profile-based scoring function. Experimental results confirm that our method generates significantly better alignments and threading results than the best profile-based methods on several large benchmarks. Our method works particularly well for distantly related proteins or proteins with sparse sequence profiles because of the effective integration of context-specific, structure and global information. AVAILABILITY: http://raptorx.uchicago.edu/download/.
Jianzhu Ma, Sheng Wang 0001, Feng Zhao 0004, Jinbo Xu
Bioinform.4
2013 Predicting protein contact map using evolutionary and physical constraints by integer programming
abstract
MOTIVATION: Protein contact map describes the pairwise spatial and functional relationship of residues in a protein and contains key information for protein 3D structure prediction. Although studied extensively, it remains challenging to predict contact map using only sequence information. Most existing methods predict the contact map matrix element-by-element, ignoring correlation among contacts and physical feasibility of the whole-contact map. A couple of recent methods predict contact map by using mutual information, taking into consideration contact correlation and enforcing a sparsity restraint, but these methods demand for a very large number of sequence homologs for the protein under consideration and the resultant contact map may be still physically infeasible. RESULTS: This article presents a novel method PhyCMAP for contact map prediction, integrating both evolutionary and physical restraints by machine learning and integer linear programming. The evolutionary restraints are much more informative than mutual information, and the physical restraints specify more concrete relationship among contacts than the sparsity restraint. As such, our method greatly reduces the solution space of the contact map matrix and, thus, significantly improves prediction accuracy. Experimental results confirm that PhyCMAP outperforms currently popular methods no matter how many sequence homologs are available for the protein under consideration. AVAILABILITY: http://raptorx.uchicago.edu.
Zhiyong Wang 0005, Jinbo Xu
Bioinform.2
2013 VLIW coprocessor for IEEE-754 quadruple-precision elementary functions
abstract
In this article, a unified VLIW coprocessor, based on a common group of atomic operation units, for Quad arithmetic and elementary functions (QP_VELP) is presented. The explicitly parallel scheme of VLIW instruction and Estrin's evaluation scheme for polynomials are used to improve the performance. A two-level VLIW instruction RAM scheme is introduced to achieve high scalability and customizability, even for more complex key program kernels. Finally, the Quad arithmetic accelerator (QAA) with the QP_VELP array is implemented on ASIC. Compared with hyper-thread software implementation on an Intel Xeon E5620, QAA with 8 QP_VELP units achieves improvement by a factor of 18X.
Yuanwu Lei, Yong Dou, Lei Guo 0029, Jinbo Xu, Jie Zhou 0007, Yazhuo Dong
ACM Trans. Archit. Code Optim.4
2012 A conditional neural fields model for protein threading
abstract
MOTIVATION: Alignment errors are still the main bottleneck for current template-based protein modeling (TM) methods, including protein threading and homology modeling, especially when the sequence identity between two proteins under consideration is low (<30%). RESULTS: We present a novel protein threading method, CNFpred, which achieves much more accurate sequence-template alignment by employing a probabilistic graphical model called a Conditional Neural Field (CNF), which aligns one protein sequence to its remote template using a non-linear scoring function. This scoring function accounts for correlation among a variety of protein sequence and structure features, makes use of information in the neighborhood of two residues to be aligned, and is thus much more sensitive than the widely used linear or profile-based scoring function. To train this CNF threading model, we employ a novel quality-sensitive method, instead of the standard maximum-likelihood method, to maximize directly the expected quality of the training set. Experimental results show that CNFpred generates significantly better alignments than the best profile-based and threading methods on several public (but small) benchmarks as well as our own large dataset. CNFpred outperforms others regardless of the lengths or classes of proteins, and works particularly well for proteins with sparse sequence profiles due to the effective utilization of structure information. Our methodology can also be adapted to protein sequence alignment.
Jianzhu Ma, Jian Peng 0001, Sheng Wang 0001, Jinbo Xu
Bioinform.4
2011 A Parallel Processing Scheme for Large-Size Sliding-Window Applications
abstract
There exists large gap between the data input speed and processing speed in large-size sliding-window applications. To shorten this gap, a parallel processing scheme is proposed, which achieves high data reusability and parallelism with memory resources as few as possible and memory access control logics as simple as possible. This scheme combines the advantages of parallelism among different sliding-windows and parallelism among different data in a single window. For different windows, they are divided into groups and mapped into multiple processing elements. And for data in a single window, multi-module memory structure is introduced to buffer them, where module assignment and addressing scheme is designed for conflict-free parallel access. Experimental results on FPGA show that this work can improve the processing speed significantly without incurring too many memory resources and too complicated memory access control logics.
Jinbo Xu, Zhengbin Pang
HPCC2
2011 Alignment of distantly related protein structures: algorithm, bound and implications to homology modeling
abstract
MOTIVATION: Building an accurate alignment of a large set of distantly related protein structures is still very challenging. RESULTS: This article presents a novel method 3DCOMB that can generate a multiple structure alignment (MSA) with not only as many conserved cores as possible, but also high-quality pairwise alignments. 3DCOMB is unique in that it makes use of both local and global structure environments, combined by a statistical learning method, to accurately identify highly similar fragment blocks (HSFBs) among all proteins to be aligned. By extending the alignments of these HSFBs, 3DCOMB can quickly generate an accurate MSA without using progressive alignment. 3DCOMB significantly excels others in aligning distantly related proteins. 3DCOMB can also generate correct alignments for functionally similar regions among proteins of very different structures while many other MSA tools fail. 3DCOMB is useful for many real-world applications. In particular, it enables us to find out that there is still large improvement room for multiple template homology modeling while several other MSA tools fail to do so. AVAILABILITY: 3DCOMB is available at http://ttic.uchicago.edu/~jinbo/software.htm. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sheng Wang 0001, Jian Peng 0001, Jinbo Xu
Bioinform.3
2011 A conditional random fields method for RNA sequence-structure relationship modeling and conformation sampling
abstract
UNLABELLED: Accurate tertiary structures are very important for the functional study of non-coding RNA molecules. However, predicting RNA tertiary structures is extremely challenging, because of a large conformation space to be explored and lack of an accurate scoring function differentiating the native structure from decoys. The fragment-based conformation sampling method (e.g. FARNA) bears shortcomings that the limited size of a fragment library makes it infeasible to represent all possible conformations well. A recent dynamic Bayesian network method, BARNACLE, overcomes the issue of fragment assembly. In addition, neither of these methods makes use of sequence information in sampling conformations. Here, we present a new probabilistic graphical model, conditional random fields (CRFs), to model RNA sequence-structure relationship, which enables us to accurately estimate the probability of an RNA conformation from sequence. Coupled with a novel tree-guided sampling scheme, our CRF model is then applied to RNA conformation sampling. Experimental results show that our CRF method can model RNA sequence-structure relationship well and sequence information is important for conformation sampling. Our method, named as TreeFolder, generates a much higher percentage of native-like decoys than FARNA and BARNACLE, although we use the same simple energy function as BARNACLE. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhiyong Wang 0005, Jinbo Xu
Bioinform.2
2010 Protein 8-class secondary structure prediction using Conditional Neural Fields
abstract
Compared to the protein 3-class secondary structure (SS) prediction, the 8-class prediction gains less attention and is also much more challenging, especially for proteins with few sequence homologs. This paper presents a new probabilistic method for 8-class SS prediction using Conditional Neural Fields (CNFs), a recently-invented probabilistic graphical model. This CNF method not only models complex relationship between sequence features and SS, but also exploits interdependency among SS types of adjacent residues. In addition to sequence profiles, our method also makes use of non-evolutionary information for SS prediction. Tested on the CB513 and RS126 datasets, our method achieves Q8 accuracy 64.9% and 64.7%, respectively, which are much better than the SSpro8 web server (51.0% and 48.0%, respectively). Our method can also be used to predict other structure properties (e.g., solvent accessibility) of a protein or the SS of RNA.
Zhiyong Wang 0005, Jian Peng 0001, Jinbo Xu
BIBM4
2010 Low-homology protein threading
abstract
MOTIVATION: The challenge of template-based modeling lies in the recognition of correct templates and generation of accurate sequence-template alignments. Homologous information has proved to be very powerful in detecting remote homologs, as demonstrated by the state-of-the-art profile-based method HHpred. However, HHpred does not fare well when proteins under consideration are low-homology. A protein is low-homology if we cannot obtain sufficient amount of homologous information for it from existing protein sequence databases. RESULTS: We present a profile-entropy dependent scoring function for low-homology protein threading. This method will model correlation among various protein features and determine their relative importance according to the amount of homologous information available. When proteins under consideration are low-homology, our method will rely more on structure information; otherwise, homologous information. Experimental results indicate that our threading method greatly outperforms the best profile-based method HHpred and all the top CASP8 servers on low-homology proteins. Tested on the CASP8 hard targets, our threading method is also better than all the top CASP8 servers but slightly worse than Zhang-Server. This is significant considering that Zhang-Server and other top CASP8 servers use a combination of multiple structure-prediction techniques including consensus method, multiple-template modeling, template-free modeling and model refinement while our method is a classical single-template-based threading method without any post-threading refinement.
Jian Peng 0001, Jinbo Xu
Bioinform.2
2010 Fragment-free approach to protein folding using conditional neural fields
abstract
MOTIVATION: One of the major bottlenecks with ab initio protein folding is an effective conformation sampling algorithm that can generate native-like conformations quickly. The popular fragment assembly method generates conformations by restricting the local conformations of a protein to short structural fragments in the PDB. This method may limit conformations to a subspace to which the native fold does not belong because (i) a protein with really new fold may contain some structural fragments not in the PDB and (ii) the discrete nature of fragments may prevent them from building a native-like fold. Previously we have developed a conditional random fields (CRF) method for fragment-free protein folding that can sample conformations in a continuous space and demonstrated that this CRF method compares favorably to the popular fragment assembly method. However, the CRF method is still limited by its capability of generating conformations compatible with a sequence. RESULTS: We present a new fragment-free approach to protein folding using a recently invented probabilistic graphical model conditional neural fields (CNF). This new CNF method is much more powerful than CRF in modeling the sophisticated protein sequence-structure relationship and thus, enables us to generate native-like conformations more easily. We show that when coupled with a simple energy function and replica exchange Monte Carlo simulation, our CNF method can generate decoys much better than CRF on a variety of test proteins including the CASP8 free-modeling targets. In particular, our CNF method can predict a correct fold for T0496_D1, one of the two CASP8 targets with truly new fold. Our predicted model for T0496 is significantly better than all the CASP8 models.
Jian Peng 0001, Jinbo Xu
Bioinform.3
2009 Implementation of Rotation Invariant Multi-View Face Detection on FPGA
Jinbo Xu, Yong Dou, Yuxing Tang, Xiaodong Wang 0002
APPT1
2009 FPGA accelerating three QR decomposition algorithms in the unified pipelined framework
abstract
Many FPGA implementations for QR decomposition have been studied on small-scale matrix and all of them are presented individually. However to the best of our knowledge, there is no FPGA-based accelerator for large-scale QR decomposition. In this paper, we propose a unified FPGA accelerator structure for large-scale QR decomposition. To exploit the computational potential of FPGA, we introduce a fine-grained parallel algorithm for QR decomposition. A scalable linear array processing elements (PEs), which is the core component of the FPGA accelerator, is proposed to implement this algorithm. A total of 15 PEs can integrated into an Altera StratixII EP2S130F1020C5 on our self-designed board. Experimental results show that a factor of 4 speedup and the maximum powerperformance of 60.9 can be achieved compare to Pentium Dual CPU with double SSE thread.
Yong Dou, Jie Zhou 0007, Yuanwu Lei, Jinbo Xu
FPL5
2009 Conditional Neural Fields
abstract
Conditional random fields (CRF) are quite successful on sequence labeling tasks such as natural language processing and biological sequence analysis. CRF models use linear potential functions to represent the relationship between input features and outputs. However, in many real-world applications such as protein structure prediction and handwriting recognition, the relationship between input features and outputs is highly complex and nonlinear, which cannot be accurately modeled by a linear function. To model the nonlinear relationship between input features and outputs we propose Conditional Neural Fields (CNF), a new conditional probabilistic graphical model for sequence labeling. Our CNF model extends CRF by adding one (or possibly several) middle layer between input features and outputs. The middle layer consists of a number of hidden parameterized gates, each acting as a local neural network node or feature extractor to capture the nonlinear relationship between input features and outputs. Therefore, conceptually this CNF model is much more expressive than the linear CRF model. To better control the complexity of the CNF model, we also present a hyperparameter optimization procedure within the evidence framework. Experiments on two widely-used benchmarks indicate that this CNF model performs significantly better than a number of popular methods. In particular, our CNF model is the best among about ten machine learning methods for protein secondary tructure prediction and also among a few of the best methods for handwriting recognition.
Jian Peng 0001, Liefeng Bo, Jinbo Xu
NIPS3
2009 Boosting Protein Threading Accuracy
Jian Peng 0001, Jinbo Xu
RECOMB2
2009 A Probabilistic Graphical Model for Ab Initio Folding
Jian Peng 0001, Joe DeBartolo, Karl F. Freed, Tobin R. Sosnick, Jinbo Xu
RECOMB6
2009 Finding compact structural motifs
Dongbo Bu, Ming Li 0001, Shuaicheng Li 0001, Jianbo Qian, Jinbo Xu
Theor. Comput. Sci.5
2008 Finding Largest Well-Predicted Subset of Protein Structure Models
Shuaicheng Li 0001, Dongbo Bu, Jinbo Xu, Ming Li 0001
CPM3
2008 Double Precision Hybrid-Mode Floating-Point FPGA CORDIC Co-processor
abstract
FPGA chips have become a promising option for accelerating scientific applications, which involve many floating-point transcendental functions, such as sin, log, exp, sqrt and etc. In this paper, we present a 64-bit ANSI/IEEE floating-point CORDIC co-processor on FPGA, providing all known CORDIC functions. And there is no 64-bit CORDIC implementation on FPGA known to us. We propose a hybrid-mode CORDIC algorithm, combining hybrid rotation angle methods with argument reduction algorithm to reduce hardware area usage and meanwhile keep unlimited convergence domain for any floating-point inputs of the functions. Our hybrid-mode CORDIC co-processor is organized into three phases, argument reduction, CORDIC calculation and normalization with 69 pipeline stages for FPGA implementation. The synthesis results show the clock frequency can reach 173 MHz on Xilinx Virtex5 FPGA. Comparing to general-purpose microprocessor in three scientific program kernels, the CORDIC co-processor can achieve a maximum speedup of 49.3 times, 28.7 times in average.
Jie Zhou 0007, Yong Dou, Yuanwu Lei, Jinbo Xu, Yazhuo Dong
HPCC4
2008 Designing succinct structural alphabets
abstract
MOTIVATION: The 3D structure of a protein sequence can be assembled from the substructures corresponding to small segments of this sequence. For each small sequence segment, there are only a few more likely substructures. We call them the 'structural alphabet' for this segment. Classical approaches such as ROSETTA used sequence profile and secondary structure information, to predict structural fragments. In contrast, we utilize more structural information, such as solvent accessibility and contact capacity, for finding structural fragments. RESULTS: Integer linear programming technique is applied to derive the best combination of these sequence and structural information items. This approach generates significantly more accurate and succinct structural alphabets with more than 50% improvement over the previous accuracies. With these novel structural alphabets, we are able to construct more accurate protein structures than the state-of-art ab initio protein structure prediction programs such as ROSETTA. We are also able to reduce the Kolodny's library size by a factor of 8, at the same accuracy. AVAILABILITY: The online FRazor server is under construction.
Shuaicheng Li 0001, Dongbo Bu, Xin Gao 0001, Jinbo Xu, Ming Li 0001
ISMB4
2008 Rapid and Accurate Protein Side Chain Prediction with Local Backbone Information
Jing Zhang 0012, Xin Gao 0001, Jinbo Xu, Ming Li 0001
RECOMB3
2008 Graph algorithms for biological systems analysis
Bonnie Berger, Rohit Singh 0001, Jinbo Xu
SODA3
2007 Finding Compact Structural Motifs
Jianbo Qian, Shuaicheng Li 0001, Dongbo Bu, Ming Li 0001, Jinbo Xu
CPM5
2007 FPGA Accelerating Algorithms of Active Shape Model in People Tracking Applications
abstract
Algorithms of Active Shape Model, as one of the most popular methods for recognizing non-rigid objects, require huge computation power for real time people tracking. After analyzing the parallel characteristics of the algorithm, we propose a deep pipelined structure for accelerating the Active Shape Model algorithm. The computing engine is organized into a deep pipeline network composing of multiple floating-point arithmetic units, including adders, multipliers, dividers and SQRT etc. In the optimization of the memory efficiency for loading random data in large images during the step of local search, we propose an on-chip buffer scheme to eliminate random accesses to off-chip memory. Experimental results show that our FPGA implementation achieves over 15 times of speedup compared with the software implementation in Pentium 4 computer.
Jinbo Xu, Yong Dou, Xingming Zhou, Qiang Dou
DSD1
2007 Pairwise Global Alignment of Protein Interaction Networks by Matching Neighborhood Topology
Rohit Singh 0001, Jinbo Xu, Bonnie Berger
RECOMB2
2006 Robust and real-time automatic target recognition using partial hausdorff distance measure on reconfigurable hardware
abstract
This paper presents a high performance FPGA-based automatic target recognition system, which matches TV templates with an image efficiently. Theoretically, image matching algorithms based on partial Hausdorff distance (HD) are more tolerant of perturbations in the locations of pixel points than other algorithms, but they are too computationally expensive to be used in embedded systems. In order to solve this problem, we present a robust and real-time implementation of the image matching algorithm based on partial HD, taking advantage of the hardware resources offered by the FPGA chips. A parallel target recognition algorithm under constraints of limited embedded memory and limited memory bandwidth is proposed first. And then the system is organized as a coarse-grained pipeline containing three stages. Each stage is implemented in highly parallel fashion. The implementation of distance transform and template matching are described in detail. Experimental results show that our work outperforms related proposals. A speedup of almost 50 is achieved while compared with the software solution in PC (Pentium 4 2.8 GHz)
Jinbo Xu, Yong Dou
FPT1
2006 A Parameterized Algorithm for Protein Structure Alignment
Jinbo Xu, Feng Jiao, Bonnie Berger
RECOMB1
2006 Fast and accurate algorithms for protein side-chain packing
abstract
This article studies the protein side-chain packing problem using the tree-decomposition of a protein structure. To obtain fast and accurate protein side-chain packing, protein structures are modeled using a geometric neighborhood graph, which can be easily decomposed into smaller blocks. Therefore, the side-chain assignment of the whole protein can be assembled from the assignment of the small blocks. Although we will show that the side-chain packing problem is stillNP-hard, we can achieve a tree-decomposition-based globally optimal algorithm with time complexity ofO(Nnrottw+ 1)and several polynomial-time approximation schemes (PTAS), whereNis the number of residues contained in the protein,nrotthe average number of rotamers for each residue, andtw=O(N2/3logN) the treewidth of the protein structure graph. Experimental results indicate that after Goldstein dead-end elimination is conducted,nrotis very small andtwis equal to 3 or 4 most of the time. Based on the globally optimal algorithm, we developed a protein side-chain assignment program TreePack, which runs up to 90 times faster than SCWRL 3.0, a widely-used side-chain packing program, on some large test proteins in the SCWRL benchmark database and an average of five times faster on all the test proteins in this database. There are also some real-world instances that TreePack can solve but that SCWRL 3.0 cannot. The TreePack program is available at http://ttic.uchicago.edu/~jinbo/TreePack.htm.
Jinbo Xu, Bonnie Berger
J. ACM1
2005 Consensus fold recognition by predicted model quality
Jinbo Xu, Libo Yu, Ming Li 0001
APBC1
2005 Rapid Protein Side-Chain Packing via Tree Decomposition
Jinbo Xu
RECOMB1
2005 Fold Recognition by Predicted Alignment Accuracy
abstract
One of the key components in protein structure prediction by protein threading technique is to choose the best overall template for a given target sequence after all the optimal sequence-template alignments are generated. The chosen template should have the best alignment with the target sequence since the three-dimensional structure of the target sequence is built on the sequence-template alignment. The traditional method for template selection is called Z-score, which uses a statistical test to rank all the sequence-template alignments and then chooses the first-ranked template for the sequence. However, the calculation of Z-score is time-consuming and not suitable for genome-scale structure prediction. Z-scores are also hard to interpret when the threading scoring function is the weighted sum of several energy items of different physical meanings. This paper presents a Support Vector Machine (SVM) regression approach to directly predict the alignment accuracy of a sequence-template alignment, which is used to rank all the templates for a specific target sequence. Experimental results on a large-scale benchmark demonstrate that SVM regression performs much better than the composition-corrected Z-score method. SVM regression also runs much faster than the Z-score method.
Jinbo Xu
IEEE ACM Trans. Comput. Biol. Bioinform.1
2004 Optimizing Multiple Spaced Seeds for Homology Search
Jinbo Xu, Dan Brown 0001, Ming Li 0001, Bin Ma 0002
CPM1
2003 Speedup LP Approach to Protein Threading via Graph Reduction
Jinbo Xu
WABI1
2003 Approximation algorithms for NMR spectral peak assignment
Zhi-Zhong Chen, Tao Jiang 0001, Guohui Lin, Jianjun Wen, Dong Xu 0002, Jinbo Xu, Ying Xu 0001
Theor. Comput. Sci.6