EDBT 2026 Demo / reviewers in the wild / expert
Dongbo Bu
dblp:24/668
· DBLP profile ↗
54ranked-venue papers
1as first author
20since 2021 · last 2025
0000-0003-4119-4238ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 40 · 10 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Systems, architecture and hardware · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A survey on deep learning-based algorithms for the traveling salesman problemabstractAbstract This paper presents an overview of deep learning (DL)-based algorithms designed for solving the traveling salesman problem (TSP), categorizing them into four categories: end-to-end construction algorithms, end-to-end improvement algorithms, direct hybrid algorithms, and large language model (LLM)-based hybrid algorithms. We introduce the principles and methodologies of these algorithms, outlining their strengths and limitations through experimental comparisons. End-to-end construction algorithms employ neural networks to generate solutions from scratch, demonstrating rapid solving speed but often yielding subpar solutions. Conversely, end-to-end improvement algorithms iteratively refine initial solutions, achieving higher-quality outcomes but necessitating longer computation times. Direct hybrid algorithms directly integrate deep learning with heuristic algorithms, showcasing robust solving performance and generalization capability. LLM-based hybrid algorithms leverage LLMs to autonomously generate and refine heuristics, showing promising performance despite being in early developmental stages. In the future, further integration of deep learning techniques, particularly LLMs, with heuristic algorithms and advancements in interpretability and generalization will be pivotal trends in TSP algorithm design. These endeavors aim to tackle larger and more complex real-world instances while enhancing algorithm reliability and practicality. This paper offers insights into the evolving landscape of DL-based TSP solving algorithms and provides a perspective for future research directions. Jingyan Sui, Shizhe Ding, Xulin Huang, Boyang Xia, Zhenxin Ding, Liming Xu, Haicang Zhang, Chungong Yu, Dongbo Bu |
Frontiers Comput. Sci. | 11 |
| 2024 | iEDA: An Open-source infrastructure of EDAabstractBy leveraging the power of open-source software, the EDA tool offers a cost-effective and flexible solution for designers, researchers, and hobbyists alike. Open-source EDA promotes collaboration, innovation, and knowledge sharing within the EDA community. It emphasizes the role of the toolchain in accelerating the development of electronic systems, reducing design costs, and improving design quality. This paper presents an open-source EDA project, iEDA, aiming to build a basic infrastructure for EDA technology evolution and closing the industrial-academic gap in the EDA area. As the foundation for developing EDA tools and researching EDA algorithms and technologies, iEDA is mainly composed of file system, database, manager, operator and interface. To demonstrate the effectiveness of iEDA, we implement and tape out four chips of different scales (from 700k to 500M gates) on different process nodes (110nm and 28nm) with iEDA. iEDA is publicly available on the project home page https://github.com/OSCC-Project/iEDA. Zengrong Huang, Simin Tao, Zhipeng Huang 0009, Chunan Zhuang, Yihang Qiu, Guojie Luo, Huawei Li 0001, Haihua Shen, Mingyu Chen 0001, Dongbo Bu, Wenxing Zhu, Ye Cai 0001, Xiaoming Xiong, Yi Heng, Peng Zhang 0007, Bei Yu 0001, Biwei Xie, Yungang Bao |
ASPDAC | 13 |
| 2024 | Accurate Interpolation of Scattered Data Via Learning Relation GraphabstractInterpolation of scattered data is crucial across various domains, and neural networks have proved effective in developing accurate interpolators. While these neural network-based approaches excel in capturing data distributions, their failure to leverage inherent locality in computations can lead to overly dense correlation modeling. This might result in capturing spurious correlations and thereby affecting accuracy. To address these shortcomings, we propose a relation-aware interpolation framework named REIN. REIN uses a relational inference module to efficiently identify neighboring observed data points for each interpolation location, and integrate the relation graph as constraints into a neural interpolator. Experimental results on both synthetic and real-world datasets show that REIN outperforms the existing interpolation methods, even those employing heuristic local constraints. The analysis also suggests that compared with the widely-used heuristic local constraints, the learned local relation graphs exhibit improved adaptability and interpretability. Shizhe Ding, Boyang Xia, Jingyan Sui, Dongbo Bu |
ICASSP | 4 |
| 2024 | Multilevel Temporal-Spectral Fusion Network for Multivariate Time Series ClassificationabstractMultivariate time series classification (MTSC) plays important roles in a large variety of applications, including human activity recognition, acoustic scene classification, and electronic health. Most of the existing approaches exploit either temporal or spectral features of the input time series data but neglect the essential correlation between these two types of features. To address this limitation, we propose a multilevel temporal-spectral fusion network (called MTSFNet) that can effectively fuse both temporal and spectral features. The main steps of MTSFNet include: i) we first extract multilevel spectral signals from the input data using wavelet transform networks, which were further encoded into embedding vectors using a reduction encoder; ii) we fuse the temporal and multilevel spectral features to exploit the correlation between using cross-attention mechanism for classification. Experimental results on ten popular datasets from the UEA archive suggest that our method outperformed the state-of-the-art methods by an average accuracy improvement of 4.3%. Deep ablation experiments show that using multilevel wavelet transform networks can effectively improve the classification accuracy, where the three-level wavelet transform has the highest average classification accuracy, reaching 76.8%. This observation clearly demonstrates the advantages of our multilevel feature extraction and temporal-spectral fusion. We anticipate that the use of MTSFNet will greatly facilitate the analysis of time-series data in practice. Xulin Huang, Shizhe Ding, Jingyan Sui, Dongbo Bu |
IJCNN | 6 |
| 2024 | NeuralSteiner: Learning Steiner Tree for Overflow-avoiding Global Routing in Chip DesignabstractGlobal routing plays a critical role in modern chip design. The routing paths generated by global routers often form a rectilinear Steiner tree (RST). Recent advances from the machine learning community have shown the power of learning-based route generation; however, the yielded routing paths by the existing approaches often suffer from considerable overflow, thus greatly hindering their application in practice.
We propose NeuralSteiner, an accurate approach to overflow-avoiding global routing in chip design. The key idea of NeuralSteiner approach is to learn Steiner trees: we first predict the locations of highly likely Steiner points by adopting a neural network considering full-net spatial and overflow information, then select appropriate points by running a graph-based post-processing algorithm, and finally connect these points with the input pins to yield overflow-avoiding RSTs. NeuralSteiner offers two advantages over previous learning-based models. First, by using the learning scheme, NeuralSteiner ensures the connectivity of generated routes while significantly reducing congestion. Second, NeuralSteiner can effectively scale to large nets and transfer to unseen chip designs without any modifications or fine-tuning. Extensive experiments over public large-scale benchmarks reveal that, compared with the state-of-the-art deep generative methods, NeuralSteiner achieves up to a 99.8\% reduction in overflow while speeding up the generation and maintaining a slight wirelength loss within only 1.8\%. Zhisheng Zeng, Shizhe Ding, Jingyan Sui, Dongbo Bu |
NeurIPS | 6 |
| 2024 | GPCR-BSD: a database of binding sites of human G-protein coupled receptors under diverse statesabstractG-protein coupled receptors (GPCRs), the largest family of membrane proteins in human body, involve a great variety of biological processes and thus have become highly valuable drug targets. By binding with ligands (e.g., drugs), GPCRs switch between active and inactive conformational states, thereby performing functions such as signal transmission. The changes in binding pockets under different states are important for a better understanding of drug-target interactions. Therefore it is critical, as well as a practical need, to obtain binding sites in human GPCR structures. We report a database (called GPCR-BSD) that collects 127,990 predicted binding sites of 803 GPCRs under active and inactive states (thus 1,606 structures in total). The binding sites were identified from the predicted GPCR structures by executing three geometric-based pocket prediction methods, fpocket, CavityPlus and GHECOM. The server provides query, visualization, and comparison of the predicted binding sites for both GPCR predicted and experimentally determined structures recorded in PDB. We evaluated the identified pockets of 132 experimentally determined human GPCR structures in terms of pocket residue coverage, pocket center distance and redocking accuracy. The evaluation showed that fpocket and CavityPlus methods performed better and successfully predicted orthosteric binding sites in over 60% of the 132 experimentally determined structures. The GPCR Binding Site database is freely accessible at https://gpcrbs.bigdata.jcmsc.cn . This study not only provides a systematic evaluation of the commonly-used fpocket and CavityPlus methods for the first time but also meets the need for binding site information in GPCR studies. Xiaonong Li, Liangliang Zhou, Chungong Yu, Haicang Zhang, Dongbo Bu, Xinmiao Liang |
BMC Bioinform. | 7 |
| 2024 | NeuralGLS: learning to guide local search with graph convolutional network for the traveling salesman problem
Jingyan Sui, Shizhe Ding, Boyang Xia, Dongbo Bu |
Neural Comput. Appl. | 5 |
| 2024 | NIERT: Accurate Numerical Interpolation Through Unifying Scattered Data Representations Using Transformer EncoderabstractInterpolation for scattered data is a classical problem in numerical analysis, with a long history of theoretical and practical contributions. Recent advances have utilized deep neural networks to construct interpolators, exhibiting excellent and generalizable performance. However, they still fall short in two aspects:1) inadequate representation learning, resulting from separate embeddings of observed and target points in popular encoder-decoder frameworks and2) limited generalization power, caused by overlooking prior interpolation knowledge shared across different domains. To overcome these limitations, we present aNumericalInterpolation approach usingEncoderRepresentation ofTransformers (calledNIERT). On one hand, NIERT utilizes an encoder-only framework rather than the encoder-decoder structure. This way, NIERT can embed observed and target points into a unified encoder representation space, thus effectively exploiting the correlations among them and obtaining more precise representations. On the other hand, we propose to pre-train NIERT on large-scale synthetic mathematical functions to acquire prior interpolation knowledge, and transfer it to multiple interpolation domains with consistent performance gain. On both synthetic and real-world datasets, NIERT outperforms the existing approaches by a large margin, i.e., 4.3$\sim 14.3\times$lower MAE on TFRD subsets, and 1.7/1.8/8.7× lower MSE on Mathit/PhysioNet/PTV datasets. The source code of NIERT is available athttps://github.com/DingShizhe/NIERT. Shizhe Ding, Boyang Xia, Milong Ren, Dongbo Bu |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Accurate Interpolation for Scattered Data through Hierarchical Residual RefinementabstractAccurate interpolation algorithms are highly desired in various theoretical and engineering scenarios. Unlike the traditional numerical algorithms that have exact zero-residual constraints on observed points, the neural network-based interpolation methods exhibit non-zero residuals at these points. These residuals, which provide observations of an underlying residual function, can guide predicting interpolation functions, but have not been exploited by the existing approaches. To fill this gap, we propose Hierarchical INTerpolation Network (HINT), which utilizes the residuals on observed points to guide target function estimation in a hierarchical fashion. HINT consists of several sequentially arranged lightweight interpolation blocks. The first interpolation block estimates the main component of the target function, while subsequent blocks predict the residual components using observed points residuals of the preceding blocks. The main component and residual components are accumulated to form the final interpolation results. Furthermore, under the assumption that finer residual prediction requires a more focused attention range on observed points, we utilize hierarchical local constraints in correlation modeling between observed and target points. Extensive experiments demonstrate that HINT outperforms existing interpolation algorithms significantly in terms of interpolation accuracy across a wide variety of datasets, which underscores its potential for practical scenarios. Shizhe Ding, Boyang Xia, Dongbo Bu |
NeurIPS | 3 |
| 2023 | Predicting mutational effects on protein-protein binding via a side-chain diffusion probabilistic modelabstractMany crucial biological processes rely on networks of protein-protein interactions. Predicting the effect of amino acid mutations on protein-protein binding is important in protein engineering, including therapeutic discovery. However, the scarcity of annotated experimental data on binding energy poses a significant challenge for developing computational approaches, particularly deep learning-based methods. In this work, we propose SidechainDiff, a novel representation learning-based approach that leverages unlabelled experimental protein structures. SidechainDiff utilizes a Riemannian diffusion model to learn the generative process of side-chain conformations and can also give the structural context representations of mutations on the protein-protein interface. Leveraging the learned representations, we achieve state-of-the-art performance in predicting the mutational effects on protein-protein binding. Furthermore, SidechainDiff is the first diffusion-based generative model for side-chains, distinguishing it from prior efforts that have predominantly focused on the generation of protein backbone structures. Milong Ren, Chungong Yu, Dongbo Bu, Haicang Zhang |
NeurIPS | 5 |
| 2023 | EMNGly: predicting N-linked glycosylation sites using the language models for feature extractionabstractMOTIVATION: N-linked glycosylation is a frequently occurring post-translational protein modification that serves critical functions in protein folding, stability, trafficking, and recognition. Its involvement spans across multiple biological processes and alterations to this process can result in various diseases. Therefore, identifying N-linked glycosylation sites is imperative for comprehending the mechanisms and systems underlying glycosylation. Due to the inherent experimental complexities, machine learning and deep learning have become indispensable tools for predicting these sites. RESULTS: In this context, a new approach called EMNGly has been proposed. The EMNGly approach utilizes pretrained protein language model (Evolutionary Scale Modeling) and pretrained protein structure model (Inverse Folding Model) for features extraction and support vector machine for classification. Ten-fold cross-validation and independent tests show that this approach has outperformed existing techniques. And it achieves Matthews Correlation Coefficient, sensitivity, specificity, and accuracy of 0.8282, 0.9343, 0.8934, and 0.9143, respectively on a benchmark independent test set. Xiaoyang Hou, Dongbo Bu, Yaojun Wang, Shiwei Sun |
Bioinform. | 3 |
| 2023 | Accurate and efficient protein sequence design through learning concise local environment of residuesabstractMOTIVATION: Computational protein sequence design has been widely applied in rational protein engineering and increasing the design accuracy and efficiency is highly desired. RESULTS: Here, we present ProDESIGN-LE, an accurate and efficient approach to protein sequence design. ProDESIGN-LE adopts a concise but informative representation of the residue's local environment and trains a transformer to learn the correlation between local environment of residues and their amino acid types. For a target backbone structure, ProDESIGN-LE uses the transformer to assign an appropriate residue type for each position based on its local environment within this structure, eventually acquiring a designed sequence with all residues fitting well with their local environments. We applied ProDESIGN-LE to design sequences for 68 naturally occurring and 129 hallucinated proteins within 20 s per protein on average. The designed proteins have their predicted structures perfectly resembling the target structures with a state-of-the-art average TM-score exceeding 0.80. We further experimentally validated ProDESIGN-LE by designing five sequences for an enzyme, chloramphenicol O-acetyltransferase type III (CAT III), and recombinantly expressing the proteins in Escherichia coli. Of these proteins, three exhibited excellent solubility, and one yielded monomeric species with circular dichroism spectra consistent with the natural CAT III protein. AVAILABILITY AND IMPLEMENTATION: The source code of ProDESIGN-LE is available at https://github.com/bigict/ProDESIGN-LE. Bin Huang 0022, Tingwen Fan, Kaiyue Wang, Haicang Zhang, Chungong Yu, Shuyu Nie, Yangshuo Qi, Wei-Mou Zheng, Shiwei Sun, Huaiyi Yang, Dongbo Bu |
Bioinform. | 14 |
| 2023 | SASA-Net: A Spatial-Aware Self-Attention Mechanism for Building Protein 3D Structure Directly From Inter- Residue DistancesabstractProtein functions are tightly related to the fine details of their 3D structures. To understand protein structures, computational prediction approaches are highly needed. Recently, protein structure prediction has achieved considerable progresses mainly due to the increased accuracy of inter-residue distance estimation and the application of deep learning techniques. Most of the distance-based ab initio prediction approaches adopt a two-step diagram: constructing a potential function based on the estimated inter-residue distances, and then build a 3D structure that minimizes the potential function. These approaches have proven very promising; however, they still suffer from several limitations, especially the inaccuracies incurred by the handcrafted potential function. Here, we present SASA-Net, a deep learning-based approach that directly learns protein 3D structure from the estimated inter-residue distances. Unlike the existing approach simply representing protein structures as coordinates of atoms, SASA-Net represents protein structures using pose of residues, i.e., the coordinate system of each individual residue in which all backbone atoms of this residue are fixed. The key element of SASA-Net is a spatial-aware self-attention mechanism, which is able to adjust a residue's pose according to all other residues' features and the estimated distances between residues. By iteratively applying the spatial-aware self-attention mechanism, SASA-Net continuously improves the structure and finally acquires a structure with high accuracy. Using the CATH35 proteins as representatives, we demonstrate that SASA-Net is able to accurately and efficiently build structures from the estimated inter-residue distances. The high accuracy and efficiency of SASA-Net enables an end-to-end neural network model for protein structure prediction through combining SASA-Net and an neural network for inter-residue distance prediction. Source code of SASA-Net is available at https://github.com/gongtiansu/SASA-Net/. Tiansu Gong, Fusong Ju, Shiwei Sun, Dongbo Bu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2022 | Seq-SetNet: directly exploiting multiple sequence alignment for protein secondary structure predictionabstractMOTIVATION: Accurate prediction of protein structure relies heavily on exploiting multiple sequence alignment (MSA) for residue mutations and correlations as this information specifies protein tertiary structure. The widely used prediction approaches usually transform MSA into inter-mediate models, say position-specific scoring matrix or profile hidden Markov model. These inter-mediate models, however, cannot fully represent residue mutations and correlations carried by MSA; hence, an effective way to directly exploit MSAs is highly desirable. RESULTS: Here, we report a novel sequence set network (called Seq-SetNet) to directly and effectively exploit MSA for protein structure prediction. Seq-SetNet uses an 'encoding and aggregation' strategy that consists of two key elements: (i) an encoding module that takes a component homologue in MSA as input, and encodes residue mutations and correlations into context-specific features for each residue; and (ii) an aggregation module to aggregate the features extracted from all component homologues, which are further transformed into structural properties for residues of the query protein. As Seq-SetNet encodes each homologue protein individually, it could consider both insertions and deletions, as well as long-distance correlations among residues, thus representing more information than the inter-mediate models. Moreover, the encoding module automatically learns effective features and thus avoids manual feature engineering. Using symmetric aggregation functions, Seq-SetNet processes the homologue proteins as a sequence set, making its prediction results invariable to the order of these proteins. On popular benchmark sets, we demonstrated the successful application of Seq-SetNet to predict secondary structure and torsion angles of residues with improved accuracy and efficiency. AVAILABILITY AND IMPLEMENTATION: The code and datasets are available through https://github.com/fusong-ju/Seq-SetNet. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Fusong Ju, Jianwei Zhu, Guozheng Wei, Shiwei Sun, Wei-Mou Zheng, Dongbo Bu |
Bioinform. | 7 |
| 2022 | Mainstream encoding-decoding methods of DNA data storageabstractAbstract DNA storage is a new digital data storage technology based on specific encoding and decoding methods between 0 and 1 binary codes of digital data and A-T-C-G quaternary codes of DNAs, which and is expected to develop into a major data storage form in the future due to its advantages (such as high data density, long storage time, low energy consumption, convenience for carrying, concealed transportation and multiple encryptions). In this review, we mainly summarize the recent research advances of four main encoding and decoding methods of DNA storage technology: direct mapping method between 0 and 1 binary and A-T-C-G quaternary codes in early-stage, fountain code for higher logical storage density, inner and outer codes for random access DNA storage data, and CRISPR mediated in vivo DNA storage method. The first three encoding/decoding methods belong to in vitro DNA storage, representing the mainstream research and application in DNA storage. Their advantages and disadvantages are also reviewed: direct mapping method is easy and efficient, but has high error rate and low logical density; fountain code can achieve higher storage density without random access; inner and outer code has error-correction design to realize random access at the expense of logic density. This review provides important references and improved understanding of DNA storage methods. Development of efficient and accurate DNA storage encoding and decoding methods will play a very important and even decisive role in the transition of DNA storage from the laboratory to practical application, which may fundamentally change the information industry in the future. Guannan Ma, Di Wei, Peihan Wang, Cuidan Li, Bo Duan, Dongxin Yang, Dongbo Bu |
CCF Trans. High Perform. Comput. | 12 |
| 2021 | Learning 3-opt heuristics for traveling salesman problem via deep reinforcement learningabstractTraveling salesman problem (TSP) is a classical combinatorial optimization problem. As it represents a large number of important practical problems, it has received extensive studies and a great variety of algorithms have been proposed to solve it, including exact and heuristic algorithms. The success of heuristic algorithms relies heavily on the design of powerful heuristic rules, and most of the existing heuristic rules were manually designed by experienced experts to model their insights and observations on TSP instances and solutions. Recent studies have shown an alternative promising design strategy that directly learns heuristic rules from TSP instances without any manual interference. Here, we report an iterative improvement approach (called Neural-3-OPT) that solves TSP through automatically learning effective 3-opt heuristics via deep reinforcement learning. In the proposed approach, we adopt a pointer network to select 3 links from the current tour,and a feature-wise linear modulation network to select an appropriate way to reconnect the segments after removing the selected 3 links. We demonstrate that our approach achieves state-of-the-art performance on both real TSP instances and randomly-generated instances than, to the best of our knowledge, the existing neural network-based approaches. Jingyan Sui, Shizhe Ding, Liming Xu, Dongbo Bu |
ACML | 5 |
| 2021 | SASA-Net: A spatial-aware self-attention mechanism for building protein 3D structure directly from inter-residue distancesabstractProtein structure prediction has achieved considerable progresses mainly due to the increased accuracy of inter-residue distance estimation and the application of deep learning techniques.Most of the distance-based ab initio prediction approaches adopt a two-step diagram: constructing a potential function based on the estimated inter-residue distances, and then build a 3D structure that minimizes the potential function. These approaches have proven very promising; however, they still suffer from several limitations, especially the inaccuracies incurred by the handcrafted potential function.Here, we present SASA-Net, a deep learning-based approach that directly learns protein 3D structure from the estimated inter-residue distances. Unlike the existing approach simply representing protein structures as coordinates of atoms, SASA-Net represents protein structures using pose of residues, i.e., the coordinate system of each individual residue in which all backbone atoms of this residue are fixed. The key element of SASA-Net is a spatial-aware self-attention mechanism, which is able to adjust a residue’s pose according to all other residues’ features and the estimated distances between residues. By iteratively applying the spatial-aware self-attention mechanism, SASA-Net continuously improves the structure and finally acquires a structure with high accuracy. Using the CATH35 proteins as representatives, we demonstrate that SASA-Net is able to accurately and efficiently build structures from the estimated inter-residue distances. The high accuracy and efficiency of SASA-Net enables an end-to-end neural network model for protein structure prediction through combining SASA-Net and an neural network for inter-residue distance prediction. Source code of SASA-Net is available at https://github.con gongtiansu/SASA-Net/ Tiansu Gong, Fusong Ju, Shiwei Sun, Dongbo Bu |
BIBM | 4 |
| 2021 | Glycan immunogenicity prediction based on Graph neural networkabstractGlycans play important roles in a great variety of biological processes, and these roles are closely determined by the details of their structures. It becomes possible to acquire hidden features from glycan structures using deep learning method with the great progress in recent years. Unlike the linear chain of proteins and DNAs, branching is a unique feature of glycan structures, which makes it very difficult to directly apply deep learning models on glycans. Thus, how to comprehensively and efficiently describe glycans and use them as input to deep learning models still remains challenging. Here, a graph neural network (GNN) called GlyNet was used to obtain high-dimensional representation of glycans and predict their immunogenicity. Our method was applied in SugarBase, and it works more superiorly than the state-of-art method with accuracy increased from 91.7% to 95.6%. Yu Wang 0225, Meijie Hou, Yaojun Wang, Dongbo Bu, Chuncui Huang, Shiwei Sun |
BIBM | 5 |
| 2021 | Filling gaps of genome scaffolds via probabilistic searching optical maps against assembly graphabstractBACKGROUND: Optical maps record locations of specific enzyme recognition sites within long genome fragments. This long-distance information enables aligning genome assembly contigs onto optical maps and ordering contigs into scaffolds. The generated scaffolds, however, often contain a large amount of gaps. To fill these gaps, a feasible way is to search genome assembly graph for the best-matching contig paths that connect boundary contigs of gaps. The combination of searching and evaluation procedures might be "searching followed by evaluation", which is infeasible for long gaps, or "searching by evaluation", which heavily relies on heuristics and thus usually yields unreliable contig paths. RESULTS: We here report an accurate and efficient approach to filling gaps of genome scaffolds with aids of optical maps. Using simulated data from 12 species and real data from 3 species, we demonstrate the successful application of our approach in gap filling with improved accuracy and completeness of genome scaffolds. CONCLUSION: Our approach applies a sequential Bayesian updating technique to measure the similarity between optical maps and candidate contig paths. Using this similarity to guide path searching, our approach achieves higher accuracy than the existing "searching by evaluation" strategy that relies on heuristics. Furthermore, unlike the "searching followed by evaluation" strategy enumerating all possible paths, our approach prunes the unlikely sub-paths and extends the highly-probable ones only, thus significantly increasing searching efficiency. Guozheng Wei, Fusong Ju, Zhuozheng Shi, Shiwei Sun, Dongbo Bu |
BMC Bioinform. | 8 |
| 2021 | FALCON2: a web server for high-quality prediction of protein tertiary structuresabstractBACKGROUND: Accurate prediction of protein tertiary structures is highly desired as the knowledge of protein structures provides invaluable insights into protein functions. We have designed two approaches to protein structure prediction, including a template-based modeling approach (called ProALIGN) and an ab initio prediction approach (called ProFOLD). Briefly speaking, ProALIGN aligns a target protein with templates through exploiting the patterns of context-specific alignment motifs and then builds the final structure with reference to the homologous templates. In contrast, ProFOLD uses an end-to-end neural network to estimate inter-residue distances of target proteins and builds structures that satisfy these distance constraints. These two approaches emphasize different characteristics of target proteins: ProALIGN exploits structure information of homologous templates of target proteins while ProFOLD exploits the co-evolutionary information carried by homologous protein sequences. Recent progress has shown that the combination of template-based modeling and ab initio approaches is promising. RESULTS: In the study, we present FALCON2, a web server that integrates ProALIGN and ProFOLD to provide high-quality protein structure prediction service. For a target protein, FALCON2 executes ProALIGN and ProFOLD simultaneously to predict possible structures and selects the most likely one as the final prediction result. We evaluated FALCON2 on widely-used benchmarks, including 104 CASP13 (the 13th Critical Assessment of protein Structure Prediction) targets and 91 CASP14 targets. In-depth examination suggests that when high-quality templates are available, ProALIGN is superior to ProFOLD and in other cases, ProFOLD shows better performance. By integrating these two approaches with different emphasis, FALCON2 server outperforms the two individual approaches and also achieves state-of-the-art performance compared with existing approaches. CONCLUSIONS: By integrating template-based modeling and ab initio approaches, FALCON2 provides an easy-to-use and high-quality protein structure prediction service for the community and we expect it to enable insights into a deep understanding of protein functions. Lupeng Kong, Fusong Ju, Haicang Zhang, Shiwei Sun, Dongbo Bu |
BMC Bioinform. | 5 |
| 2020 | Mut-Detecter: An EGFR activating mutation type classification method with a deep convolutional neural networkabstractEpidermal growth factor receptor (EGFR) plays an essential role in tumor cell proliferation, angiogenesis and apoptosis inhibition; it is a crucial factor leading to cancer occurrence. For example, EGFR tyrosine kinase inhibitors in treating lung cancer patients have an excellent therapeutic effect. Targeted therapy based on EGFR gene mutation is one of the mainstream lung cancer treatment methods. Recent studies have shown that pulmonary nodules' characteristics are associated with the mutant status of EGFR, which provides the possibility of using CT images of patients with pulmonary nodules to predict the mutant status of EGFR. This study used the deep learning algorithm to establish the EGFR mutation type prediction model based on CT image recognition. The data sets used for model training and testing included 121 labeled CT images from hospital patients with pulmonary nodules. The research results showed that the model could be used for the non-invasive EGFR mutation type based on CT images. Yaojun Wang, Xinyu Hua, Dongbo Bu, Shiwei Sun, Xingce Wang |
BIBM | 5 |
| 2020 | ISSEC: inferring contacts among protein secondary structure elements using deep object detectionabstractBACKGROUND: The formation of contacts among protein secondary structure elements (SSEs) is an important step in protein folding as it determines topology of protein tertiary structure; hence, inferring inter-SSE contacts is crucial to protein structure prediction. One of the existing strategies infers inter-SSE contacts directly from the predicted possibilities of inter-residue contacts without any preprocessing, and thus suffers from the excessive noises existing in the predicted inter-residue contacts. Another strategy defines SSEs based on protein secondary structure prediction first, and then judges whether each candidate SSE pair could form contact or not. However, it is difficult to accurately determine boundary of SSEs due to the errors in secondary structure prediction. The incorrectly-deduced SSEs definitely hinder subsequent prediction of the contacts among them. RESULTS: We here report an accurate approach to infer the inter-SSE contacts (thus called as ISSEC) using the deep object detection technique. The design of ISSEC is based on the observation that, in the inter-residue contact map, the contacting SSEs usually form rectangle regions with characteristic patterns. Therefore, ISSEC infers inter-SSE contacts through detecting such rectangle regions. Unlike the existing approach directly using the predicted probabilities of inter-residue contact, ISSEC applies the deep convolution technique to extract high-level features from the inter-residue contacts. More importantly, ISSEC does not rely on the pre-defined SSEs. Instead, ISSEC enumerates multiple candidate rectangle regions in the predicted inter-residue contact map, and for each region, ISSEC calculates a confidence score to measure whether it has characteristic patterns or not. ISSEC employs greedy strategy to select non-overlapping regions with high confidence score, and finally infers inter-SSE contacts according to these regions. CONCLUSIONS: Comprehensive experimental results suggested that ISSEC outperformed the state-of-the-art approaches in predicting inter-SSE contacts. We further demonstrated the successful applications of ISSEC to improve prediction of both inter-residue contacts and tertiary structure as well. Jianwei Zhu, Fusong Ju, Lupeng Kong, Shiwei Sun, Wei-Mou Zheng, Dongbo Bu |
BMC Bioinform. | 7 |
| 2019 | DIFFUSE: predicting isoform functions from sequences and expression profiles via deep learningabstractMOTIVATION: Alternative splicing generates multiple isoforms from a single gene, greatly increasing the functional diversity of a genome. Although gene functions have been well studied, little is known about the specific functions of isoforms, making accurate prediction of isoform functions highly desirable. However, the existing approaches to predicting isoform functions are far from satisfactory due to at least two reasons: (i) unlike genes, isoform-level functional annotations are scarce. (ii) The information of isoform functions is concealed in various types of data including isoform sequences, co-expression relationship among isoforms, etc. RESULTS: In this study, we present a novel approach, DIFFUSE (Deep learning-based prediction of IsoForm FUnctions from Sequences and Expression), to predict isoform functions. To integrate various types of data, our approach adopts a hybrid framework by first using a deep neural network (DNN) to predict the functions of isoforms from their genomic sequences and then refining the prediction using a conditional random field (CRF) based on co-expression relationship. To overcome the lack of isoform-level ground truth labels, we further propose an iterative semi-supervised learning algorithm to train both the DNN and CRF together. Our extensive computational experiments demonstrate that DIFFUSE could effectively predict the functions of isoforms and genes. It achieves an average area under the receiver operating characteristics curve of 0.840 and area under the precision-recall curve of 0.581 over 4184 GO functional categories, which are significantly higher than the state-of-the-art methods. We further validate the prediction results by analyzing the correlation between functional similarity, sequence similarity, expression similarity and structural similarity, as well as the consistency between the predicted functions and some well-studied functional features of isoform sequences. AVAILABILITY AND IMPLEMENTATION: https://github.com/haochenucr/DIFFUSE. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hao Chen 0097, Dipan Shaw, Jianyang Zeng 0001, Dongbo Bu, Tao Jiang 0001 |
Bioinform. | 4 |
| 2019 | Best-first search guided multistage mass spectrometry-based glycan identificationabstractMOTIVATION: Glycan identification has long been hampered by complicated branching patterns and various isomeric structures of glycans. Multistage mass spectrometry (MSn) is a promising glycan identification technique as it generates multiple-level fragments of a glycan, which can be explored to deduce branching pattern of the glycan and further distinguish it from other candidates with identical mass. However, the automatic glycan identification still remains a challenge since it mainly relies on expertise to guide a MSn instrument to generate spectra. RESULTS: Here, we proposed a novel method, named bestFSA, based on a best-first search algorithm to guide the process of spectrum producing in glycan identification using MSn. BestFSA is able to select the most appropriate peaks for next round of experiments and complete the identification using as few experimental rounds. Our analysis of seven representative glycans shows that bestFSA correctly distinguishes actual glycans efficiently and suggested bestFSA could be used in practical glycan identification. The combination of the MSn technology coupled with bestFSA should greatly facilitate the automatic identification of glycan branching patterns, with significantly improved identification sensitivity, and reduce time and cost of MSn experiments. AVAILABILITY AND IMPLEMENTATION: http://glycan.ict.ac.cn. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yaojun Wang, Dongbo Bu, Chuncui Huang, Junchuan Dong, Weiyi Pan, Shiwei Sun |
Bioinform. | 2 |
| 2019 | Constructing effective energy functions for protein structure prediction through broadening attraction-basin and reverse Monte Carlo samplingabstractBACKGROUND: The ab initio approaches to protein structure prediction usually employ the Monte Carlo technique to search the structural conformation that has the lowest energy. However, the widely-used energy functions are usually ineffective for conformation search. How to construct an effective energy function remains a challenging task. RESULTS: Here, we present a framework to construct effective energy functions for protein structure prediction. Unlike existing energy functions only requiring the native structure to be the lowest one, we attempt to maximize the attraction-basin where the native structure lies in the energy landscape. The underlying rationale is that each energy function determines a specific energy landscape together with a native attraction-basin, and the larger the attraction-basin is, the more likely for the Monte Carlo search procedure to find the native structure. Following this rationale, we constructed effective energy functions as follows: i) To explore the native attraction-basin determined by a certain energy function, we performed reverse Monte Carlo sampling starting from the native structure, identifying the structural conformations on the edge of attraction-basin. ii) To broaden the native attraction-basin, we smoothened the edge points of attraction-basin through tuning weights of energy terms, thus acquiring an improved energy function. Our framework alternates the broadening attraction-basin and reverse sampling steps (thus called BARS) until the native attraction-basin is sufficiently large. We present extensive experimental results to show that using the BARS framework, the constructed energy functions could greatly facilitate protein structure prediction in improving the quality of predicted structures and speeding up conformation search. CONCLUSION: Using the BARS framework, we constructed effective energy functions for protein structure prediction, which could improve the quality of predicted structures and speed up conformation search as well. Haicang Zhang, Lupeng Kong, Shiwei Sun, Wei-Mou Zheng, Dongbo Bu |
BMC Bioinform. | 7 |
| 2019 | Predicting protein inter-residue contacts using composite likelihood maximization and deep learningabstractBACKGROUND: Accurate prediction of inter-residue contacts of a protein is important to calculating its tertiary structure. Analysis of co-evolutionary events among residues has been proved effective in inferring inter-residue contacts. The Markov random field (MRF) technique, although being widely used for contact prediction, suffers from the following dilemma: the actual likelihood function of MRF is accurate but time-consuming to calculate; in contrast, approximations to the actual likelihood, say pseudo-likelihood, are efficient to calculate but inaccurate. Thus, how to achieve both accuracy and efficiency simultaneously remains a challenge. RESULTS: In this study, we present such an approach (called clmDCA) for contact prediction. Unlike plmDCA using pseudo-likelihood, i.e., the product of conditional probability of individual residues, our approach uses composite-likelihood, i.e., the product of conditional probability of all residue pairs. Composite likelihood has been theoretically proved as a better approximation to the actual likelihood function than pseudo-likelihood. Meanwhile, composite likelihood is still efficient to maximize, thus ensuring the efficiency of clmDCA. We present comprehensive experiments on popular benchmark datasets, including PSICOV dataset and CASP-11 dataset, to show that: i) clmDCA alone outperforms the existing MRF-based approaches in prediction accuracy. ii) When equipped with deep learning technique for refinement, the prediction accuracy of clmDCA was further significantly improved, suggesting the suitability of clmDCA for subsequent refinement procedure. We further present a successful application of the predicted contacts to accurately build tertiary structures for proteins in the PSICOV dataset. CONCLUSIONS: Composite likelihood maximization algorithm can efficiently estimate the parameters of Markov Random Fields and can improve the prediction accuracy of protein inter-residue contacts. Haicang Zhang, Fusong Ju, Jianwei Zhu, Yujuan Gao, Ziwei Xie, Minghua Deng, Shiwei Sun, Wei-Mou Zheng, Dongbo Bu |
BMC Bioinform. | 10 |
| 2019 | Correction to: Predicting protein inter-residue contacts using composite likelihood maximization and deep learningabstractFollowing publication of the original article [1], the author explained that there are several errors in the original article. Haicang Zhang, Fusong Ju, Jianwei Zhu, Yujuan Gao, Ziwei Xie, Minghua Deng, Shiwei Sun, Wei-Mou Zheng, Dongbo Bu |
BMC Bioinform. | 10 |
| 2019 | Constrained maximum weighted bipartite matching: a novel approach to radio broadcast scheduling
Shaojiang Wang, Tianyong Wu, Dongbo Bu, Shaowei Cai 0001 |
Sci. China Inf. Sci. | 4 |
| 2018 | Protein threading using residue co-variation and deep learningabstractMotivation: Template-based modeling, including homology modeling and protein threading, is a popular method for protein 3D structure prediction. However, alignment generation and template selection for protein sequences without close templates remain very challenging. Results: We present a new method called DeepThreader to improve protein threading, including both alignment generation and template selection, by making use of deep learning (DL) and residue co-variation information. Our method first employs DL to predict inter-residue distance distribution from residue co-variation and sequential information (e.g. sequence profile and predicted secondary structure), and then builds sequence-template alignment by integrating predicted distance information and sequential features through an ADMM algorithm. Experimental results suggest that predicted inter-residue distance is helpful to both protein alignment and template selection especially for protein sequences without very close templates, and that our method outperforms currently popular homology modeling method HHpred and threading method CNFpred by a large margin and greatly outperforms the latest contact-assisted protein threading method EigenTHREADER. Availability and implementation: http://raptorx.uchicago.edu/. Supplementary information: Supplementary data are available at Bioinformatics online. Jianwei Zhu, Sheng Wang 0001, Dongbo Bu, Jinbo Xu |
Bioinform. | 3 |
| 2017 | Improving protein fold recognition by extracting fold-specific features from predicted residue-residue contactsabstractMOTIVATION: Accurate recognition of protein fold types is a key step for template-based prediction of protein structures. The existing approaches to fold recognition mainly exploit the features derived from alignments of query protein against templates. These approaches have been shown to be successful for fold recognition at family level, but usually failed at superfamily/fold levels. To overcome this limitation, one of the key points is to explore more structurally informative features of proteins. Although residue-residue contacts carry abundant structural information, how to thoroughly exploit these information for fold recognition still remains a challenge. RESULTS: In this study, we present an approach (called DeepFR) to improve fold recognition at superfamily/fold levels. The basic idea of our approach is to extract fold-specific features from predicted residue-residue contacts of proteins using deep convolutional neural network (DCNN) technique. Based on these fold-specific features, we calculated similarity between query protein and templates, and then assigned query protein with fold type of the most similar template. DCNN has showed excellent performance in image feature extraction and image recognition; the rational underlying the application of DCNN for fold recognition is that contact likelihood maps are essentially analogy to images, as they both display compositional hierarchy. Experimental results on the LINDAHL dataset suggest that even using the extracted fold-specific features alone, our approach achieved success rate comparable to the state-of-the-art approaches. When further combining these features with traditional alignment-related features, the success rate of our approach increased to 92.3%, 82.5% and 78.8% at family, superfamily and fold levels, respectively, which is about 18% higher than the state-of-the-art approach at fold level, 6% higher at superfamily level and 1% higher at family level. An independent assessment on SCOP_TEST dataset showed consistent performance improvement, indicating robustness of our approach. Furthermore, bi-clustering results of the extracted features are compatible with fold hierarchy of proteins, implying that these features are fold-specific. Together, these results suggest that the features extracted from predicted contacts are orthogonal to alignment-related features, and the combination of them could greatly facilitate fold recognition at superfamily/fold levels and template-based prediction of protein structures. AVAILABILITY AND IMPLEMENTATION: Source code of DeepFR is freely available through https://github.com/zhujianwei31415/deepfr, and a web server is available through http://protein.ict.ac.cn/deepfr. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jianwei Zhu, Haicang Zhang, Shuaicheng Li 0001, Lupeng Kong, Shiwei Sun, Wei-Mou Zheng, Dongbo Bu |
Bioinform. | 8 |
| 2017 | Improving prediction of burial state of residues by exploiting correlation among residuesabstractBACKGROUND: Residues in a protein might be buried inside or exposed to the solvent surrounding the protein. The buried residues usually form hydrophobic cores to maintain the structural integrity of proteins while the exposed residues are tightly related to protein functions. Thus, the accurate prediction of solvent accessibility of residues will greatly facilitate our understanding of both structure and functionalities of proteins. Most of the state-of-the-art prediction approaches consider the burial state of each residue independently, thus neglecting the correlations among residues. RESULTS: In this study, we present a high-order conditional random field model that considers burial states of all residues in a protein simultaneously. Our approach exploits not only the correlation among adjacent residues but also the correlation among long-range residues. Experimental results showed that by exploiting the correlation among residues, our approach outperformed the state-of-the-art approaches in prediction accuracy. In-depth case studies also showed that by using the high-order statistical model, the errors committed by the bidirectional recurrent neural network and chain conditional random field models were successfully corrected. CONCLUSIONS: Our methods enable the accurate prediction of residue burial states, which should greatly facilitate protein structure prediction and evaluation. Hai'e Gong, Haicang Zhang, Jianwei Zhu, Shiwei Sun, Wei-Mou Zheng, Dongbo Bu |
BMC Bioinform. | 7 |
| 2016 | FALCON@home: a high-throughput protein structure prediction server based on remote homologue recognitionabstractSUMMARY: The protein structure prediction approaches can be categorized into template-based modeling (including homology modeling and threading) and free modeling. However, the existing threading tools perform poorly on remote homologous proteins. Thus, improving fold recognition for remote homologous proteins remains a challenge. Besides, the proteome-wide structure prediction poses another challenge of increasing prediction throughput. In this study, we presented FALCON@home as a protein structure prediction server focusing on remote homologue identification. The design of FALCON@home is based on the observation that a structural template, especially for remote homologous proteins, consists of conserved regions interweaved with highly variable regions. The highly variable regions lead to vague alignments in threading approaches. Thus, FALCON@home first extracts conserved regions from each template and then aligns a query protein with conserved regions only rather than the full-length template directly. This helps avoid the vague alignments rooted in highly variable regions, improving remote homologue identification. We implemented FALCON@home using the Berkeley Open Infrastructure of Network Computing (BOINC) volunteer computing protocol. With computation power donated from over 20,000 volunteer CPUs, FALCON@home shows a throughput as high as processing of over 1000 proteins per day. In the Critical Assessment of protein Structure Prediction (CASP11), the FALCON@home-based prediction was ranked the 12th in the template-based modeling category. As an application, the structures of 880 mouse mitochondria proteins were predicted, which revealed the significant correlation between protein half-lives and protein structural factors. AVAILABILITY AND IMPLEMENTATION: FALCON@home is freely available at http://protein.ict.ac.cn/FALCON/. CONTACT: [email protected], [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Haicang Zhang, Wei-Mou Zheng, Dong Xu 0002, Jianwei Zhu, Kang Ning 0001, Shiwei Sun, Shuaicheng Li 0001, Dongbo Bu |
Bioinform. | 10 |
| 2015 | Condensing Raman spectrum for single-cell phenotype analysisabstractBACKGROUND: In recent years, high throughput and non-invasive Raman spectrometry technique has matured as an effective approach to identification of individual cells by species, even in complex, mixed populations. Raman profiling is an appealing optical microscopic method to achieve this. To fully utilize Raman proling for single-cell analysis, an extensive understanding of Raman spectra is necessary to answer questions such as which filtering methodologies are effective for pre-processing of Raman spectra, what strains can be distinguished by Raman spectra, and what features serve best as Raman-based biomarkers for single-cells, etc. RESULTS: In this work, we have proposed an approach called rDisc to discretize the original Raman spectrum into only a few (usually less than 20) representative peaks (Raman shifts). The approach has advantages in removing noises, and condensing the original spectrum. In particular, effective signal processing procedures were designed to eliminate noise, utilising wavelet transform denoising, baseline correction, and signal normalization. In the discretizing process, representative peaks were selected to signicantly decrease the Raman data size. More importantly, the selected peaks are chosen as suitable to serve as key biological markers to differentiate species and other cellular features. Additionally, the classication performance of discretized spectra was found to be comparable to full spectrum having more than 1000 Raman shifts. Overall, the discretized spectrum needs about 5storage space of a full spectrum and the processing speed is considerably faster. This makes rDisc clearly superior to other methods for single-cell classication. Shiwei Sun, Xuetao Wang, Xin Gao 0001, Lihui Ren, Xiaoquan Su, Dongbo Bu, Kang Ning 0001 |
BMC Bioinform. | 6 |
| 2015 | OpenMS-Simulator: an open-source software for theoretical tandem mass spectrum predictionabstractBACKGROUND: Tandem mass spectrometry (MS/MS) acts as a key technique for peptide identification. The MS/MS-based peptide identification approaches can be categorized into two families, namely, de novo and database search. Both of the two types of approaches can benefit from an accurate prediction of theoretical spectrum. A theoretical spectrum consists of m/z and intensity of possibly occurring ions, which are estimated via simulating the spectrum generating process. Extensive researches have been conducted for theoretical spectrum prediction; however, the prediction methods suffer from low prediciton accuracy due to oversimplifications in the spectrum simulation process. RESULTS: In the study, we present an open-source software package, called OpenMS-Simulator, to predict theoretical spectrum for a given peptide sequence. Based on the mobile-proton hypothesis for peptide fragmentation, OpenMS-Simulator trained a closed-form model for the intensity ratio of adjacent y ions, from which the whole theoretical spectrum can be constructed. On a collection of representative spectra datasets with annotated peptide sequences, experimental results suggest that OpenMS-Simulator can predict theoretical spectra with considerable accuracy. The study also presents an application of OpenMS-Simulator: the similarity between theoretical spectra and query spectra can be used to re-rank the peptide sequence reported by SEQUEST/X!Tandem. CONCLUSIONS: OpenMS-Simulator implements a novel model to predict theoretical spectrum for a given peptide sequence. Compared with existing theoretical spectrum prediction tools, say MassAnalyzer and MSSimulator, our method not only simplifies the computation process, but also improves the prediction accuracy. Currently, OpenMS-Simulator supports the prediction of CID and HCD spectrum for peptides with double charges. The extension to cover more fragmentation models and support multiple-charged peptides remains as one of the future works. Yaojun Wang, Dongbo Bu, Shiwei Sun |
BMC Bioinform. | 4 |
| 2015 | Parameterized BLOSUM Matrices for Protein AlignmentabstractProtein alignment is a basic step for many molecular biology researches. The BLOSUM matrices, especially BLOSUM62, are the de facto standard matrices for protein alignments. However, after widely utilization of the matrices for 15 years, programming errors were surprisingly found in the initial version of source codes for their generation. And amazingly, after bug correction, the "intended" BLOSUM62 matrix performs consistently worse than the "miscalculated" one. In this paper, we find linear relationships among the eigenvalues of the matrices and propose an algorithm to find optimal unified eigenvectors. With them, we can parameterize matrix BLOSUMx for any given variable x that could change continuously. We compare the effectiveness of our parameterized isentropic matrix with BLOSUM62. Furthermore, an iterative alignment and matrix selection process is proposed to adaptively find the best parameter and globally align two sequences. Experiments are conducted on aligning 13,667 families of Pfam database and on clustering MHC II protein sequences, whose improved accuracy demonstrates the effectiveness of our proposed method. Dongbo Bu, Shuaicheng Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2013 | Towards Reliable Automatic Protein Structure Alignment
Xuefeng Cui, Shuaicheng Li 0001, Dongbo Bu, Ming Li 0001 |
WABI | 3 |
| 2012 | How Accurately Can We Model Protein Structures with Dihedral Angles?
Xuefeng Cui, Shuaicheng Li 0001, Dongbo Bu, Babak Alipanahi, Ming Li 0001 |
WABI | 3 |
| 2012 | Residues with Similar Hexagon Neighborhoods Share Similar Side-Chain ConformationsabstractWe present in this study a new approach to code protein side-chain conformations into hexagon substructures. Classical side-chain packing methods consist of two steps: first, side-chain conformations, known as rotamers, are extracted from known protein structures as candidates for each residue; second, a searching method along with an energy function is used to resolve conflicts among residues and to optimize the combinations of side chain conformations for all residues. These methods benefit from the fact that the number of possible side-chain conformations is limited, and the rotamer candidates are readily extracted; however, these methods also suffer from the inaccuracy of energy functions. Inspired by threading and Ab Initio approaches to protein structure prediction, we propose to use hexagon substructures to implicitly capture subtle issues of energy functions. Our initial results indicate that even without guidance from an energy function, hexagon structures alone can capture side-chain conformations at an accuracy of 83.8 percent, higher than 82.6 percent by the state-of-art side-chain packing methods. Shuaicheng Li 0001, Dongbo Bu, Ming Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2012 | Clustering 100, 000 Protein Structure Decoys in MinutesabstractAb initio protein structure prediction methods first generate large sets of structural conformations as candidates (called decoys), and then select the most representative decoys through clustering techniques. Classical clustering methods are inefficient due to the pairwise distance calculation, and thus become infeasible when the number of decoys is large. In addition, the existing clustering approaches suffer from the arbitrariness in determining a distance threshold for proteins within a cluster: a small distance threshold leads to many small clusters, while a large distance threshold results in the merging of several independent clusters into one cluster. In this paper, we propose an efficient clustering method through fast estimating cluster centroids and efficient pruning rotation spaces. The number of clusters is automatically detected by information distance criteria. A package named ONION, which can be downloaded freely, is implemented accordingly. Experimental results on benchmark data sets suggest that ONION is 14 times faster than existing tools, and ONION obtains better selections for 31 targets, and worse selection for 19 targets compared to SPICKER’s selections. On an average PC, ONION can cluster 100,000 decoys in around 12 minutes. Shuaicheng Li 0001, Dongbo Bu, Ming Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2011 | PI: An open-source software package for validation of the SEQUEST result and visualization of mass spectrumabstractBACKGROUND: Tandem mass spectrometry (MS/MS) has emerged as the leading method for high- throughput protein identification in proteomics. Recent technological breakthroughs have dramatically increased the efficiency of MS/MS data generation. Meanwhile, sophisticated algorithms have been developed for identifying proteins from peptide MS/MS data by searching available protein sequence databases for the peptide that is most likely to have produced the observed spectrum. The popular SEQUEST algorithm relies on the cross-correlation between the experimental mass spectrum and the theoretical spectrum of a peptide. It utilizes a simplified fragmentation model that assigns a fixed and identical intensity for all major ions and fixed and lower intensity for their neutral losses. In this way, the common issues involved in predicting theoretical spectra are circumvented. In practice, however, an experimental spectrum is usually not similar to its SEQUEST -predicted theoretical one, and as a result, incorrect identifications are often generated. RESULTS: Better understanding of peptide fragmentation is required to produce more accurate and sensitive peptide sequencing algorithms. Here, we designed the software PI of novel and exquisite algorithms that make a good use of intensity property of a spectrum. CONCLUSIONS: We designed the software PI with the novel and effective algorithms which made a good use of intensity property of the spectrum. Experiments have shown that PI was able to validate and improve the results of SEQUEST to a more satisfactory degree. Yantao Qiao, Dongbo Bu, Shiwei Sun |
BMC Bioinform. | 3 |
| 2011 | Incorporating Ab Initio energy into threading approaches for protein structure predictionabstractBACKGROUND: Native structures of proteins are formed essentially due to the combining effects of local and distant (in the sense of sequence) interactions among residues. These interaction information are, explicitly or implicitly, encoded into the scoring function in protein structure prediction approaches--threading approaches usually measure an alignment in the sense that how well a sequence adopts an existing structure; while the energy functions in Ab Initio methods are designed to measure how likely a conformation is near-native. Encouraging progress has been observed in structure refinement where knowledge-based or physics-based potentials are designed to capture distant interactions. Thus, it is interesting to investigate whether distant interaction information captured by the Ab Initio energy function can be used to improve threading, especially for the weakly/distant homologous templates. RESULTS: In this paper, we investigate the possibility to improve alignment-generating through incorporating distant interaction information into the alignment scoring function in a nontrivial approach. Specifically, the distant interaction information is introduced through employing an Ab Initio energy function to evaluate the "partial" decoy built from an alignment. Subsequently, a local search algorithm is utilized to optimize the scoring function.Experimental results demonstrate that with distant interaction items, the quality of generated alignments are improved on 68 out of 127 query-template pairs in Prosup benchmark. In addition, compared with state-to-art threading methods, our method performs better on alignment accuracy comparison. CONCLUSIONS: Incorporating Ab Initio energy functions into threading can greatly improve alignment accuracy. Mingfu Shao, Sheng Wang 0001, Xiongying Yuan, Wei-Mou Zheng, Dongbo Bu |
BMC Bioinform. | 7 |
| 2011 | ProbPS: A new model for peak selection based on quantifying the dependence of the existence of derivative peaks on primary ion intensityabstractBACKGROUND: The analysis of mass spectra suggests that the existence of derivative peaks is strongly dependent on the intensity of the primary peaks. Peak selection from tandem mass spectrum is used to filter out noise and contaminant peaks. It is widely accepted that a valid primary peak tends to have high intensity and is accompanied by derivative peaks, including isotopic peaks, neutral loss peaks, and complementary peaks. Existing models for peak selection ignore the dependence between the existence of the derivative peaks and the intensity of the primary peaks. Simple models for peak selection assume that these two attributes are independent; however, this assumption is contrary to real data and prone to error. RESULTS: In this paper, we present a statistical model to quantitatively measure the dependence of the derivative peak's existence on the primary peak's intensity. Here, we propose a statistical model, named ProbPS, to capture the dependence in a quantitative manner and describe a statistical model for peak selection. Our results show that the quantitative understanding can successfully guide the peak selection process. By comparing ProbPS with AuDeNS we demonstrate the advantages of our method in both filtering out noise peaks and in improving de novo identification. In addition, we present a tag identification approach based on our peak selection method. Our results, using a test data set, suggest that our tag identification method (876 correct tags in 1000 spectra) outperforms PepNovoTag (790 correct tags in 1000 spectra). CONCLUSIONS: We have shown that ProbPS improves the accuracy of peak selection which further enhances the performance of de novo sequencing and tag identification. Thus, our model saves valuable computation time and improving the accuracy of the results. Shenghui Zhang, Yaojun Wang, Dongbo Bu, Shiwei Sun |
BMC Bioinform. | 3 |
| 2011 | Revisiting Multiple Pattern Matching Algorithms for Multi-Core Architecture
Guangming Tan, Ping Liu 0001, Dongbo Bu, Yanbing Liu 0007 |
J. Comput. Sci. Technol. | 3 |
| 2009 | Finding compact structural motifs
Dongbo Bu, Ming Li 0001, Shuaicheng Li 0001, Jianbo Qian, Jinbo Xu |
Theor. Comput. Sci. | 1 |
| 2008 | Finding Largest Well-Predicted Subset of Protein Structure Models
Shuaicheng Li 0001, Dongbo Bu, Jinbo Xu, Ming Li 0001 |
CPM | 2 |
| 2008 | Designing succinct structural alphabetsabstractMOTIVATION: The 3D structure of a protein sequence can be assembled from the substructures corresponding to small segments of this sequence. For each small sequence segment, there are only a few more likely substructures. We call them the 'structural alphabet' for this segment. Classical approaches such as ROSETTA used sequence profile and secondary structure information, to predict structural fragments. In contrast, we utilize more structural information, such as solvent accessibility and contact capacity, for finding structural fragments. RESULTS: Integer linear programming technique is applied to derive the best combination of these sequence and structural information items. This approach generates significantly more accurate and succinct structural alphabets with more than 50% improvement over the previous accuracies. With these novel structural alphabets, we are able to construct more accurate protein structures than the state-of-art ab initio protein structure prediction programs such as ROSETTA. We are also able to reduce the Kolodny's library size by a factor of 8, at the same accuracy. AVAILABILITY: The online FRazor server is under construction. Shuaicheng Li 0001, Dongbo Bu, Xin Gao 0001, Jinbo Xu, Ming Li 0001 |
ISMB | 2 |
| 2008 | A Fragmentation Event Model for Peptide Identification by Mass Spectrometry
Yu Lin 0001, Yantao Qiao, Shiwei Sun, Chungong Yu, Gongjin Dong, Dongbo Bu |
RECOMB | 6 |
| 2008 | FlexStem: improving predictions of RNA secondary structures with pseudoknots by reducing the search spaceabstractMOTIVATION: RNA secondary structures with pseudoknots are often predicted by minimizing free energy, which is proved to be NP-hard. Due to kinetic reasons the real RNA secondary structure often has local instead of global minimum free energy. This implies that we may improve the performance of RNA secondary structure prediction by taking kinetics into account and minimize free energy in a local area. RESULT: we propose a novel algorithm named FlexStem to predict RNA secondary structures with pseudoknots. Still based on MFE criterion, FlexStem adopts comprehensive energy models that allow complex pseudoknots. Unlike classical thermodynamic methods, our approach aims to simulate the RNA folding process by successive addition of maximal stems, reducing the search space while maintaining or even improving the prediction accuracy. This reduced space is constructed by our maximal stem strategy and stem-adding rule induced from elaborate statistical experiments on real RNA secondary structures. The strategy and the rule also reflect the folding characteristic of RNA from a new angle and help compensate for the deficiency of merely relying on MFE in RNA structure prediction. We validate FlexStem by applying it to tRNAs, 5SrRNAs and a large number of pseudoknotted structures and compare it with the well-known algorithms such as RNAfold, PKNOTS, PknotsRG, HotKnots and ILM according to their overall sensitivities and specificities, as well as positive and negative controls on pseudoknots. The results show that FlexStem significantly increases the prediction accuracy through its local search strategy. AVAILABILITY: Software is available at http://pfind.ict.ac.cn/FlexStem/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Simin He 0001, Dongbo Bu, Runsheng Chen, Wen Gao 0001 |
Bioinform. | 3 |
| 2007 | A Flexible Stem-Based Local Search Algorithm for Predicting RNA Secondary Structures Including PseudoknotsabstractWe demonstrate a new algorithm named FlexStem to predict RNA secondary structures with pseudoknots. Our approach is based on the free energy minimization criterion, and utilizes a sophisticated energy model that is more accurate and supports more types of pseudoknots. By using a "maximal stem" and "stem merging" strategy, the search space for RNA structure prediction is significantly reduced. We have also developed a greedy search algorithm with perturbation on stems. The FlexStem algorithm is applied to a large number of sequences taken from Sprinzl, Pseudobase and other reliable resource. We find that FlexStem outperforms the well-known optimal and heuristic algorithms such as Mfold, PKNOTS, HotKnots and ILM in overall sensitivity and has comparable results to those algorithms in overall specificity. Performance evaluation demonstrate that FlexStem requires significantly less time than the optimal algorithm PKNOTS. In addition, our algorithm has better prediction results on pseudoknot-free sequences compared to competing algorithms. Simin He 0001, Dongbo Bu, Runsheng Chen, Wen Gao 0001 |
BIBE | 3 |
| 2007 | Finding Compact Structural Motifs
Jianbo Qian, Shuaicheng Li 0001, Dongbo Bu, Ming Li 0001, Jinbo Xu |
CPM | 3 |
| 2006 | Improving locality of nonserial polyadic dynamic programming
Guangming Tan, Ninghui Sun, Dongbo Bu |
IPDPS | 3 |
| 2006 | A novel scoring schema for peptide identification by searching protein sequence databases using tandem mass spectrometry dataabstractBACKGROUND: Tandem mass spectrometry (MS/MS) is a powerful tool for protein identification. Although great efforts have been made in scoring the correlation between tandem mass spectra and an amino acid sequence database, improvements could be made in three aspects, including characterization ofpeaks in spectra, adoption of effective scoring functions and access to thereliability of matching between peptides and spectra. RESULTS: A novel scoring function is presented, along with criteria to estimate the performance confidence of the function. Through learning the typesof product ions and the probability of generating them, a hypothetic spectrum was generated for each candidate peptide. Then relative entropy was introduced to measure the similarity between the hypothetic and the observed spectra. Based on the extreme value distribution (EVD) theory, a threshold was chosen to distinguish a true peptide assignment from a random one. Tests on a public MS/MS dataset demonstrated that this method performs better than the well-known SEQUEST. CONCLUSION: A reliable identification of proteins from the spectra promises a more efficient application of tandem mass spectrometry to proteomes with high complexity. Shiwei Sun, Suhua Chang, Chungong Yu, Dongbo Bu, Runsheng Chen |
BMC Bioinform. | 7 |
| 2006 | Improvement of Performance of MegaBlast Algorithm for DNA Sequence Alignment
Guangming Tan, Dongbo Bu, Shengzhong Feng, Ninghui Sun |
J. Comput. Sci. Technol. | 3 |
| 2004 | A Note on the Single Genotype Resolution Problem
Qiang-Feng Zhang, Dongbo Bu, Ming Li 0001 |
J. Comput. Sci. Technol. | 4 |