VLDB 2026 Research / reviewers in the wild / expert
Shiwei Sun
dblp:93/6964
· DBLP profile ↗
36ranked-venue papers
3as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 32 · 1 first-author · 13 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual-Axis Message Passing on Molecular Graphs: Layerwise Multi-hop Convolution Meets Depthwise Dense Mixing
Xuefeng Cui, Shiwei Sun |
ISBRA (1) | 2 |
| 2024 | Deep Representation Learning for Electron Ionization Mass Spectra RetrievalabstractThe task of retrieving and analyzing mass spectra is indispensable for the identification of compounds in mass spectrometry (MS). This methodology is of critical importance as it enables researchers to correlate observed spectra with established databases, thereby precisely determining the chemical composition of samples. The primary challenges to its efficacy lie in optimizing the balance between retrieval accuracy and processing speed. Empirical studies have demonstrated that by converting mass spectra into embeddings via deep learning, it is possible to achieve both high accuracy and speed in retrieval. Nevertheless, there are complex challenges associated with employing deep learning for spectral embedding, particularly within the domain of electron ionization mass spectrometry (EI-MS). In this paper, we introduce a novel representation learning technique termed EI-MS2VEC for EI-MS retrieval. Our spectrum retrieval methodology surpasses current state-of-the-art techniques such as FastEI. For the in-silico library, we attain hit rate@1 and hit rate@10 of 43.6% and 84.5%, respectively, compared to FastEI’s 36.7% and 80.4%. Moreover, our retrieval approach operates with an order of magnitude greater speed than FastEI. The source code is available on Github (https://github.com/xfcui/EI-MS2VEC). Anlei Jiao, Shiwei Sun, Longyang Dian, Xuefeng Cui |
BIBM | 4 |
| 2024 | How to Train Your Neural Network for Molecular Structure Generation from Mass Spectra?abstractMass spectrometry serves as a pivotal tool for the analysis of small molecules through an examination of their mass-to-charge ratios. Recent advancements in deep learning have markedly enhanced the analysis of mass spectrometric data, facilitating the prediction of novel small molecule structures without the necessity of extensive databases. Nonetheless, the paucity of annotated datasets impedes the efficacious training of molecular generation models predicated on MS2spectra. To mitigate this limitation, we introduce ctMSNovelist, an avant-garde method that amalgamates pre-training, fine-tuning, and co-training techniques to construct a more precise model for the generation of molecular structures from tandem mass spectrometry data. This novel approach augments both the training regimen and the predictive accuracy of the MSNovelist model, thereby surmounting the obstacle of limited data. The methodology commences with the pretraining of a Variational Autoencoder (VAE) to generate molecular fingerprints derived from SMILES strings. Subsequently, it undergoes fine-tuning to emulate noisy fingerprints originating from mass spectrometry (MS) data. Concurrently, MSNovelist is co-trained utilizing these simulated fingerprints, inclusive of the highly noisy variants produced in the early stages of VAE training. The incorporation of a substantial volume of noisy data serves to enhance model accuracy and avert overfitting. We evaluated ctMSNovelist using the GNPS dataset and attained a SMILES prediction accuracy of 48.8%, representing a 4.1% enhancement over MSNovelist. It is pertinent to note that the sole distinction between ctMSNovelist and MSNovelist in this experiment was the training process. The code and models are publicly available at https://github.com/xfcui/ctMSNovelist. Yanmin Liu, Longyang Dian, Shiwei Sun, Xuefeng Cui |
BIBM | 4 |
| 2023 | Integrative Drug Discovery Platform: A Modular Approach for Efficient and Automated Virtual ScreeningabstractThis paper presents a drug development platform based on virtual screening technology. The platform integrates key components such as pocket prediction, molecular docking, molecular dynamics simulation, and ADMET evaluation to achieve an efficient and automated drug virtual screening process. The platform utilizes Docker for modular encapsulation, ensuring environment isolation and convenient deployment. It also provides standardized input-output formats and a task allocation system, enabling users to quickly deploy and customize the workflow. Experimental results demonstrate the effectiveness of the platform in identifying real drugs and evaluating virtual screening results, providing an efficient and reliable solution for drug development. The platform features easy deployment and migration, independent module execution, automated workflow implementation, personalized customization and replacement, task allocation for computationally intensive steps, and complex operations in molecular dynamics simulation. Lulu Xie, Zhonghai Zhang, Bo Duan, Gang Niu 0008, Shiwei Sun, Fa Zhang 0001, Runting Zhang, Guangming Tan |
BIBM | 7 |
| 2023 | Learned Fingerprint Embedding for Large-Scale Peptide Mass Spectra RetrievalabstractTandem mass spectrometry (MS/MS) is a widely used technique for protein identification, post-translational modifications, immunotherapy, and other applications. As the amount of MS/MS spectra data increases, new computational methods are needed to efficiently search through these databases. This study introduces MS2VEC, a novel fingerprint embedding model designed to facilitate large-scale retrieval of peptide mass spectra. MS2VEC captures the relationships between distant peaks and incorporates position-aware fingerprint features from all peaks. To do this, dilated convolutions are used to capture remote relationships, and a novel position-aware multi-head attention pooling mechanism is used to abstract fingerprint features. The results demonstrate that MS2VEC achieves a top-1 retrieval accuracy of 0.810, outperforming existing methods by 5.1%. Interestingly, the precursor charge is not essential for the retrieval task, as the spectra itself contains enough information to accurately predict the charge. Additionally, the results suggest that weight-balanced fragment ions and water losses are important contributors to fingerprint features. Yongshuai Wang, Xiaojun Cai, Defeng Li, Shiwei Sun, Xuefeng Cui |
BIBM | 4 |
| 2023 | EMNGly: predicting N-linked glycosylation sites using the language models for feature extractionabstractMOTIVATION: N-linked glycosylation is a frequently occurring post-translational protein modification that serves critical functions in protein folding, stability, trafficking, and recognition. Its involvement spans across multiple biological processes and alterations to this process can result in various diseases. Therefore, identifying N-linked glycosylation sites is imperative for comprehending the mechanisms and systems underlying glycosylation. Due to the inherent experimental complexities, machine learning and deep learning have become indispensable tools for predicting these sites. RESULTS: In this context, a new approach called EMNGly has been proposed. The EMNGly approach utilizes pretrained protein language model (Evolutionary Scale Modeling) and pretrained protein structure model (Inverse Folding Model) for features extraction and support vector machine for classification. Ten-fold cross-validation and independent tests show that this approach has outperformed existing techniques. And it achieves Matthews Correlation Coefficient, sensitivity, specificity, and accuracy of 0.8282, 0.9343, 0.8934, and 0.9143, respectively on a benchmark independent test set. Xiaoyang Hou, Dongbo Bu, Yaojun Wang, Shiwei Sun |
Bioinform. | 5 |
| 2023 | Accurate and efficient protein sequence design through learning concise local environment of residuesabstractMOTIVATION: Computational protein sequence design has been widely applied in rational protein engineering and increasing the design accuracy and efficiency is highly desired. RESULTS: Here, we present ProDESIGN-LE, an accurate and efficient approach to protein sequence design. ProDESIGN-LE adopts a concise but informative representation of the residue's local environment and trains a transformer to learn the correlation between local environment of residues and their amino acid types. For a target backbone structure, ProDESIGN-LE uses the transformer to assign an appropriate residue type for each position based on its local environment within this structure, eventually acquiring a designed sequence with all residues fitting well with their local environments. We applied ProDESIGN-LE to design sequences for 68 naturally occurring and 129 hallucinated proteins within 20 s per protein on average. The designed proteins have their predicted structures perfectly resembling the target structures with a state-of-the-art average TM-score exceeding 0.80. We further experimentally validated ProDESIGN-LE by designing five sequences for an enzyme, chloramphenicol O-acetyltransferase type III (CAT III), and recombinantly expressing the proteins in Escherichia coli. Of these proteins, three exhibited excellent solubility, and one yielded monomeric species with circular dichroism spectra consistent with the natural CAT III protein. AVAILABILITY AND IMPLEMENTATION: The source code of ProDESIGN-LE is available at https://github.com/bigict/ProDESIGN-LE. Bin Huang 0022, Tingwen Fan, Kaiyue Wang, Haicang Zhang, Chungong Yu, Shuyu Nie, Yangshuo Qi, Wei-Mou Zheng, Shiwei Sun, Huaiyi Yang, Dongbo Bu |
Bioinform. | 11 |
| 2023 | SASA-Net: A Spatial-Aware Self-Attention Mechanism for Building Protein 3D Structure Directly From Inter- Residue DistancesabstractProtein functions are tightly related to the fine details of their 3D structures. To understand protein structures, computational prediction approaches are highly needed. Recently, protein structure prediction has achieved considerable progresses mainly due to the increased accuracy of inter-residue distance estimation and the application of deep learning techniques. Most of the distance-based ab initio prediction approaches adopt a two-step diagram: constructing a potential function based on the estimated inter-residue distances, and then build a 3D structure that minimizes the potential function. These approaches have proven very promising; however, they still suffer from several limitations, especially the inaccuracies incurred by the handcrafted potential function. Here, we present SASA-Net, a deep learning-based approach that directly learns protein 3D structure from the estimated inter-residue distances. Unlike the existing approach simply representing protein structures as coordinates of atoms, SASA-Net represents protein structures using pose of residues, i.e., the coordinate system of each individual residue in which all backbone atoms of this residue are fixed. The key element of SASA-Net is a spatial-aware self-attention mechanism, which is able to adjust a residue's pose according to all other residues' features and the estimated distances between residues. By iteratively applying the spatial-aware self-attention mechanism, SASA-Net continuously improves the structure and finally acquires a structure with high accuracy. Using the CATH35 proteins as representatives, we demonstrate that SASA-Net is able to accurately and efficiently build structures from the estimated inter-residue distances. The high accuracy and efficiency of SASA-Net enables an end-to-end neural network model for protein structure prediction through combining SASA-Net and an neural network for inter-residue distance prediction. Source code of SASA-Net is available at https://github.com/gongtiansu/SASA-Net/. Tiansu Gong, Fusong Ju, Shiwei Sun, Dongbo Bu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | Automated Patching for Unreproducible BuildsabstractSoftware reproducibility plays an essential role in establishing trust between source code and the built artifacts, by comparing compilation outputs acquired from independent users. Although the testing for unreproducible builds could be automated, fixing unreproducible build issues poses a set of challenges within the reproducible builds practice, among which we consider the localization granularity and the historical knowledge utilization as the most significant ones. To tackle these challenges, we propose a novel approach RepFix that combines tracing-based fine-grained localization with history-based patch generation mechanisms. Zhilei Ren, Shiwei Sun, Jifeng Xuan, Zhide Zhou, He Jiang 0001 |
ICSE | 2 |
| 2022 | Seq-SetNet: directly exploiting multiple sequence alignment for protein secondary structure predictionabstractMOTIVATION: Accurate prediction of protein structure relies heavily on exploiting multiple sequence alignment (MSA) for residue mutations and correlations as this information specifies protein tertiary structure. The widely used prediction approaches usually transform MSA into inter-mediate models, say position-specific scoring matrix or profile hidden Markov model. These inter-mediate models, however, cannot fully represent residue mutations and correlations carried by MSA; hence, an effective way to directly exploit MSAs is highly desirable. RESULTS: Here, we report a novel sequence set network (called Seq-SetNet) to directly and effectively exploit MSA for protein structure prediction. Seq-SetNet uses an 'encoding and aggregation' strategy that consists of two key elements: (i) an encoding module that takes a component homologue in MSA as input, and encodes residue mutations and correlations into context-specific features for each residue; and (ii) an aggregation module to aggregate the features extracted from all component homologues, which are further transformed into structural properties for residues of the query protein. As Seq-SetNet encodes each homologue protein individually, it could consider both insertions and deletions, as well as long-distance correlations among residues, thus representing more information than the inter-mediate models. Moreover, the encoding module automatically learns effective features and thus avoids manual feature engineering. Using symmetric aggregation functions, Seq-SetNet processes the homologue proteins as a sequence set, making its prediction results invariable to the order of these proteins. On popular benchmark sets, we demonstrated the successful application of Seq-SetNet to predict secondary structure and torsion angles of residues with improved accuracy and efficiency. AVAILABILITY AND IMPLEMENTATION: The code and datasets are available through https://github.com/fusong-ju/Seq-SetNet. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Fusong Ju, Jianwei Zhu, Guozheng Wei, Shiwei Sun, Wei-Mou Zheng, Dongbo Bu |
Bioinform. | 5 |
| 2022 | When Job Candidates Experience Social Media Privacy Violations: A Cross-Culture StudyabstractThis study uses a cross-cultural sample from the U.S. and China to compare information privacy-protective responses to a breach in privacy during a job interview. Using a job recruitment scenario, the relationships among individuals' concern for information privacy, disposition to trust, judgment of moral issues, and their information privacy-protective responses were examined. Based on the multiple group analysis results, this paper find that the privacy-protective responses significantly vary between the American and Chinese cultures. The findings shed light on individuals' responses to privacy issues in the United States and China. Shiwei Sun, John R. Drake, Dianne J. Hall |
J. Glob. Inf. Manag. | 1 |
| 2021 | SASA-Net: A spatial-aware self-attention mechanism for building protein 3D structure directly from inter-residue distancesabstractProtein structure prediction has achieved considerable progresses mainly due to the increased accuracy of inter-residue distance estimation and the application of deep learning techniques.Most of the distance-based ab initio prediction approaches adopt a two-step diagram: constructing a potential function based on the estimated inter-residue distances, and then build a 3D structure that minimizes the potential function. These approaches have proven very promising; however, they still suffer from several limitations, especially the inaccuracies incurred by the handcrafted potential function.Here, we present SASA-Net, a deep learning-based approach that directly learns protein 3D structure from the estimated inter-residue distances. Unlike the existing approach simply representing protein structures as coordinates of atoms, SASA-Net represents protein structures using pose of residues, i.e., the coordinate system of each individual residue in which all backbone atoms of this residue are fixed. The key element of SASA-Net is a spatial-aware self-attention mechanism, which is able to adjust a residue’s pose according to all other residues’ features and the estimated distances between residues. By iteratively applying the spatial-aware self-attention mechanism, SASA-Net continuously improves the structure and finally acquires a structure with high accuracy. Using the CATH35 proteins as representatives, we demonstrate that SASA-Net is able to accurately and efficiently build structures from the estimated inter-residue distances. The high accuracy and efficiency of SASA-Net enables an end-to-end neural network model for protein structure prediction through combining SASA-Net and an neural network for inter-residue distance prediction. Source code of SASA-Net is available at https://github.con gongtiansu/SASA-Net/ Tiansu Gong, Fusong Ju, Shiwei Sun, Dongbo Bu |
BIBM | 3 |
| 2021 | Glycan immunogenicity prediction based on Graph neural networkabstractGlycans play important roles in a great variety of biological processes, and these roles are closely determined by the details of their structures. It becomes possible to acquire hidden features from glycan structures using deep learning method with the great progress in recent years. Unlike the linear chain of proteins and DNAs, branching is a unique feature of glycan structures, which makes it very difficult to directly apply deep learning models on glycans. Thus, how to comprehensively and efficiently describe glycans and use them as input to deep learning models still remains challenging. Here, a graph neural network (GNN) called GlyNet was used to obtain high-dimensional representation of glycans and predict their immunogenicity. Our method was applied in SugarBase, and it works more superiorly than the state-of-art method with accuracy increased from 91.7% to 95.6%. Yu Wang 0225, Meijie Hou, Yaojun Wang, Dongbo Bu, Chuncui Huang, Shiwei Sun |
BIBM | 8 |
| 2021 | Filling gaps of genome scaffolds via probabilistic searching optical maps against assembly graphabstractBACKGROUND: Optical maps record locations of specific enzyme recognition sites within long genome fragments. This long-distance information enables aligning genome assembly contigs onto optical maps and ordering contigs into scaffolds. The generated scaffolds, however, often contain a large amount of gaps. To fill these gaps, a feasible way is to search genome assembly graph for the best-matching contig paths that connect boundary contigs of gaps. The combination of searching and evaluation procedures might be "searching followed by evaluation", which is infeasible for long gaps, or "searching by evaluation", which heavily relies on heuristics and thus usually yields unreliable contig paths. RESULTS: We here report an accurate and efficient approach to filling gaps of genome scaffolds with aids of optical maps. Using simulated data from 12 species and real data from 3 species, we demonstrate the successful application of our approach in gap filling with improved accuracy and completeness of genome scaffolds. CONCLUSION: Our approach applies a sequential Bayesian updating technique to measure the similarity between optical maps and candidate contig paths. Using this similarity to guide path searching, our approach achieves higher accuracy than the existing "searching by evaluation" strategy that relies on heuristics. Furthermore, unlike the "searching followed by evaluation" strategy enumerating all possible paths, our approach prunes the unlikely sub-paths and extends the highly-probable ones only, thus significantly increasing searching efficiency. Guozheng Wei, Fusong Ju, Zhuozheng Shi, Shiwei Sun, Dongbo Bu |
BMC Bioinform. | 7 |
| 2021 | FALCON2: a web server for high-quality prediction of protein tertiary structuresabstractBACKGROUND: Accurate prediction of protein tertiary structures is highly desired as the knowledge of protein structures provides invaluable insights into protein functions. We have designed two approaches to protein structure prediction, including a template-based modeling approach (called ProALIGN) and an ab initio prediction approach (called ProFOLD). Briefly speaking, ProALIGN aligns a target protein with templates through exploiting the patterns of context-specific alignment motifs and then builds the final structure with reference to the homologous templates. In contrast, ProFOLD uses an end-to-end neural network to estimate inter-residue distances of target proteins and builds structures that satisfy these distance constraints. These two approaches emphasize different characteristics of target proteins: ProALIGN exploits structure information of homologous templates of target proteins while ProFOLD exploits the co-evolutionary information carried by homologous protein sequences. Recent progress has shown that the combination of template-based modeling and ab initio approaches is promising. RESULTS: In the study, we present FALCON2, a web server that integrates ProALIGN and ProFOLD to provide high-quality protein structure prediction service. For a target protein, FALCON2 executes ProALIGN and ProFOLD simultaneously to predict possible structures and selects the most likely one as the final prediction result. We evaluated FALCON2 on widely-used benchmarks, including 104 CASP13 (the 13th Critical Assessment of protein Structure Prediction) targets and 91 CASP14 targets. In-depth examination suggests that when high-quality templates are available, ProALIGN is superior to ProFOLD and in other cases, ProFOLD shows better performance. By integrating these two approaches with different emphasis, FALCON2 server outperforms the two individual approaches and also achieves state-of-the-art performance compared with existing approaches. CONCLUSIONS: By integrating template-based modeling and ab initio approaches, FALCON2 provides an easy-to-use and high-quality protein structure prediction service for the community and we expect it to enable insights into a deep understanding of protein functions. Lupeng Kong, Fusong Ju, Haicang Zhang, Shiwei Sun, Dongbo Bu |
BMC Bioinform. | 4 |
| 2020 | Mut-Detecter: An EGFR activating mutation type classification method with a deep convolutional neural networkabstractEpidermal growth factor receptor (EGFR) plays an essential role in tumor cell proliferation, angiogenesis and apoptosis inhibition; it is a crucial factor leading to cancer occurrence. For example, EGFR tyrosine kinase inhibitors in treating lung cancer patients have an excellent therapeutic effect. Targeted therapy based on EGFR gene mutation is one of the mainstream lung cancer treatment methods. Recent studies have shown that pulmonary nodules' characteristics are associated with the mutant status of EGFR, which provides the possibility of using CT images of patients with pulmonary nodules to predict the mutant status of EGFR. This study used the deep learning algorithm to establish the EGFR mutation type prediction model based on CT image recognition. The data sets used for model training and testing included 121 labeled CT images from hospital patients with pulmonary nodules. The research results showed that the model could be used for the non-invasive EGFR mutation type based on CT images. Yaojun Wang, Xinyu Hua, Dongbo Bu, Shiwei Sun, Xingce Wang |
BIBM | 6 |
| 2020 | SVLR: Genome Structure Variant Detection Using Long Read Sequencing Data
Wenyan Gu, Aizhong Zhou, Lusheng Wang 0001, Shiwei Sun, Xuefeng Cui, Daming Zhu |
ISBRA | 4 |
| 2020 | ISSEC: inferring contacts among protein secondary structure elements using deep object detectionabstractBACKGROUND: The formation of contacts among protein secondary structure elements (SSEs) is an important step in protein folding as it determines topology of protein tertiary structure; hence, inferring inter-SSE contacts is crucial to protein structure prediction. One of the existing strategies infers inter-SSE contacts directly from the predicted possibilities of inter-residue contacts without any preprocessing, and thus suffers from the excessive noises existing in the predicted inter-residue contacts. Another strategy defines SSEs based on protein secondary structure prediction first, and then judges whether each candidate SSE pair could form contact or not. However, it is difficult to accurately determine boundary of SSEs due to the errors in secondary structure prediction. The incorrectly-deduced SSEs definitely hinder subsequent prediction of the contacts among them. RESULTS: We here report an accurate approach to infer the inter-SSE contacts (thus called as ISSEC) using the deep object detection technique. The design of ISSEC is based on the observation that, in the inter-residue contact map, the contacting SSEs usually form rectangle regions with characteristic patterns. Therefore, ISSEC infers inter-SSE contacts through detecting such rectangle regions. Unlike the existing approach directly using the predicted probabilities of inter-residue contact, ISSEC applies the deep convolution technique to extract high-level features from the inter-residue contacts. More importantly, ISSEC does not rely on the pre-defined SSEs. Instead, ISSEC enumerates multiple candidate rectangle regions in the predicted inter-residue contact map, and for each region, ISSEC calculates a confidence score to measure whether it has characteristic patterns or not. ISSEC employs greedy strategy to select non-overlapping regions with high confidence score, and finally infers inter-SSE contacts according to these regions. CONCLUSIONS: Comprehensive experimental results suggested that ISSEC outperformed the state-of-the-art approaches in predicting inter-SSE contacts. We further demonstrated the successful applications of ISSEC to improve prediction of both inter-residue contacts and tertiary structure as well. Jianwei Zhu, Fusong Ju, Lupeng Kong, Shiwei Sun, Wei-Mou Zheng, Dongbo Bu |
BMC Bioinform. | 5 |
| 2019 | Cerebrovascular Segmentation Algorithm Based on Focused Multi-Gaussians Model and Weighted 3D Markov Random FieldabstractSegmenting the cerebral vessels precisely from the time-of-flight magnetic resonance angiography (TOF-MRA) images is important for the diagnosis and therapy of the cerebrovascular diseases. Since the complex structures of cerebral vessels, the current cerebrovascular segmentation algorithms based on statistical model have less accuracy for stenotic vessels and are quite time-consuming. In this paper, we propose a novel automatic cerebrovascular segmentation algorithm based on focused Multi-Gaussians (FMG) model and weighted 3D Markov Random Field. As far as our knowledge, this is the first time to adopt multi-Gaussians distributions as vascular model with the purpose of modeling the vascular tissue more accurately. Furthermore, the fitting range is narrowed to local region related to vessels in order to make the model focus on the vascular tissue and simplify the finite mixture model. To incorporate precise local character of images to the model, we design a new weighted 3D MRF by a weighted neighborhood system (W-NBS). Finally, the particle swarm optimization (PSO) algorithm of parameter estimation has been implemented parallelly based on GPUs and the execution speed was improved by about 70 times. The experimental results show that the algorithm can produce detailed segmentation results especially for stenotic vessels. Zhilong Lv, Rui Yan 0009, Xinyu Liu 0008, Zhongke Wu, Yicheng Zhu, Shiwei Sun, Fa Zhang 0001, Xingce Wang |
BIBM | 6 |
| 2019 | Best-first search guided multistage mass spectrometry-based glycan identificationabstractMOTIVATION: Glycan identification has long been hampered by complicated branching patterns and various isomeric structures of glycans. Multistage mass spectrometry (MSn) is a promising glycan identification technique as it generates multiple-level fragments of a glycan, which can be explored to deduce branching pattern of the glycan and further distinguish it from other candidates with identical mass. However, the automatic glycan identification still remains a challenge since it mainly relies on expertise to guide a MSn instrument to generate spectra. RESULTS: Here, we proposed a novel method, named bestFSA, based on a best-first search algorithm to guide the process of spectrum producing in glycan identification using MSn. BestFSA is able to select the most appropriate peaks for next round of experiments and complete the identification using as few experimental rounds. Our analysis of seven representative glycans shows that bestFSA correctly distinguishes actual glycans efficiently and suggested bestFSA could be used in practical glycan identification. The combination of the MSn technology coupled with bestFSA should greatly facilitate the automatic identification of glycan branching patterns, with significantly improved identification sensitivity, and reduce time and cost of MSn experiments. AVAILABILITY AND IMPLEMENTATION: http://glycan.ict.ac.cn. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yaojun Wang, Dongbo Bu, Chuncui Huang, Junchuan Dong, Weiyi Pan, Shiwei Sun |
Bioinform. | 11 |
| 2019 | Constructing effective energy functions for protein structure prediction through broadening attraction-basin and reverse Monte Carlo samplingabstractBACKGROUND: The ab initio approaches to protein structure prediction usually employ the Monte Carlo technique to search the structural conformation that has the lowest energy. However, the widely-used energy functions are usually ineffective for conformation search. How to construct an effective energy function remains a challenging task. RESULTS: Here, we present a framework to construct effective energy functions for protein structure prediction. Unlike existing energy functions only requiring the native structure to be the lowest one, we attempt to maximize the attraction-basin where the native structure lies in the energy landscape. The underlying rationale is that each energy function determines a specific energy landscape together with a native attraction-basin, and the larger the attraction-basin is, the more likely for the Monte Carlo search procedure to find the native structure. Following this rationale, we constructed effective energy functions as follows: i) To explore the native attraction-basin determined by a certain energy function, we performed reverse Monte Carlo sampling starting from the native structure, identifying the structural conformations on the edge of attraction-basin. ii) To broaden the native attraction-basin, we smoothened the edge points of attraction-basin through tuning weights of energy terms, thus acquiring an improved energy function. Our framework alternates the broadening attraction-basin and reverse sampling steps (thus called BARS) until the native attraction-basin is sufficiently large. We present extensive experimental results to show that using the BARS framework, the constructed energy functions could greatly facilitate protein structure prediction in improving the quality of predicted structures and speeding up conformation search. CONCLUSION: Using the BARS framework, we constructed effective energy functions for protein structure prediction, which could improve the quality of predicted structures and speed up conformation search as well. Haicang Zhang, Lupeng Kong, Shiwei Sun, Wei-Mou Zheng, Dongbo Bu |
BMC Bioinform. | 5 |
| 2019 | Predicting protein inter-residue contacts using composite likelihood maximization and deep learningabstractBACKGROUND: Accurate prediction of inter-residue contacts of a protein is important to calculating its tertiary structure. Analysis of co-evolutionary events among residues has been proved effective in inferring inter-residue contacts. The Markov random field (MRF) technique, although being widely used for contact prediction, suffers from the following dilemma: the actual likelihood function of MRF is accurate but time-consuming to calculate; in contrast, approximations to the actual likelihood, say pseudo-likelihood, are efficient to calculate but inaccurate. Thus, how to achieve both accuracy and efficiency simultaneously remains a challenge. RESULTS: In this study, we present such an approach (called clmDCA) for contact prediction. Unlike plmDCA using pseudo-likelihood, i.e., the product of conditional probability of individual residues, our approach uses composite-likelihood, i.e., the product of conditional probability of all residue pairs. Composite likelihood has been theoretically proved as a better approximation to the actual likelihood function than pseudo-likelihood. Meanwhile, composite likelihood is still efficient to maximize, thus ensuring the efficiency of clmDCA. We present comprehensive experiments on popular benchmark datasets, including PSICOV dataset and CASP-11 dataset, to show that: i) clmDCA alone outperforms the existing MRF-based approaches in prediction accuracy. ii) When equipped with deep learning technique for refinement, the prediction accuracy of clmDCA was further significantly improved, suggesting the suitability of clmDCA for subsequent refinement procedure. We further present a successful application of the predicted contacts to accurately build tertiary structures for proteins in the PSICOV dataset. CONCLUSIONS: Composite likelihood maximization algorithm can efficiently estimate the parameters of Markov Random Fields and can improve the prediction accuracy of protein inter-residue contacts. Haicang Zhang, Fusong Ju, Jianwei Zhu, Yujuan Gao, Ziwei Xie, Minghua Deng, Shiwei Sun, Wei-Mou Zheng, Dongbo Bu |
BMC Bioinform. | 8 |
| 2019 | Correction to: Predicting protein inter-residue contacts using composite likelihood maximization and deep learningabstractFollowing publication of the original article [1], the author explained that there are several errors in the original article. Haicang Zhang, Fusong Ju, Jianwei Zhu, Yujuan Gao, Ziwei Xie, Minghua Deng, Shiwei Sun, Wei-Mou Zheng, Dongbo Bu |
BMC Bioinform. | 8 |
| 2018 | Fiducial marker detection via deep learning approach for electron tomography
Renmin Han, Fa Zhang 0001, Shiwei Sun |
BIBM | 5 |
| 2018 | Understanding the Factors Affecting the Organizational Adoption of Big DataabstractBig data is rapidly becoming a major driver for firms seeking to gain a competitive advantage. Grounded in the Diffusion of Innovation theory (DOI), the institutional theory, and the Tech-nology–Organization–Environment (TOE) framework, this study applies the results of a content analysis to develop a framework to identify the main factors affecting the organizational adoption of big data. The content analysis is based on the retrieval and review of relevant papers in the business intelligence & analytics (BI&A) literature published during the period 2009–2015. The 26 factors identified by this review are then integrated into a TOE framework. The findings of this research enrich the current big data literature and enhance practitioners’ understanding of the decision-making processes involved in a firm’s adoption of big data. Shiwei Sun, Casey G. Cegielski, Dianne J. Hall |
J. Comput. Inf. Syst. | 1 |
| 2017 | Improving protein fold recognition by extracting fold-specific features from predicted residue-residue contactsabstractMOTIVATION: Accurate recognition of protein fold types is a key step for template-based prediction of protein structures. The existing approaches to fold recognition mainly exploit the features derived from alignments of query protein against templates. These approaches have been shown to be successful for fold recognition at family level, but usually failed at superfamily/fold levels. To overcome this limitation, one of the key points is to explore more structurally informative features of proteins. Although residue-residue contacts carry abundant structural information, how to thoroughly exploit these information for fold recognition still remains a challenge. RESULTS: In this study, we present an approach (called DeepFR) to improve fold recognition at superfamily/fold levels. The basic idea of our approach is to extract fold-specific features from predicted residue-residue contacts of proteins using deep convolutional neural network (DCNN) technique. Based on these fold-specific features, we calculated similarity between query protein and templates, and then assigned query protein with fold type of the most similar template. DCNN has showed excellent performance in image feature extraction and image recognition; the rational underlying the application of DCNN for fold recognition is that contact likelihood maps are essentially analogy to images, as they both display compositional hierarchy. Experimental results on the LINDAHL dataset suggest that even using the extracted fold-specific features alone, our approach achieved success rate comparable to the state-of-the-art approaches. When further combining these features with traditional alignment-related features, the success rate of our approach increased to 92.3%, 82.5% and 78.8% at family, superfamily and fold levels, respectively, which is about 18% higher than the state-of-the-art approach at fold level, 6% higher at superfamily level and 1% higher at family level. An independent assessment on SCOP_TEST dataset showed consistent performance improvement, indicating robustness of our approach. Furthermore, bi-clustering results of the extracted features are compatible with fold hierarchy of proteins, implying that these features are fold-specific. Together, these results suggest that the features extracted from predicted contacts are orthogonal to alignment-related features, and the combination of them could greatly facilitate fold recognition at superfamily/fold levels and template-based prediction of protein structures. AVAILABILITY AND IMPLEMENTATION: Source code of DeepFR is freely available through https://github.com/zhujianwei31415/deepfr, and a web server is available through http://protein.ict.ac.cn/deepfr. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jianwei Zhu, Haicang Zhang, Shuaicheng Li 0001, Lupeng Kong, Shiwei Sun, Wei-Mou Zheng, Dongbo Bu |
Bioinform. | 6 |
| 2017 | Improving prediction of burial state of residues by exploiting correlation among residuesabstractBACKGROUND: Residues in a protein might be buried inside or exposed to the solvent surrounding the protein. The buried residues usually form hydrophobic cores to maintain the structural integrity of proteins while the exposed residues are tightly related to protein functions. Thus, the accurate prediction of solvent accessibility of residues will greatly facilitate our understanding of both structure and functionalities of proteins. Most of the state-of-the-art prediction approaches consider the burial state of each residue independently, thus neglecting the correlations among residues. RESULTS: In this study, we present a high-order conditional random field model that considers burial states of all residues in a protein simultaneously. Our approach exploits not only the correlation among adjacent residues but also the correlation among long-range residues. Experimental results showed that by exploiting the correlation among residues, our approach outperformed the state-of-the-art approaches in prediction accuracy. In-depth case studies also showed that by using the high-order statistical model, the errors committed by the bidirectional recurrent neural network and chain conditional random field models were successfully corrected. CONCLUSIONS: Our methods enable the accurate prediction of residue burial states, which should greatly facilitate protein structure prediction and evaluation. Hai'e Gong, Haicang Zhang, Jianwei Zhu, Shiwei Sun, Wei-Mou Zheng, Dongbo Bu |
BMC Bioinform. | 5 |
| 2016 | FALCON@home: a high-throughput protein structure prediction server based on remote homologue recognitionabstractSUMMARY: The protein structure prediction approaches can be categorized into template-based modeling (including homology modeling and threading) and free modeling. However, the existing threading tools perform poorly on remote homologous proteins. Thus, improving fold recognition for remote homologous proteins remains a challenge. Besides, the proteome-wide structure prediction poses another challenge of increasing prediction throughput. In this study, we presented FALCON@home as a protein structure prediction server focusing on remote homologue identification. The design of FALCON@home is based on the observation that a structural template, especially for remote homologous proteins, consists of conserved regions interweaved with highly variable regions. The highly variable regions lead to vague alignments in threading approaches. Thus, FALCON@home first extracts conserved regions from each template and then aligns a query protein with conserved regions only rather than the full-length template directly. This helps avoid the vague alignments rooted in highly variable regions, improving remote homologue identification. We implemented FALCON@home using the Berkeley Open Infrastructure of Network Computing (BOINC) volunteer computing protocol. With computation power donated from over 20,000 volunteer CPUs, FALCON@home shows a throughput as high as processing of over 1000 proteins per day. In the Critical Assessment of protein Structure Prediction (CASP11), the FALCON@home-based prediction was ranked the 12th in the template-based modeling category. As an application, the structures of 880 mouse mitochondria proteins were predicted, which revealed the significant correlation between protein half-lives and protein structural factors. AVAILABILITY AND IMPLEMENTATION: FALCON@home is freely available at http://protein.ict.ac.cn/FALCON/. CONTACT: [email protected], [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Haicang Zhang, Wei-Mou Zheng, Dong Xu 0002, Jianwei Zhu, Kang Ning 0001, Shiwei Sun, Shuaicheng Li 0001, Dongbo Bu |
Bioinform. | 8 |
| 2015 | Condensing Raman spectrum for single-cell phenotype analysisabstractBACKGROUND: In recent years, high throughput and non-invasive Raman spectrometry technique has matured as an effective approach to identification of individual cells by species, even in complex, mixed populations. Raman profiling is an appealing optical microscopic method to achieve this. To fully utilize Raman proling for single-cell analysis, an extensive understanding of Raman spectra is necessary to answer questions such as which filtering methodologies are effective for pre-processing of Raman spectra, what strains can be distinguished by Raman spectra, and what features serve best as Raman-based biomarkers for single-cells, etc. RESULTS: In this work, we have proposed an approach called rDisc to discretize the original Raman spectrum into only a few (usually less than 20) representative peaks (Raman shifts). The approach has advantages in removing noises, and condensing the original spectrum. In particular, effective signal processing procedures were designed to eliminate noise, utilising wavelet transform denoising, baseline correction, and signal normalization. In the discretizing process, representative peaks were selected to signicantly decrease the Raman data size. More importantly, the selected peaks are chosen as suitable to serve as key biological markers to differentiate species and other cellular features. Additionally, the classication performance of discretized spectra was found to be comparable to full spectrum having more than 1000 Raman shifts. Overall, the discretized spectrum needs about 5storage space of a full spectrum and the processing speed is considerably faster. This makes rDisc clearly superior to other methods for single-cell classication. Shiwei Sun, Xuetao Wang, Xin Gao 0001, Lihui Ren, Xiaoquan Su, Dongbo Bu, Kang Ning 0001 |
BMC Bioinform. | 1 |
| 2015 | OpenMS-Simulator: an open-source software for theoretical tandem mass spectrum predictionabstractBACKGROUND: Tandem mass spectrometry (MS/MS) acts as a key technique for peptide identification. The MS/MS-based peptide identification approaches can be categorized into two families, namely, de novo and database search. Both of the two types of approaches can benefit from an accurate prediction of theoretical spectrum. A theoretical spectrum consists of m/z and intensity of possibly occurring ions, which are estimated via simulating the spectrum generating process. Extensive researches have been conducted for theoretical spectrum prediction; however, the prediction methods suffer from low prediciton accuracy due to oversimplifications in the spectrum simulation process. RESULTS: In the study, we present an open-source software package, called OpenMS-Simulator, to predict theoretical spectrum for a given peptide sequence. Based on the mobile-proton hypothesis for peptide fragmentation, OpenMS-Simulator trained a closed-form model for the intensity ratio of adjacent y ions, from which the whole theoretical spectrum can be constructed. On a collection of representative spectra datasets with annotated peptide sequences, experimental results suggest that OpenMS-Simulator can predict theoretical spectra with considerable accuracy. The study also presents an application of OpenMS-Simulator: the similarity between theoretical spectra and query spectra can be used to re-rank the peptide sequence reported by SEQUEST/X!Tandem. CONCLUSIONS: OpenMS-Simulator implements a novel model to predict theoretical spectrum for a given peptide sequence. Compared with existing theoretical spectrum prediction tools, say MassAnalyzer and MSSimulator, our method not only simplifies the computation process, but also improves the prediction accuracy. Currently, OpenMS-Simulator supports the prediction of CID and HCD spectrum for peptides with double charges. The extension to cover more fragmentation models and support multiple-charged peptides remains as one of the future works. Yaojun Wang, Dongbo Bu, Shiwei Sun |
BMC Bioinform. | 5 |
| 2015 | User ratings analysis in social networks through a hypernetwork method
Qi Suo, Shiwei Sun, Mahmood Hajli, Peter E. D. Love |
Expert Syst. Appl. | 2 |
| 2011 | PI: An open-source software package for validation of the SEQUEST result and visualization of mass spectrumabstractBACKGROUND: Tandem mass spectrometry (MS/MS) has emerged as the leading method for high- throughput protein identification in proteomics. Recent technological breakthroughs have dramatically increased the efficiency of MS/MS data generation. Meanwhile, sophisticated algorithms have been developed for identifying proteins from peptide MS/MS data by searching available protein sequence databases for the peptide that is most likely to have produced the observed spectrum. The popular SEQUEST algorithm relies on the cross-correlation between the experimental mass spectrum and the theoretical spectrum of a peptide. It utilizes a simplified fragmentation model that assigns a fixed and identical intensity for all major ions and fixed and lower intensity for their neutral losses. In this way, the common issues involved in predicting theoretical spectra are circumvented. In practice, however, an experimental spectrum is usually not similar to its SEQUEST -predicted theoretical one, and as a result, incorrect identifications are often generated. RESULTS: Better understanding of peptide fragmentation is required to produce more accurate and sensitive peptide sequencing algorithms. Here, we designed the software PI of novel and exquisite algorithms that make a good use of intensity property of a spectrum. CONCLUSIONS: We designed the software PI with the novel and effective algorithms which made a good use of intensity property of the spectrum. Experiments have shown that PI was able to validate and improve the results of SEQUEST to a more satisfactory degree. Yantao Qiao, Dongbo Bu, Shiwei Sun |
BMC Bioinform. | 4 |
| 2011 | ProbPS: A new model for peak selection based on quantifying the dependence of the existence of derivative peaks on primary ion intensityabstractBACKGROUND: The analysis of mass spectra suggests that the existence of derivative peaks is strongly dependent on the intensity of the primary peaks. Peak selection from tandem mass spectrum is used to filter out noise and contaminant peaks. It is widely accepted that a valid primary peak tends to have high intensity and is accompanied by derivative peaks, including isotopic peaks, neutral loss peaks, and complementary peaks. Existing models for peak selection ignore the dependence between the existence of the derivative peaks and the intensity of the primary peaks. Simple models for peak selection assume that these two attributes are independent; however, this assumption is contrary to real data and prone to error. RESULTS: In this paper, we present a statistical model to quantitatively measure the dependence of the derivative peak's existence on the primary peak's intensity. Here, we propose a statistical model, named ProbPS, to capture the dependence in a quantitative manner and describe a statistical model for peak selection. Our results show that the quantitative understanding can successfully guide the peak selection process. By comparing ProbPS with AuDeNS we demonstrate the advantages of our method in both filtering out noise peaks and in improving de novo identification. In addition, we present a tag identification approach based on our peak selection method. Our results, using a test data set, suggest that our tag identification method (876 correct tags in 1000 spectra) outperforms PepNovoTag (790 correct tags in 1000 spectra). CONCLUSIONS: We have shown that ProbPS improves the accuracy of peak selection which further enhances the performance of de novo sequencing and tag identification. Thus, our model saves valuable computation time and improving the accuracy of the results. Shenghui Zhang, Yaojun Wang, Dongbo Bu, Shiwei Sun |
BMC Bioinform. | 5 |
| 2008 | A Fragmentation Event Model for Peptide Identification by Mass Spectrometry
Yu Lin 0001, Yantao Qiao, Shiwei Sun, Chungong Yu, Gongjin Dong, Dongbo Bu |
RECOMB | 3 |
| 2006 | Phylophenetic properties of metabolic pathway topologies as revealed by global analysisabstractBACKGROUND: As phenotypic features derived from heritable characters, the topologies of metabolic pathways contain both phylogenetic and phenetic components. In the post-genomic era, it is possible to measure the "phylophenetic" contents of different pathways topologies from a global perspective. RESULTS: We reconstructed phylophenetic trees for all available metabolic pathways based on topological similarities, and compared them to the corresponding 16S rRNA-based trees. Similarity values for each pair of trees ranged from 0.044 to 0.297. Using the quartet method, single pathways trees were merged into a comprehensive tree containing information from a large part of the entire metabolic networks. This tree showed considerably higher similarity (0.386) to the corresponding 16S rRNA-based tree than any tree based on a single pathway, but was, on the other hand, sufficiently distinct to preserve unique phylogenetic information not reflected by the 16S rRNA tree. CONCLUSION: We observed that the topology of different metabolic pathways provided different phylogenetic and phenetic information, depicting the compromise between phylogenetic information and varying evolutionary pressures forming metabolic pathway topologies in different organisms. The phylogenetic information content of the comprehensive tree is substantially higher than that of any tree based on a single pathway, which also gave clues to constraints working on the topology of the global metabolic networks, information that is only partly reflected by the topologies of individual metabolic pathways. Yong Zhang 0006, Shaojuan Li, Geir Skogerbø, Shiwei Sun, Hongchao Lu, Baochen Shi, Runsheng Chen |
BMC Bioinform. | 7 |
| 2006 | A novel scoring schema for peptide identification by searching protein sequence databases using tandem mass spectrometry dataabstractBACKGROUND: Tandem mass spectrometry (MS/MS) is a powerful tool for protein identification. Although great efforts have been made in scoring the correlation between tandem mass spectra and an amino acid sequence database, improvements could be made in three aspects, including characterization ofpeaks in spectra, adoption of effective scoring functions and access to thereliability of matching between peptides and spectra. RESULTS: A novel scoring function is presented, along with criteria to estimate the performance confidence of the function. Through learning the typesof product ions and the probability of generating them, a hypothetic spectrum was generated for each candidate peptide. Then relative entropy was introduced to measure the similarity between the hypothetic and the observed spectra. Based on the extreme value distribution (EVD) theory, a threshold was chosen to distinguish a true peptide assignment from a random one. Tests on a public MS/MS dataset demonstrated that this method performs better than the well-known SEQUEST. CONCLUSION: A reliable identification of proteins from the spectra promises a more efficient application of tandem mass spectrometry to proteomes with high complexity. Shiwei Sun, Suhua Chang, Chungong Yu, Dongbo Bu, Runsheng Chen |
BMC Bioinform. | 2 |