EDBT 2026 Demo / reviewers in the wild / expert
Qiwen Dong
dblp:01/4765
· DBLP profile ↗
36ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0002-3166-0541ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 7 first-author · 1 since 2021Artificial intelligence and machine learning · 10 · 5 since 2021Databases, data management, data science and information retrieval · 7 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accuracy-Aware Log Replay with Fine-Grained Prioritization for Real-Time Prediction Queries
Jing Jiang 0025, Peng Cai 0001, Qiwen Dong, Huiqi Hu |
DASFAA (2) | 4 |
| 2025 | Model-Accuracy Aware Query Routing for Smart Logistics ServiceabstractData driven business applications in logistics industry often issue prediction queries over relational databases to retrieve the newly generated transaction data for feature computations. In some cases even using slightly outdated data can result in significant inaccuracies in predictions. In another cases, we also observed the accuracy of model prediction is not sensitive to the data freshness. In the setting of primary-backup databases, one may choose to fetch the freshest data from the primary database to ensure model accuracy. However, this may hurt the performance of read-write transactions on the primary, especially when subjected to a high volume of prediction requests. In this work, we propose a Model-Accuracy Aware Service that facilitates a flexible trade-off between model prediction accuracy and primary database performance. This service implements an automated routing strategy aimed at minimizing the impact on the primary database's performance while meeting the requirements of model accuracy. It achieves this by leveraging the maintained database freshness information and the predictive results feedback from the model to learn the relationship between data discrepancy and prediction discrepancy in primary-backup scenarios. We report the experimental results on a real logistics application and also show its effectiveness on a public dataset. Zhiwei Ye, Peng Cai 0001, Qiwen Dong |
ICDE | 4 |
| 2024 | Job Title Prediction as a Dual Task of Expertise Prediction in Open Source Software
Xin Liu 0151, Yu Wang 0215, Qiwen Dong |
ECML/PKDD (10) | 3 |
| 2024 | AutoTable: Effective and Efficient Automated Feature Transformation for Tabular Data
Junpeng Zhu, Fengyan Zhang, Qiwen Dong |
WISE (1) | 5 |
| 2022 | Supervised Multi-view Latent Space Learning by Jointly Preserving Similarities Across Views and Samples
Martin Pavlovski, Qiwen Dong, Weining Qian, Zoran Obradovic |
DASFAA (2) | 4 |
| 2022 | An Exploratory Approach to Intelligent Quiz Question RecommendationabstractWith the rapid advancement of ICT, the digital transformation on education is greatly accelerating in various applications. As a particularly prominent application of digital education, quiz question recommendation is playing a vital role in precision teaching, smart tutoring, and personalized learning. However, the looming challenge of quiz question recommender for students is to satisfy the question diversity demands for students ZPD (the zone of proximal development) stage dynamically online. Therefore, we propose to formalize quiz question recommendation with a novel approach of reinforcement learning based two-sided recommender system. We develop a recommendation framework RTR (Reinforcement-Learning based Two-sided Recommender Systems) for taking into account the interests of both sides of the system, learning and adapting to those interests in real time, and resulting in more satisfactory recommended content. This established recommendation framework captures question characters and student dynamic preferences by considering the emergence of both sides of the system, and it yields a better learning experience in the context of practical quiz question generation. Kejie Mao, Qiwen Dong, Daocheng Hong |
KES | 2 |
| 2022 | Dynamic self-paced sampling ensemble for highly imbalanced and class-overlapped data classification
Suting Gao, Lyu Ni, Martin Pavlovski, Qiwen Dong, Zoran Obradovic, Weining Qian |
Data Min. Knowl. Discov. | 5 |
| 2022 | Scalable and adaptive log manager in distributed systems
Weining Qian, Xuan Zhou 0001, Qiwen Dong, Aoying Zhou, Wenrong Tan |
Frontiers Comput. Sci. | 4 |
| 2021 | Service-Oriented Data Processing for Dynamic SchemaabstractIn recent years, with the rapid development of big data technology, more and more Internet enterprises have started to transform and upgrade into big data driven enterprises. In addition, in the industrial field, information services driven by industrial big data have also attracted widespread attention. Some traditional industries, such as state-owned chemical enterprises are also actively transforming to information management and big data management. These traditional chemical enterprises have been using simple data processing tools internally in the past, such as Microsoft Excel. Although these tools are relatively simple to use and do not involve much learning cost, they cannot support situations where the volume of data increases and data types become more complex. These data processing methods suffer from insufficient capacity, performance degradation, management confusion and other problems, so that they are no longer suitable for modern data management work. For these traditional chemical companies, on the one hand they have many types of specialist data to store, such as simulation data, measurement data, formulation data, component data, process data, etc. On the other hand, they are unable to determine the full database table structure from the outset, and need to change it dynamically during use. This paper designs a new service-oriented data processing for dynamic schema to meet these needs, and applies it to the development of a data center web application platform for a state-owned chemical company. Zixin Chen, Zeqiu Fan, Daocheng Hong, Qiwen Dong |
KES | 6 |
| 2021 | Design and Implementation of Scientific Research Big Data Service Platform for Experimental Data ManagingabstractThe goal of the scientific research big data service platform for experimental data managing is to solve the problem of data silos in the state-owned scientific research management system caused by backward informatization. The new data service platform integrates data collection, data analysis, data governance, monitoring and management, prediction and early warning, and visualization platform. We are committed to improving data management and service capability with informatization, and to grasp the material development and design situation timely and accurately. The goal of our platform is to truly use data to speak, manage and make decisions with data. Zeqiu Fan, Zixin Chen, Daocheng Hong, Qiwen Dong |
KES | 6 |
| 2021 | Application of access control model for confidential dataabstractIn any field, the security of data is extremely important, and it is even related to national security and personal privacy. Within a mature system framework, the design of data security is the most basic and challenging task, and access control is one of the main strategies for Network security prevention. Log as an indispensable part of a secure system can help us to complete traceability after a data breach and to monitor the operation of the application at any time. However, in the existing confidential data management systems, the existing access control methods are not friendly to confidential data, and there are problems of excessive administrator privileges and no confidentiality restrictions. Considering of the fact that the authority and log Module is not well implemented in most confidential data management system, we propose to design a general access control model application. We propose an access control model based on roles and object domains, combined with a security level. Through this model, we can implement three-layer filtering when users access data, thereby ensuring data security and avoiding data leakage problems. At the same time, by implementing the log module, some deficiencies in the log analysis and monitoring of existing confidential data management system can be solved. Lumin Shan, Daocheng Hong, Qiwen Dong, Shubing Song |
KES | 4 |
| 2020 | Role and Object Domain-Based Access Control Model for Graduate Education Information SystemabstractWith the booming of Chinese education informatization 2.0, East China Normal University proposes to design a new generation system of graduate education to provide better services for teachers and students. Within the graduate education system, ensuring system service availability and data security has become the primary challenge, and access control is one of the main strategies for Network security prevention and protection [1]. Hence, we proposed the role and object domain-based access control model (RDBAC) which specifies the object domain category for each role based on prior works. In the new model, when the account is assigned a role, the system specifies a specific object domain instance to achieve more fine-grained access control to student objects. Besides, on the basis of the formulation of RESTful [2] API specification and Trie tree, a matching algorithm is proposed to optimize the matching efficiency between access requests and URL patterns for more efficient system authorization. Furthermore, a comparison experiment with the regular method verifies that the Trie tree method has good performance on graduate education system including URL pattern construction, matching, and scalability. Our research also establishes that future advance of access control is a valuable avenue for education system development and will inspire much more design research for education information systems. Gangzeng Jin, Daojiang Wang, Daocheng Hong, Qiwen Dong |
KES | 5 |
| 2020 | Humming-Query and Reinforcement-Learning based Modeling Approach for Personalized Music RecommendationabstractMusic recommendation is a prominent application of recommender systems, which has been attracting more and more attentions. There are two research streams of music recommender systems: one is static recommendation based on learning user’s preference according to historical data, and the other is dynamic recommendation considering user’s feedback. But the individual music preference for a certain moment is closely related to personal experience of the music and music literacy, as well as temporal scenario with diversity. Thus, it’s necessary to design a new music recommendation framework by integrating static recommendation and dynamic recommendation. Therefore, we propose a novel approach for music recommendation HRRS (Humming-Query and Reinforcement-Learning based Recommender Systems) by integrating prior two research streams. This novel recommendation framework HRRS based on humming query and reinforcement learning is learning and adapting to user’s current preference continually by collecting interactive data in real time. This preliminary recommendation framework captures song characters, personal dynamic preferences, and yields a better listening experience with proper interaction. Dezhuang Miao, Qiwen Dong, Daocheng Hong |
KES | 3 |
| 2020 | A Novel Application of Educational Management Information System based on Micro FrontendsabstractWith the launch of the Education Informatization 2.0 action plan by the Ministry of Education, a large number of college information systems have been born in China. Most of these systems are single page web applications (SPA) based on traditional MVC structures. Due to the complex logic and high coupling between educational businesses, developers need to write a lot of code. The education information system has many businesses and high coupling between businesses that the system often face problems such as bloated frontend businesses, iterative system updates, and difficult incremental function developments. Combined with the idea of service-oriented architecture, this paper proposes a micro frontends solution and applies it to the new generation of graduate information platform of East China Normal University, which has better agile development capabilities. From the aspects of service separation, efficient development, and incremental upgrade, this paper verifies that the architecture can well adapt to the needs of future educational management information system. The design of the micro frontends provides a new idea for the development of a new generation of education information system. Daojiang Wang, Dongming Yang, Daocheng Hong, Qiwen Dong, Shubing Song |
KES | 6 |
| 2020 | DevOps in Practice for Education Management Information System at ECNUabstractWith the rapid development of the Internet, the education information systems have become more prevalent aligning with better management to produce better education. However, the limitations of prior education systems development are gradually exposed, which ignore the changing requirements, the high concurrency bottlenecks and lean development of education information systems. Therefore, we develop and build a novel education information system at ECNU based on DevOps and related techniques. This paper reveals the practice of DevOps for new education information system from four aspects: CI (Continuous Integration), CD (Continuous Deployment), log management, and code quality. Meanwhile, brief technical explanations include Git, Jenkins, Kubernetes, ELK, SonarQube, etc. Through our continuous engineering practice, the new education information system has been developed and implemented at ECNU. The DevOps practice for information system establishes that it is so convenient for developing, testing and release of education information systems, and it also improves reliability, availability and scalability of information platform especially considering the guarantee of efficiency. Daojiang Wang, Dongming Yang, Qiwen Dong, Daocheng Hong |
KES | 4 |
| 2020 | A novel application integration architecture for the education industryabstractSince the Ministry of Education launched Education Informatization 2.0, the digitalization of colleges and universities has entered a stage of rapid growth. However, after more than 20 years of construction, problems such as system barriers and information islands have emerged in the digital construction of university systems. In order to solve such problems between the university systems, this paper proposes an easily expandable and configurable open information integration architecture by considering traditional information integration methods and combining with Web service technology. The architecture handles user service invocation information through a service layer, and manages the registration and invocation of services through a service module. The permission module manages user permissions to prevent information leakage and security issues. The data module abstracts data-related services to provide a basis for the deep use of data. And other optional development services are designed to satisfy special requirements for different platforms. The architecture proposed in this paper can integrate different heterogeneous subsystems in colleges and universities, eliminating the problem of system barriers and information islands, and providing specifications for the construction of new applications. Dongming Yang, Daojiang Wang, Shubing Song, Qiwen Dong |
KES | 6 |
| 2020 | Nonintrusive-Sensing and Reinforcement-Learning Based Adaptive Personalized Music RecommendationabstractAs a particularly prominent application of recommender systems on automated personalized service, the music recommendation has been widely used in various music network platforms, music education and music therapy. Importantly, the individual music preference for a certain moment is closely related to personal experience of the music and music literacy, as well as temporal scenario without any interruption. Therefore, this paper proposes a novel policy for music recommendation NRRS (Nonintrusive-Sensing and Reinforcement-Learning based Recommender Systems) by integrating prior research streams. Specifically, we develop a novel recommendation framework for sensing, learning and adaptation to user's current preference based on wireless sensing and reinforcement learning in real time during a listening session. The established music recommendation prototype monitors individual vital signals for listening music, and captures song characters, individual dynamic preferences, and that it can yield better listening experience for users. Daocheng Hong, Qiwen Dong |
SIGIR | 3 |
| 2020 | Amino Acid Encoding Methods for Protein Sequences: A Comprehensive Review and AssessmentabstractAs the first step of machine-learning based protein structure and function prediction, the amino acid encoding play a fundamental role in the final success of those methods. Different from the protein sequence encoding, the amino acid encoding can be used in both residue-level and sequence-level prediction of protein properties by combining them with different algorithms. However, it has not attracted enough attention in the past decades, and there are no comprehensive reviews and assessments about encoding methods so far. In this article, we make a systematic classification and propose a comprehensive review and assessment for various amino acid encoding methods. Those methods are grouped into five categories according to their information sources and information extraction methodologies, including binary encoding, physicochemical properties encoding, evolution-based encoding, structure-based encoding, and machine-learning encoding. Then, 16 representative methods from five categories are selected and compared on protein secondary structure prediction and protein fold recognition tasks by using large-scale benchmark datasets. The results show that the evolution-based position-dependent encoding method PSSM achieved the best performance, and the structure-based and machine-learning encoding methods also show some potential for further application, the neural network based distributed representation of amino acids in particular may bring new light to this area. We hope that the review and assessment are useful for future studies in amino acid encoding. Xiaoyang Jing, Qiwen Dong, Daocheng Hong, Ruqian Lu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2019 | Predicting protein-ligand binding residues with deep convolutional neural networksabstractBACKGROUND: Ligand-binding proteins play key roles in many biological processes. Identification of protein-ligand binding residues is important in understanding the biological functions of proteins. Existing computational methods can be roughly categorized as sequence-based or 3D-structure-based methods. All these methods are based on traditional machine learning. In a series of binding residue prediction tasks, 3D-structure-based methods are widely superior to sequence-based methods. However, due to the great number of proteins with known amino acid sequences, sequence-based methods have considerable room for improvement with the development of deep learning. Therefore, prediction of protein-ligand binding residues with deep learning requires study. RESULTS: In this study, we propose a new sequence-based approach called DeepCSeqSite for ab initio protein-ligand binding residue prediction. DeepCSeqSite includes a standard edition and an enhanced edition. The classifier of DeepCSeqSite is based on a deep convolutional neural network. Several convolutional layers are stacked on top of each other to extract hierarchical features. The size of the effective context scope is expanded as the number of convolutional layers increases. The long-distance dependencies between residues can be captured by the large effective context scope, and stacking several layers enables the maximum length of dependencies to be precisely controlled. The extracted features are ultimately combined through one-by-one convolution kernels and softmax to predict whether the residues are binding residues. The state-of-the-art ligand-binding method COACH and some of its submethods are selected as baselines. The methods are tested on a set of 151 nonredundant proteins and three extended test sets. Experiments show that the improvement of the Matthews correlation coefficient (MCC) is no less than 0.05. In addition, a training data augmentation method that slightly improves the performance is discussed in this study. CONCLUSIONS: Without using any templates that include 3D-structure data, DeepCSeqSite significantlyoutperforms existing sequence-based and 3D-structure-based methods, including COACH. Augmentation of the training sets slightly improves the performance. The model, code and datasets are available at https://github.com/yfCuiFaith/DeepCSeqSite . Yifeng Cui, Qiwen Dong, Daocheng Hong, Xikun Wang |
BMC Bioinform. | 2 |
| 2017 | MQAPRank: improved global protein model quality assessment by learning-to-rankabstractBACKGROUND: Protein structure prediction has achieved a lot of progress during the last few decades and a greater number of models for a certain sequence can be predicted. Consequently, assessing the qualities of predicted protein models in perspective is one of the key components of successful protein structure prediction. Over the past years, a number of methods have been developed to address this issue, which could be roughly divided into three categories: single methods, quasi-single methods and clustering (or consensus) methods. Although these methods achieve much success at different levels, accurate protein model quality assessment is still an open problem. RESULTS: Here, we present the MQAPRank, a global protein model quality assessment program based on learning-to-rank. The MQAPRank first sorts the decoy models by using single method based on learning-to-rank algorithm to indicate their relative qualities for the target protein. And then it takes the first five models as references to predict the qualities of other models by using average GDT_TS scores between reference models and other models. Benchmarked on CASP11 and 3DRobot datasets, the MQAPRank achieved better performances than other leading protein model quality assessment methods. Recently, the MQAPRank participated in the CASP12 under the group name FDUBio and achieved the state-of-the-art performances. CONCLUSIONS: The MQAPRank provides a convenient and powerful tool for protein model quality assessment with the state-of-the-art performances, it is useful for protein structure prediction and model quality assessment usages. Xiaoyang Jing, Qiwen Dong |
BMC Bioinform. | 2 |
| 2017 | RRCRank: a fusion method using rank strategy for residue-residue contact predictionabstractBACKGROUND: In structural biology area, protein residue-residue contacts play a crucial role in protein structure prediction. Some researchers have found that the predicted residue-residue contacts could effectively constrain the conformational search space, which is significant for de novo protein structure prediction. In the last few decades, related researchers have developed various methods to predict residue-residue contacts, especially, significant performance has been achieved by using fusion methods in recent years. In this work, a novel fusion method based on rank strategy has been proposed to predict contacts. Unlike the traditional regression or classification strategies, the contact prediction task is regarded as a ranking task. First, two kinds of features are extracted from correlated mutations methods and ensemble machine-learning classifiers, and then the proposed method uses the learning-to-rank algorithm to predict contact probability of each residue pair. RESULTS: First, we perform two benchmark tests for the proposed fusion method (RRCRank) on CASP11 dataset and CASP12 dataset respectively. The test results show that the RRCRank method outperforms other well-developed methods, especially for medium and short range contacts. Second, in order to verify the superiority of ranking strategy, we predict contacts by using the traditional regression and classification strategies based on the same features as ranking strategy. Compared with these two traditional strategies, the proposed ranking strategy shows better performance for three contact types, in particular for long range contacts. Third, the proposed RRCRank has been compared with several state-of-the-art methods in CASP11 and CASP12. The results show that the RRCRank could achieve comparable prediction precisions and is better than three methods in most assessment metrics. CONCLUSIONS: The learning-to-rank algorithm is introduced to develop a novel rank-based method for the residue-residue contact prediction of proteins, which achieves state-of-the-art performance based on the extensive assessment. Xiaoyang Jing, Qiwen Dong, Ruqian Lu |
BMC Bioinform. | 2 |
| 2016 | Improved protein residue-residue contacts prediction using learning-to-rankabstractProtein residue-residue contacts dictate the topology of protein structure and play an important role in structural biology, especially in de novo protein structure prediction. Accurate prediction of residue contacts could improve the performance of de novo protein structure prediction methods. In this study, a novel method based on learning-to-rank (RRCRank) has been presented to predict protein residue-residue contacts. The proposed method formulates the contacts prediction problem as a ranking problem. Firstly, the contact probabilities of residue pairs are predicted by ensemble machine-learning classifiers and correlated mutations approaches. And then, the proposed method integrates the complementary outputs of machine-learning and correlated mutations approaches and uses the learning-to-rank algorithm to rank residue pairs based on their probabilities to be contacts. Benchmarked on the CASP11 dataset, the proposed method achieves an improved performance for all three categories of contacts (short-range, medium-range and long-range contacts), which shows the proposed method based on learning-to-rank could take advantage of machine-learning and correlated mutations approaches and could provide the state-of-the-art performance. Xiaoyang Jing, Qiwen Dong |
BIBM | 2 |
| 2016 | Recognizing metal and acid radical ion-binding sites by integrating ab initio modeling with template-based transferalsabstractMOTIVATION: More than half of proteins require binding of metal and acid radical ions for their structure and function. Identification of the ion-binding locations is important for understanding the biological functions of proteins. Due to the small size and high versatility of the metal and acid radical ions, however, computational prediction of their binding sites remains difficult. RESULTS: ) that are most frequently seen in protein databases. A sequence-based ab initio model is first trained on sequence profiles, where a modified AdaBoost algorithm is extended to balance binding and non-binding residue samples. A composite method IonCom is then developed to combine the ab initio model with multiple threading alignments for further improving the robustness of the binding site predictions. The pipeline was tested using 5-fold cross validations on a comprehensive set of 2,100 non-redundant proteins bound with 3,075 small ion ligands. Significant advantage was demonstrated compared with the state of the art ligand-binding methods including COACH and TargetS for high-accuracy ion-binding site identification. Detailed data analyses show that the major advantage of IonCom lies at the integration of complementary ab initio and template-based components. Ion-specific feature design and binding library selection also contribute to the improvement of small ion ligand binding predictions. AVAILABILITY AND IMPLEMENTATION: http://zhanglab.ccmb.med.umich.edu/IonCom CONTACT: [email protected] or [email protected] information: Supplementary data are available at Bioinformatics online. Xiuzhen Hu, Qiwen Dong, Jianyi Yang 0002, Yang Zhang 0040 |
Bioinform. | 2 |
| 2016 | Protein ligand-specific binding residue predictions by an ensemble classifierabstractBACKGROUND: Prediction of ligand binding sites is important to elucidate protein functions and is helpful for drug design. Although much progress has been made, many challenges still need to be addressed. Prediction methods need to be carefully developed to account for chemical and structural differences between ligands. RESULTS: In this study, we present ligand-specific methods to predict the binding sites of protein-ligand interactions. First, a sequence-based method is proposed that only extracts features from protein sequence information, including evolutionary conservation scores and predicted structure properties. An improved AdaBoost algorithm is applied to address the serious imbalance problem between the binding and non-binding residues. Then, a combined method is proposed that combines the current template-free method and four other well-established template-based methods. The above two methods predict the ligand binding sites along the sequences using a ligand-specific strategy that contains metal ions, acid radical ions, nucleotides and ferroheme. Testing on a well-established dataset showed that the proposed sequence-based method outperformed the profile-based method by 4-19% in terms of the Matthews correlation coefficient on different ligands. The combined method outperformed each of the individual methods, with an improvement in the average Matthews correlation coefficients of 5.55% over all ligands. The results also show that the ligand-specific methods significantly outperform the general-purpose methods, which confirms the necessity of developing elaborate ligand-specific methods for ligand binding site prediction. CONCLUSIONS: Two efficient ligand-specific binding site predictors are presented. The standalone package is freely available for academic usage at http://dase.ecnu.edu.cn/qwdong/TargetCom/TargetCom_standalone.tar.gz or request upon the corresponding author. Xiuzhen Hu, Qiwen Dong |
BMC Bioinform. | 3 |
| 2015 | Identification of DNA-binding proteins by auto-cross covariance transformationabstractDNA-binding proteins play a pivotal role in various intra- and extra-cellular activities ranging from DNA replication to gene expression control. With the rapid development of next generation of sequencing technique, the number of protein sequences are unprecedentedly increasing. Thus it is necessary to develop computational methods to identify the DNA-binding protein from the protein sequence information. In this study, a novel method is presented which combines the support vector machine and the auto-cross covariance transformation. The protein sequence represented in the form of amino acids or the physical-chemical properties of amino acids are converted into a series of fixed-length vectors by Kmer composition and the auto-cross covariance transformation. The sequence order effect can be effectively capture by this scheme. These vectors are then inputted to support vector machine to discriminate the DNA-binding proteins from the non DNA-binding ones. The proposed method achieves the overall accuracy of 75.23% and Matthew correlation coefficient of 0.5 by a rigorous jackknife test. The independent test shows that the proposed method outperforms most of the existing methods. These results demonstrate that the proposed method provides the state-of-the-art performance for the prediction of DNA-binding proteins. Qiwen Dong, Shanyi Wang |
BIBM | 1 |
| 2015 | Protein model quality assessment by learning-to-rankabstractProtein structures are essential to understand the function. The predicted models have a broad range of the accuracy. Reliable estimates of the model quality are critical in determining the usefulness of the model to address a specific problem. In this study, a novel method has been presented to rank the models by their relative qualities. The proposed method first extracts various features from the three dimensional structures of proteins and then the learning-to-rank algorithm is used to rank the models based on their similarities with the native structures. Furthermore, a quasi single-model method is presented, which uses the top five identified models as references and ranks the other models by the average similarity with the reference models. Benchmark test is performed on a newly developed, template-based decoy generators which covers all the main structure classes of proteins. The proposed learning-to-rank method achieves an average Pearson correlation coefficient of 0.94 and a AUC value of 0.97, which consistently outperform all other well-developed methods. The quasi single-model can further improves the performance and achieve nearly perfect results with both PCC and AUC value of 0.99. The results demonstrate that the proposed method is an effective methodology for model quality assessment and provides the state-of-the-art performance. Xiaoyang Jing, Qiwen Dong |
BIBM | 2 |
| 2014 | Combining evolutionary information extracted from frequency profiles with sequence-based kernels for protein remote homology detectionabstractMOTIVATION: Owing to its importance in both basic research (such as molecular evolution and protein attribute prediction) and practical application (such as timely modeling the 3D structures of proteins targeted for drug development), protein remote homology detection has attracted a great deal of interest. It is intriguing to note that the profile-based approach is promising and holds high potential in this regard. To further improve protein remote homology detection, a key step is how to find an optimal means to extract the evolutionary information into the profiles. RESULTS: Here, we propose a novel approach, the so-called profile-based protein representation, to extract the evolutionary information via the frequency profiles. The latter can be calculated from the multiple sequence alignments generated by PSI-BLAST. Three top performing sequence-based kernels (SVM-Ngram, SVM-pairwise and SVM-LA) were combined with the profile-based protein representation. Various tests were conducted on a SCOP benchmark dataset that contains 54 families and 23 superfamilies. The results showed that the new approach is promising, and can obviously improve the performance of the three kernels. Furthermore, our approach can also provide useful insights for studying the features of proteins in various families. It has not escaped our notice that the current approach can be easily combined with the existing sequence-based methods so as to improve their performance as well. AVAILABILITY AND IMPLEMENTATION: For users' convenience, the source code of generating the profile-based proteins and the multiple kernel learning was also provided at http://bioinformatics.hitsz.edu.cn/main/~binliu/remote/ Bin Liu 0014, Deyuan Zhang, Ruifeng Xu 0001, Jinghao Xu, Xiaolong Wang 0001, Qingcai Chen, Qiwen Dong, Kuo-Chen Chou |
Bioinform. | 7 |
| 2011 | Novel Nonlinear Knowledge-Based Mean Force Potentials Based on Machine LearningabstractThe prediction of 3D structures of proteins from amino acid sequences is one of the most challenging problems in molecular biology. An essential task for solving this problem with coarse-grained models is to deduce effective interaction potentials. The development and evaluation of new energy functions is critical to accurately modeling the properties of biological macromolecules. Knowledge-based mean force potentials are derived from statistical analysis of proteins of known structures. Current knowledge-based potentials are almost in the form of weighted linear sum of interaction pairs. In this study, a class of novel nonlinear knowledge-based mean force potentials is presented. The potential parameters are obtained by nonlinear classifiers, instead of relative frequencies of interaction pairs against a reference state or linear classifiers. The support vector machine is used to derive the potential parameters on data sets that contain both native structures and decoy structures. Five knowledge-based mean force Boltzmann-based or linear potentials are introduced and their corresponding nonlinear potentials are implemented. They are the DIH potential (single-body residue-level Boltzmann-based potential), the DFIRE-SCM potential (two-body residue-level Boltzmann-based potential), the FS potential (two-body atom-level Boltzmann-based potential), the HR potential (two-body residue-level linear potential), and the T32S3 potential (two-body atom-level linear potential). Experiments are performed on well-established decoy sets, including the LKF data set, the CASP7 data set, and the Decoys “R”Us data set. The evaluation metrics include the energy Z score and the ability of each potential to discriminate native structures from a set of decoy structures. Experimental results show that all nonlinear potentials significantly outperform the corresponding Boltzmann-based or linear potentials, and the proposed discriminative framework is effective in developing knowledge-based mean force potentials. The nonlinear potentials can be widely used for ab initio protein structure prediction, model quality assessment, protein docking, and other challenging problems in computational biology. Qiwen Dong, Shuigeng Zhou |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2009 | Empirical Probability Functions Derived from Dihedral Angles for Protein Structure PredictionabstractThe development and evaluation of functions for protein energetics is an important part of current research aiming at understanding protein structures and functions. Knowledgebase mean force potentials are derived from statistical analysis of interacting groups in experimentally determined protein structures. Current knowledge-based mean force potentials are based on the inverse Boltzmannpsilas law, which calculate the ratio of the observed probability with respect to the probability of the reference state. In this study, a general probability framework is presented with the aim to develop novel energy scores. A class of empirical probability functions is derived by decomposing the joint probability of backbone dihedral angles and amino acid sequences. The neighboring interactions are modeled by conditional probabilities. Such probability functions are based on the strict probability theory and some suitable suppositions for convenience of computation. Experiments are performed on several well-constructed decoy sets and the results show that the empirical probability functions presented here outperform previous statistical potentials based on dihedral angles. Such probability functions will be helpful for protein structure prediction,model quality evaluation, transcription factors identification and other challenging problems in computational biology. Qiwen Dong, Shuigeng Zhou, Jihong Guan |
BIBE | 1 |
| 2009 | A new taxonomy-based protein fold recognition approach based on autocross-covariance transformationabstractMOTIVATION: Fold recognition is an important step in protein structure and function prediction. Traditional sequence comparison methods fail to identify reliable homologies with low sequence identity, while the taxonomic methods are effective alternatives, but their prediction accuracies are around 70%, which are still relatively low for practical usage. RESULTS: In this study, a simple and powerful method is presented for taxonomic fold recognition, which combines support vector machine (SVM) with autocross-covariance (ACC) transformation. The evolutionary information represented in the form of position-specific score matrices is converted into a series of fixed-length vectors by ACC transformation and these vectors are then input to a SVM classifier for fold recognition. The sequence-order effect can be effectively captured by this scheme. Experiments are performed on the widely used D-B dataset and the corresponding extended dataset, respectively. The proposed method, called ACCFold, gets an overall accuracy of 70.1% on the D-B dataset, which is higher than major existing taxonomic methods by 2-14%. Furthermore, the method achieves an overall accuracy of 87.6% on the extended dataset, which surpasses major existing taxonomic methods by 9-17%. Additionally, our method obtains an overall accuracy of 80.9% for 86-folds and 77.2% for 199-folds. These results demonstrate that the ACCFold method provides the state-of-the-art performance for taxonomic fold recognition. AVAILABILITY: The source code for ACC transformation is freely available at http://www.iipl.fudan.edu.cn/demo/accpkg.html. Qiwen Dong, Shuigeng Zhou, Jihong Guan |
Bioinform. | 1 |
| 2009 | Prediction of protein-protein interaction sites using an ensemble methodabstractBACKGROUND: Prediction of protein-protein interaction sites is one of the most challenging and intriguing problems in the field of computational biology. Although much progress has been achieved by using various machine learning methods and a variety of available features, the problem is still far from being solved. RESULTS: In this paper, an ensemble method is proposed, which combines bootstrap resampling technique, SVM-based fusion classifiers and weighted voting strategy, to overcome the imbalanced problem and effectively utilize a wide variety of features. We evaluate the ensemble classifier using a dataset extracted from 99 polypeptide chains with 10-fold cross validation, and get a AUC score of 0.86, with a sensitivity of 0.76 and a specificity of 0.78, which are better than that of the existing methods. To improve the usefulness of the proposed method, two special ensemble classifiers are designed to handle the cases of missing homologues and structural information respectively, and the performance is still encouraging. The robustness of the ensemble method is also evaluated by effectively classifying interaction sites from surface residues as well as from all residues in proteins. Moreover, we demonstrate the applicability of the proposed method to identify interaction sites from the non-structural proteins (NS) of the influenza A virus, which may be utilized as potential drug target sites. CONCLUSION: Our experimental results show that the ensemble classifiers are quite effective in predicting protein interaction sites. The Sub-EnClassifiers with resampling technique can alleviate the imbalanced problem and the combination of Sub-EnClassifiers with a wide variety of feature groups can significantly improve prediction performance. Lei Deng 0002, Jihong Guan, Qiwen Dong, Shuigeng Zhou |
BMC Bioinform. | 3 |
| 2009 | Prediction of protein binding sites in protein structures using hidden Markov support vector machineabstractBACKGROUND: Predicting the binding sites between two interacting proteins provides important clues to the function of a protein. Recent research on protein binding site prediction has been mainly based on widely known machine learning techniques, such as artificial neural networks, support vector machines, conditional random field, etc. However, the prediction performance is still too low to be used in practice. It is necessary to explore new algorithms, theories and features to further improve the performance. RESULTS: In this study, we introduce a novel machine learning model hidden Markov support vector machine for protein binding site prediction. The model treats the protein binding site prediction as a sequential labelling task based on the maximum margin criterion. Common features derived from protein sequences and structures, including protein sequence profile and residue accessible surface area, are used to train hidden Markov support vector machine. When tested on six data sets, the method based on hidden Markov support vector machine shows better performance than some state-of-the-art methods, including artificial neural networks, support vector machines and conditional random field. Furthermore, its running time is several orders of magnitude shorter than that of the compared methods. CONCLUSION: The improved prediction performance and computational efficiency of the method based on hidden Markov support vector machine can be attributed to the following three factors. Firstly, the relation between labels of neighbouring residues is useful for protein binding site prediction. Secondly, the kernel trick is very advantageous to this field. Thirdly, the complexity of the training step for hidden Markov support vector machine is linear with the number of training samples by using the cutting-plane algorithm. Bin Liu 0014, Xiaolong Wang 0001, Lei Lin 0001, Buzhou Tang, Qiwen Dong, Xuan Wang 0002 |
BMC Bioinform. | 5 |
| 2008 | A discriminative method for protein remote homology detection and fold recognition combining Top-n-grams and latent semantic analysisabstractBACKGROUND: Protein remote homology detection and fold recognition are central problems in bioinformatics. Currently, discriminative methods based on support vector machine (SVM) are the most effective and accurate methods for solving these problems. A key step to improve the performance of the SVM-based methods is to find a suitable representation of protein sequences. RESULTS: In this paper, a novel building block of proteins called Top-n-grams is presented, which contains the evolutionary information extracted from the protein sequence frequency profiles. The protein sequence frequency profiles are calculated from the multiple sequence alignments outputted by PSI-BLAST and converted into Top-n-grams. The protein sequences are transformed into fixed-dimension feature vectors by the occurrence times of each Top-n-gram. The training vectors are evaluated by SVM to train classifiers which are then used to classify the test protein sequences. We demonstrate that the prediction performance of remote homology detection and fold recognition can be improved by combining Top-n-grams and latent semantic analysis (LSA), which is an efficient feature extraction technique from natural language processing. When tested on superfamily and fold benchmarks, the method combining Top-n-grams and LSA gives significantly better results compared to related methods. CONCLUSION: The method based on Top-n-grams significantly outperforms the methods based on many other building blocks including N-grams, patterns, motifs and binary profiles. Therefore, Top-n-gram is a good building block of the protein sequences and can be widely used in many tasks of the computational biology, such as the sequence alignment, the prediction of domain boundary, the designation of knowledge-based potentials and the prediction of protein binding sites. Bin Liu 0014, Xiaolong Wang 0001, Lei Lin 0001, Qiwen Dong, Xuan Wang 0002 |
BMC Bioinform. | 4 |
| 2007 | Exploiting residue-level and profile-level interface propensities for usage in binding sites prediction of proteinsabstractBACKGROUND: Recognition of binding sites in proteins is a direct computational approach to the characterization of proteins in terms of biological and biochemical function. Residue preferences have been widely used in many studies but the results are often not satisfactory. Although different amino acid compositions among the interaction sites of different complexes have been observed, such differences have not been integrated into the prediction process. Furthermore, the evolution information has not been exploited to achieve a more powerful propensity. RESULT: In this study, the residue interface propensities of four kinds of complexes (homo-permanent complexes, homo-transient complexes, hetero-permanent complexes and hetero-transient complexes) are investigated. These propensities, combined with sequence profiles and accessible surface areas, are inputted to the support vector machine for the prediction of protein binding sites. Such propensities are further improved by taking evolutional information into consideration, which results in a class of novel propensities at the profile level, i.e. the binary profiles interface propensities. Experiment is performed on the 1139 non-redundant protein chains. Although different residue interface propensities among different complexes are observed, the improvement of the classifier with residue interface propensities can be negligible in comparison with that without propensities. The binary profile interface propensities can significantly improve the performance of binding sites prediction by about ten percent in term of both precision and recall. CONCLUSION: Although there are minor differences among the four kinds of complexes, the residue interface propensities cannot provide efficient discrimination for the complicated interfaces of proteins. The binary profile interface propensities can significantly improve the performance of binding sites prediction of protein, which indicates that the propensities at the profile level are more accurate than those at the residue level. Qiwen Dong, Xiaolong Wang 0001, Lei Lin 0001, Yi Guan |
BMC Bioinform. | 1 |
| 2006 | Application of latent semantic analysis to protein remote homology detectionabstractMOTIVATION: Remote homology detection between protein sequences is a central problem in computational biology. The discriminative method such as the support vector machine (SVM) is one of the most effective methods. Many of the SVM-based methods focus on finding useful representations of protein sequence, using either explicit feature vector representations or kernel functions. Such representations may suffer from the peaking phenomenon in many machine-learning methods because the features are usually very large and noise data may be introduced. Based on these observations, this research focuses on feature extraction and efficient representation of protein vectors for SVM protein classification. RESULTS: In this study, a latent semantic analysis (LSA) model, which is an efficient feature extraction technique from natural language processing, has been introduced in protein remote homology detection. Several basic building blocks of protein sequences have been investigated as the 'words' of 'protein sequence language', including N-grams, patterns and motifs. Each protein sequence is taken as a 'document' that is composed of bags-of-word. The word-document matrix is constructed first. The LSA is performed on the matrix to produce the latent semantic representation vectors of protein sequences, leading to noise-removal and smart description of protein sequences. The latent semantic representation vectors are then evaluated by SVM. The method is tested on the SCOP 1.53 database. The results show that the LSA model significantly improves the performance of remote homology detection in comparison with the basic formalisms. Furthermore, the performance of this method is comparable with that of the complex kernel methods such as SVM-LA and better than that of other sequence-based methods such as PSI-BLAST and SVM-pairwise. Qiwen Dong, Xiaolong Wang 0001, Lei Lin 0001 |
Bioinform. | 1 |
| 2006 | Novel knowledge-based mean force potential at the profile levelabstractBACKGROUND: The development and testing of functions for the modeling of protein energetics is an important part of current research aimed at understanding protein structure and function. Knowledge-based mean force potentials are derived from statistical analyses of interacting groups in experimentally determined protein structures. Current knowledge-based mean force potentials are developed at the atom or amino acid level. The evolutionary information contained in the profiles is not investigated. Based on these observations, a class of novel knowledge-based mean force potentials at the profile level has been presented, which uses the evolutionary information of profiles for developing more powerful statistical potentials. RESULTS: The frequency profiles are directly calculated from the multiple sequence alignments outputted by PSI-BLAST and converted into binary profiles with a probability threshold. As a result, the protein sequences are represented as sequences of binary profiles rather than sequences of amino acids. Similar to the knowledge-based potentials at the residue level, a class of novel potentials at the profile level is introduced. We develop four types of profile-level statistical potentials including distance-dependent, contact, Phi/Psi dihedral angle and accessible surface statistical potentials. These potentials are first evaluated by the fold assessment between the correct and incorrect models generated by comparative modeling from our own and other groups. They are then used to recognize the native structures from well-constructed decoy sets. Experimental results show that all the knowledge-base mean force potentials at the profile level outperform those at the residue level. Significant improvements are obtained for the distance-dependent and accessible surface potentials (5-6%). The contact and Phi/Psi dihedral angle potential only get a slight improvement (1-2%). Decoy set evaluation results show that the distance-dependent profile-level potentials even outperform other atom-level potentials. We also demonstrate that profile-level statistical potentials can improve the performance of threading. CONCLUSION: The knowledge-base mean force potentials at the profile level can provide better discriminatory ability than those at the residue level, so they will be useful for protein structure prediction and model refinement. Qiwen Dong, Xiaolong Wang 0001, Lei Lin 0001 |
BMC Bioinform. | 1 |