VLDB 2026 Research / reviewers in the wild / expert
Zhiqiang Ma 0003
dblp:62/3653-3
· DBLP profile ↗
26ranked-venue papers
1as first author
13since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 12 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Evolving pathway activation from cancer gene expression data using nature-inspired ensemble optimization
Xubin Wang 0001, Yunhe Wang 0002, Zhiqiang Ma 0003, Ka-Chun Wong, Xiangtao Li |
Expert Syst. Appl. | 3 |
| 2024 | Exhaustive Exploitation of Nature-Inspired Computation for Cancer Screening in an Ensemble MannerabstractAccurate screening of cancer types is crucial for effective cancer detection and precise treatment selection. However, the association between gene expression profiles and tumors is often limited to a small number of biomarker genes. While computational methods using nature-inspired algorithms have shown promise in selecting predictive genes, existing techniques are limited by inefficient search and poor generalization across diverse datasets. This study presents a framework termed Evolutionary Optimized Diverse Ensemble Learning (EODE) to improve ensemble learning for cancer classification from gene expression data. The EODE methodology combines an intelligent grey wolf optimization algorithm for selective feature space reduction, guided random injection modeling for ensemble diversity enhancement, and subset model optimization for synergistic classifier combinations. Extensive experiments were conducted across 35 gene expression benchmark datasets encompassing varied cancer types. Results demonstrated that EODE obtained significantly improved screening accuracy over individual and conventionally aggregated models. The integrated optimization of advanced feature selection, directed specialized modeling, and cooperative classifier ensembles helps address key challenges in current nature-inspired approaches. This provides an effective framework for robust and generalized ensemble learning with gene expression biomarkers. Xubin Wang 0001, Yunhe Wang 0002, Zhiqiang Ma 0003, Ka-Chun Wong, Xiangtao Li |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | DSHP: A Novel Sequence Based Deep Learning Prediction Model for HPV Integration SiteabstractHPV, a significant hazard to human health, is the primary cause of many cancers. The proteins E6 and E7 of HPV are known to damage oncogenes, but many mechanisms are still unknown. Research reveals that HPV can integrate its genome into host genes, and the integration mechanism strongly depends on the local genomic environment. Research on the integration mechanism can deepen the understanding of HPV and the development of vaccines, thus further affecting the cure of cancers and other related diseases. However, the research on HPV integration sites in silico experiments is in its infancy, and improving the model performance of HPV integration site predictors is challenging. In this work, we propose a novel deep learning model for HPV integration site prediction named DSHP. DSHP uses a variety of features of DNA sequences as input. In the 5-fold cross-validation, the ACC and AUC of DSHP are 0.914 and 0.934; in the 10-fold cross-validation, the ACC and AUC of DSHP are 0.933 and 0.941. The performance fully illustrates the effectiveness of DSHP. Moreover, our ablation experiments further explain the importance of features in the prediction process, and provide a reference for future prediction research. The data and code are available at: https://github.com/xtnenu/DSHP. Xian Tan, Yiping Sun, Shijie Fan, Yanhe Wang, Zhiqiang Ma 0003 |
BIBM | 7 |
| 2022 | HCRNet: high-throughput circRNA-binding event identification from CLIP-seq data using deep temporal convolutional networkabstractIdentifying genome-wide binding events between circular RNAs (circRNAs) and RNA-binding proteins (RBPs) can greatly facilitate our understanding of functional mechanisms within circRNAs. Thanks to the development of cross-linked immunoprecipitation sequencing technology, large amounts of genome-wide circRNA binding event data have accumulated, providing opportunities for designing high-performance computational models to discriminate RBP interaction sites and thus to interpret the biological significance of circRNAs. Unfortunately, there are still no computational models sufficiently flexible to accommodate circRNAs from different data scales and with various degrees of feature representation. Here, we present HCRNet, a novel end-to-end framework for identification of circRNA-RBP binding events. To capture the hierarchical relationships, the multi-source biological information is fused to represent circRNAs, including various natural language sequence features. Furthermore, a deep temporal convolutional network incorporating global expectation pooling was developed to exploit the latent nucleotide dependencies in an exhaustive manner. We benchmarked HCRNet on 37 circRNA datasets and 31 linear RNA datasets to demonstrate the effectiveness of our proposed method. To evaluate further the model's robustness, we performed HCRNet on a full-length dataset containing 740 circRNAs. Results indicate that HCRNet generally outperforms existing methods. In addition, motif analyses were conducted to exhibit the interpretability of HCRNet on circRNAs. All supporting source code and data can be downloaded from https://github.com/yangyn533/HCRNet and https://doi.org/10.6084/m9.figshare.16943722.v1. And the web server of HCRNet is publicly accessible at http://39.104.118.143:5001/. Zilong Hou, Hongli Ma, Zhiqiang Ma 0003, Ka-Chun Wong, Xiangtao Li |
Briefings Bioinform. | 6 |
| 2022 | GMHCC: high-throughput analysis of biomolecular data using graph-based multiple hierarchical consensus clusteringabstractMOTIVATION: Thanks to the development of high-throughput sequencing technologies, massive amounts of various biomolecular data have been accumulated to revolutionize the study of genomics and molecular biology. One of the main challenges in analyzing this biomolecular data is to cluster their subtypes into subpopulations to facilitate subsequent downstream analysis. Recently, many clustering methods have been developed to address the biomolecular data. However, the computational methods often suffer from many limitations such as high dimensionality, data heterogeneity and noise. RESULTS: In our study, we develop a novel Graph-based Multiple Hierarchical Consensus Clustering (GMHCC) method with an unsupervised graph-based feature ranking (FR) and a graph-based linking method to explore the multiple hierarchical information of the underlying partitions of the consensus clustering for multiple types of biomolecular data. Indeed, we first propose to use a graph-based unsupervised FR model to measure each feature by building a graph over pairwise features and then providing each feature with a rank. Subsequently, to maintain the diversity and robustness of basic partitions (BPs), we propose multiple diverse feature subsets to generate several BPs and then explore the hierarchical structures of the multiple BPs by refining the global consensus function. Finally, we develop a new graph-based linking method, which explicitly considers the relationships between clusters to generate the final partition. Experiments on multiple types of biomolecular data including 35 cancer gene expression datasets and eight single-cell RNA-seq datasets validate the effectiveness of our method over several state-of-the-art consensus clustering approaches. Furthermore, differential gene analysis, gene ontology enrichment analysis and KEGG pathway analysis are conducted, providing novel insights into cell developmental lineages and characterization mechanisms. AVAILABILITY AND IMPLEMENTATION: The source code is available at GitHub: https://github.com/yifuLu/GMHCC. The software and the supporting data can be downloaded from: https://figshare.com/articles/software/GMHCC/17111291. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yifu Lu, Zhuohan Yu, Yunhe Wang 0006, Zhiqiang Ma 0003, Ka-Chun Wong, Xiangtao Li |
Bioinform. | 4 |
| 2022 | EDCNN: identification of genome-wide RNA-binding proteins using evolutionary deep convolutional neural networkabstractMOTIVATION: RNA-binding proteins (RBPs) are a group of proteins associated with RNA regulation and metabolism, and play an essential role in mediating the maturation, transport, localization and translation of RNA. Recently, Genome-wide RNA-binding event detection methods have been developed to predict RBPs. Unfortunately, the existing computational methods usually suffer some limitations, such as high-dimensionality, data sparsity and low model performance. RESULTS: Deep convolution neural network has a useful advantage for solving high-dimensional and sparse data. To improve further the performance of deep convolution neural network, we propose evolutionary deep convolutional neural network (EDCNN) to identify protein-RNA interactions by synergizing evolutionary optimization with gradient descent to enhance deep conventional neural network. In particular, EDCNN combines evolutionary algorithms and different gradient descent models in a complementary algorithm, where the gradient descent and evolution steps can alternately optimize the RNA-binding event search. To validate the performance of EDCNN, an experiment is conducted on two large-scale CLIP-seq datasets, and results reveal that EDCNN provides superior performance to other state-of-the-art methods. Furthermore, time complexity analysis, parameter analysis and motif analysis are conducted to demonstrate the effectiveness of our proposed algorithm from several perspectives. AVAILABILITY AND IMPLEMENTATION: The EDCNN algorithm is available at GitHub: https://github.com/yaweiwang1232/EDCNN. Both the software and the supporting data can be downloaded from: https://figshare.com/articles/software/EDCNN/16803217. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhiqiang Ma 0003, Ka-Chun Wong, Xiangtao Li |
Bioinform. | 3 |
| 2022 | SSKM_Succ: A Novel Succinylation Sites Prediction Method Incorporating K-Means Clustering With a New Semi-Supervised Learning AlgorithmabstractProtein succinylation is a type of post-translational modification (PTM) that occurs on lysine sites and plays a key role in protein conformation regulation and cellular function control. When training in computational method, it is difficult to designate negative samples because of the uncertainty of non-succinylation lysine sites, and if not handled properly, it may affect the performance of computational models dramatically. Therefore, we propose a new semi-supervised learning method to identify reliable non-succinylation lysine sites as negative samples. This method, named SSKM_Succ, also employs K-means clustering to divide data into 5 clusters. Besides, information of proximal PTMs and three kinds of sequence features (grey pseudo amino acid composition, K-space and position-special amino acid propensity) are utilized to formulate protein. Then, we perform a two-step feature selection to remove redundant features and construct the optimization model for each cluster. Finally, support vector machine is applied to construct a prediction model for each cluster. Promising results are obtained by this method with an accuracy of 80.18 percent for succinylation sites on the independent testing dataset. Meanwhile, we compare the result with other existing tools, and it shows that our method is promising for predicting succinylation sites. Through analysis, we further verify that succinylated protein has potential effects on amino acid degradation and fatty acid metabolism, and speculate that protein succinylation may be closely related to neurodegenerative diseases. The code of SSKM_Succ is available on the web https://github.com/yangyq505/SSKM_Succ.git. Zhiqiang Ma 0003, Xiaowei Zhao 0004, Minghao Yin |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | A Novel Method for Identification of Glutarylation Sites Combining Borderline-SMOTE With Tomek Links Technique in Imbalanced DataabstractGlutarylation is a type of post-translational modification that occurs on lysine residues. It plays an irreplaceable role in various cellular functions. Therefore, identification of glutarylation sites is significant for understanding the molecular mechanism of glutarylation. In this study, we proposed a method named DEXGB_Glu to identify lysine glutarylation sites using XGBoost as classifier which was optimized by differential evolution algorithm. Aiming at the imbalance between positive samples and negative samples, Borderline-SMOTE method was employed to synthesize positive samples, increasing their amount equal to negative samples. Then, Tomek links technique was applied to filter out noise data. Analysis of this method and its results showed that differential evolution algorithm obviously improved the performance and the combination of Borderline-SMOTE and Tomek links effectively solved the imbalance between positive samples and negative samples. Finally, the performance of this method was much better than other methods in prediction of glutarylation sites. The data and code are available on https://github.com/ningq669/DEXGB_Glu. Xiaowei Zhao 0004, Zhiqiang Ma 0003 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | An Improved Topology Prediction of Alpha-Helical Transmembrane Protein Based on Deep Multi-Scale Convolutional Neural NetworkabstractAlpha-helical proteins ( αTMPs) are essential in various biological processes. Despite their tertiary structures are crucial for revealing complex functions, experimental structure determination remains challenging and costly. In the past decades, various sequence-based topology prediction methods have been developed to bridge the gap between the sequences and structures by characterizing the structural features, but significant improvements are still required. Deep learning brings a great opportunity for its powerful representation learning capability from limited original data. In this work, we improved our αTMP topology prediction method DMCTOP using deep learning, which composed of two deep convolutional blocks to simultaneously extract local and global contextual features. Consequently, the inputs were simplified to reflect the original features of the sequence, including a protein sequence feature and an evolutionary conservation feature. DMCTOP can efficiently and accurately identify all topological types and the N-terminal orientation for an αTMP sequence. To validate the effectiveness of our method, we benchmarked DMCTOP against 13 peer methods according to the whole sequence, the transmembrane segment and the traditional criterion in testing experiments. All the results reveal that our method achieved the highest prediction accuracy and outperformed all the previous methods. The method is available at https://icdtools.nenu.edu.cn/dmctop. Jiawen Yu, Zhe Liu 0030, Han Wang 0028, Zhiqiang Ma 0003, Dong Xu 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2021 | A Novel Approach for LncRNA Function Prediction Based on Deep LearningabstractLncRNAs have been discovered to play important roles in biology processes and progression of some severe genetic diseases, such as cancer. However, a number of lncRNAs is lack of function annotations. To understand the functions of lncRNAs is a key step for biology and medical research. Computational method is more suitable for large-scale experiment than biological experimental method due to time and expense, and it can provide reference and basis for biological experiment. Nonetheless, developing computational method to predict function of a newly detected lncRNA is still a challenging task. Some lncRNAs work according to binding with proteins, so calculating interaction preferences of lncRNA-protein may provide novel breakthrough to develop new computational approaches. In this work, we propose a novel lncRNA function prediction method to link function of lncRNAs and RNA binding protein. The method consists of two steps: firstly, we improve the CNN models of DLPRB, a high-performance deep neural network for protein-RNA binding preferences prediction, secondly, we apply a voting algorithm to predict lncRNA function by integrating binding preference predicting models. As we know, it is the first time that using RNA-protein preferences to predict function of lncRNAs. We test the predicting approach with 25 lncRNAs from LncBook, and the results confirm our approach makes sense. The approach only uses sequence-based features, so it can be a novel choice for newly detected lncRNA function prediction. We think this work can provide new insights for lncRNA research. Xian Tan, Minghang Zou, Zhiqiang Ma 0003 |
BIBM | 6 |
| 2021 | iCircRBP-DHN: identification of circRNA-RBP interaction sites using deep hierarchical networkabstractCircular RNAs (circRNAs) are widely expressed in eukaryotes. The genome-wide interactions between circRNAs and RNA-binding proteins (RBPs) can be probed from cross-linking immunoprecipitation with sequencing data. Therefore, computational methods have been developed for identifying RBP binding sites on circRNAs. Unfortunately, those computational methods often suffer from the low discriminative power of feature representations, numerical instability and poor scalability. To address those limitations, we propose a novel computational method called iCircRBP-DHN using deep hierarchical network for discriminating circRNA-RBP binding sites. The network architecture can be regarded as a deep multi-scale residual network followed by bidirectional gated recurrent units (BiGRUs) with the self-attention mechanism, which can simultaneously extract local and global contextual information. Meanwhile, we propose novel encoding schemes by integrating CircRNA2Vec and the K-tuple nucleotide frequency pattern to represent different degrees of nucleotide dependencies. To validate the effectiveness of our proposed iCircRBP-DHN, we compared its performance with other computational methods on 37 circRNAs datasets and 31 linear RNAs datasets, respectively. The experimental results reveal that iCircRBP-DHN can achieve superior performance over those state-of-the-art algorithms. Moreover, we perform motif analysis on circRNAs bound by those different RBPs, demonstrating that our proposed CircRNA2Vec encoding scheme can be promising. The iCircRBP-DHN method is made available at https://github.com/houzl3416/iCircRBP-DHN. Zilong Hou, Zhiqiang Ma 0003, Xiangtao Li, Ka-Chun Wong |
Briefings Bioinform. | 3 |
| 2021 | Identification of haploinsufficient genes from epigenomic data using deep forestabstractHaploinsufficiency, wherein a single allele is not enough to maintain normal functions, can lead to many diseases including cancers and neurodevelopmental disorders. Recently, computational methods for identifying haploinsufficiency have been developed. However, most of those computational methods suffer from study bias, experimental noise and instability, resulting in unsatisfactory identification of haploinsufficient genes. To address those challenges, we propose a deep forest model, called HaForest, to identify haploinsufficient genes. The multiscale scanning is proposed to extract local contextual representations from input features under Linear Discriminant Analysis. After that, the cascade forest structure is applied to obtain the concatenated features directly by integrating decision-tree-based forests. Meanwhile, to exploit the complex dependency structure among haploinsufficient genes, the LightGBM library is embedded into HaForest to reveal the highly expressive features. To validate the effectiveness of our method, we compared it to several computational methods and four deep learning algorithms on five epigenomic data sets. The results reveal that HaForest achieves superior performance over the other algorithms, demonstrating its unique and complementary performance in identifying haploinsufficient genes. The standalone tool is available at https://github.com/yangyn533/HaForest. Shaochuan Li, Yunhe Wang 0002, Zhiqiang Ma 0003, Ka-Chun Wong, Xiangtao Li |
Briefings Bioinform. | 4 |
| 2021 | Evolving Multiobjective Cancer Subtype Diagnosis From Cancer Gene Expression DataabstractDetection and diagnosis of cancer are especially essential for early prevention and effective treatments. Many studies have been proposed to tackle the subtype diagnosis problems with those data, which often suffer from low diagnostic ability and bad generalization. This article studies a multiobjective PSO-based hybrid algorithm (MOPSOHA) to optimize four objectives including the number of features, the accuracy, and two entropy-based measures: the relevance and the redundancy simultaneously, diagnosing the cancer data with high classification power and robustness. First, we propose a novel binary encoding strategy to choose informative gene subsets to optimize those objective functions. Second, a mutation operator is designed to enhance the exploration capability of the swarm. Finally, a local search method based on the "best/1" mutation operator of differential evolutionary algorithm (DE) is employed to exploit the neighborhood area with sparse high-quality solutions since the base vector always approaches to some good promising areas. In order to demonstrate the effectiveness of MOPSOHA, it is tested on 41 cancer datasets including thirty-five cancer gene expression datasets and six independent disease datasets. Compared MOPSOHA with other state-of-the-art algorithms, the performance of MOPSOHA is superior to other algorithms in most of the benchmark datasets. Yunhe Wang 0002, Zhiqiang Ma 0003, Ka-Chun Wong, Xiangtao Li |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2020 | SeqTMPPI: Sequence-Based Transmembrane Protein Interaction PredictionabstractTransmembrane proteins (TMPs) play important roles in diverse cellular processes, they are the most common drug targets. The interaction between TMPs and nontransmembrane proteins (nonTMPs) is an important part to form the pathway crossing biomembranes, which is directly related to TMPs associated signal transduction, substance transport, and drug metabolism. However, it is difficult for biologists to identify interactions between TMPs and nonTMPs by wet-lab experiments since TMPs are embedded in the phospholipid bilayer while most nonTMPs are soluble proteins. Predicting protein-protein interactions (PPIs) has always been a hot topic in bioinformatics. However, TMP related PPIs occupy a small proportion of those researches, and those methods designed for soluble protein PPI may have limitations to predict the interaction of TMPs and non-TMPs due to the difference between the aqueous and lipid senvironments. There still lack of deep-learning-based predictors using primary sequence information on TMP-nonTMP interactions. In this work, we constructed a benchmark dataset for TMP-nonTMP interactions, adopted the one-hot vector to encode protein sequence pairs and built a Convolutional Neural Network (CNN) based model, SeqTMPPI, predicting the TMP-nonTMP interactions. The experimental results indicated that our method achieved a good performance on an independent testing where the Matthews Correlation Coefficient (MCC) achieved at 0.7066. The predicted interactions were further analyzed in the scope of distribution of the protein family and species. Materials related are available in https://github.com/JulseJiang/SeqTMPPI. Han Wang 0028, Jiuhong Jiang, Qiufen Chen, Chang Lu 0015, Zhiqiang Ma 0003 |
BIBM | 6 |
| 2020 | Corrigendum to: Comprehensive review and empirical analysis of hallmarks of DNA-, RNA- and protein-binding residues in protein chainsabstractBriefings in Bioinformatics, 2017. https://doi.org/10.1093/bib/bbx168 In the original version of this article, reference 60 incorrectly listed Ben-Tal N as the third author of the paper ‘Local geometry and evolutionary conservation of protein surfaces reveal the multiple recognition patches in protein-protein interactions’. This reference has now been corrected to the below: Laine E, Carbone A. Local geometry and evolutionary conservation of protein surfaces reveal the multiple recognition patches in protein-protein interactions. PLoS Comput Biol 2015;11(12):e1004580. The authors would like to apologize for this error. Jian Zhang 0020, Zhiqiang Ma 0003, Lukasz A. Kurgan |
Briefings Bioinform. | 2 |
| 2020 | Nature-inspired multiobjective patient stratification from cancer gene expression data
Yunhe Wang 0002, Zhiqiang Ma 0003, Ka-Chun Wong, Xiangtao Li |
Inf. Sci. | 2 |
| 2020 | Cancer molecular subtype classification from hypervolume-based discrete evolutionary optimization
Yunhe Wang 0002, Shaochuan Li, Zhiqiang Ma 0003, Xiangtao Li |
Neural Comput. Appl. | 4 |
| 2019 | Comprehensive review and empirical analysis of hallmarks of DNA-, RNA- and protein-binding residues in protein chainsabstractProteins interact with a variety of molecules including proteins and nucleic acids. We review a comprehensive collection of over 50 studies that analyze and/or predict these interactions. While majority of these studies address either solely protein-DNA or protein-RNA binding, only a few have a wider scope that covers both protein-protein and protein-nucleic acid binding. Our analysis reveals that binding residues are typically characterized with three hallmarks: relative solvent accessibility (RSA), evolutionary conservation and propensity of amino acids (AAs) for binding. Motivated by drawbacks of the prior studies, we perform a large-scale analysis to quantify and contrast the three hallmarks for residues that bind DNA-, RNA-, protein- and (for the first time) multi-ligand-binding residues that interact with DNA and proteins, and with RNA and proteins. Results generated on a well-annotated data set of over 23 000 proteins show that conservation of binding residues is higher for nucleic acid- than protein-binding residues. Multi-ligand-binding residues are more conserved and have higher RSA than single-ligand-binding residues. We empirically show that each hallmark discriminates between binding and nonbinding residues, even predicted RSA, and that combining them improves discriminatory power for each of the five types of interactions. Linear scoring functions that combine these hallmarks offer good predictive performance of residue-level propensity for binding and provide intuitive interpretation of predictions. Better understanding of these residue-level interactions will facilitate development of methods that accurately predict binding in the exponentially growing databases of protein sequences. Jian Zhang 0020, Zhiqiang Ma 0003, Lukasz A. Kurgan |
Briefings Bioinform. | 2 |
| 2019 | Analysis and prediction of human acetylation using a cascade classifier based on support vector machineabstractBACKGROUND: Acetylation on lysine is a widespread post-translational modification which is reversible and plays a crucial role in some biological activities. To better understand the mechanism, it is necessary to identify acetylation sites in proteins accurately. Computational methods are popular because they are more convenient and faster than experimental methods. In this study, we proposed a new computational method to predict acetylation sites in human by combining sequence features and structural features including physicochemical property (PCP), position specific score matrix (PSSM), auto covariation (AC), residue composition (RC), secondary structure (SS) and accessible surface area (ASA), which can well characterize the information of acetylated lysine sites. Besides, a two-step feature selection was applied, which combined mRMR and IFS. It finally trained a cascade classifier based on SVM, which successfully solved the imbalance between positive samples and negative samples and covered all negative sample information. RESULTS: The performance of this method is measured with a specificity of 72.19% and a sensibility of 76.71% on independent dataset which shows that a cascade SVM classifier outperforms single SVM classifier. CONCLUSIONS: In addition to the analysis of experimental results, we also made a systematic and comprehensive analysis of the acetylation data. Jinchao Ji, Zhiqiang Ma 0003, Xiaowei Zhao 0004 |
BMC Bioinform. | 4 |
| 2018 | Detecting Succinylation sites from protein sequences using ensemble support vector machineabstractBACKGROUND: Lysine succinylation is a new kind of post-translational modification which plays a key role in protein conformation regulation and cellular function control. To understand the mechanism of succinylation profoundly, it is necessary to identify succinylation sites in proteins accurately. However, traditional methods, experimental approaches, are labor-intensive and time-consuming. Computational prediction methods have been proposed recent years, and they are popular because of their convenience and high speed. In this study, we developed a new method to predict succinylation sites in protein combining multiple features, including amino acid composition, binary encoding, physicochemical property and grey pseudo amino acid composition, with a feature selection scheme (information gain). And then, it was trained using SVM (Support Vector Machine) and an ensemble learning algorithm. RESULTS: The performance of this method was measured with an accuracy of 89.14% and a MCC (Matthew Correlation Coefficient) of 0.79 using 10-fold cross validation on training dataset and an accuracy of 84.5% and a MCC of 0.2 on independent dataset. CONCLUSIONS: The conclusions made from this study can help to understand more of the succinylation mechanism. These results suggest that our method was very promising for predicting succinylation sites. The source code and data of this paper are freely available at https://github.com/ningq669/PSuccE . Xiaosa Zhao, Lingling Bao, Zhiqiang Ma 0003, Xiaowei Zhao 0004 |
BMC Bioinform. | 4 |
| 2018 | HEMEsPred: Structure-Based Ligand-Specific Heme Binding Residues Prediction by Using Fast-Adaptive Ensemble Learning SchemeabstractHeme is an essential biomolecule that widely exists in numerous extant organisms. Accurately identifying heme binding residues (HEMEs) is of great importance in disease progression and drug development. In this study, a novel predictor named HEMEsPred was proposed for predicting HEMEs. First, several sequence- and structure-based features, including amino acid composition, motifs, surface preferences, and secondary structure, were collected to construct feature matrices. Second, a novel fast-adaptive ensemble learning scheme was designed to overcome the serious class-imbalance problem as well as to enhance the prediction performance. Third, we further developed ligand-specific models considering that different heme ligands varied significantly in their roles, sizes, and distributions. Statistical test proved the effectiveness of ligand-specific models. Experimental results on benchmark datasets demonstrated good robustness of our proposed method. Furthermore, our method also showed good generalization capability and outperformed many state-of-art predictors on two independent testing datasets. HEMEsPred web server was available at http://www.inforstation.com/HEMEsPred/ for free academic use. Jian Zhang 0020, Haiting Chai, Guifu Yang, Zhiqiang Ma 0003 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2017 | Prediction of bioluminescent proteins by using sequence-derived features and lineage-specific schemeabstractBACKGROUND: Bioluminescent proteins (BLPs) widely exist in many living organisms. As BLPs are featured by the capability of emitting lights, they can be served as biomarkers and easily detected in biomedical research, such as gene expression analysis and signal transduction pathways. Therefore, accurate identification of BLPs is important for disease diagnosis and biomedical engineering. In this paper, we propose a novel accurate sequence-based method named PredBLP (Prediction of BioLuminescent Proteins) to predict BLPs. RESULTS: We collect a series of sequence-derived features, which have been proved to be involved in the structure and function of BLPs. These features include amino acid composition, dipeptide composition, sequence motifs and physicochemical properties. We further prove that the combination of four types of features outperforms any other combinations or individual features. To remove potential irrelevant or redundant features, we also introduce Fisher Markov Selector together with Sequential Backward Selection strategy to select the optimal feature subsets. Additionally, we design a lineage-specific scheme, which is proved to be more effective than traditional universal approaches. CONCLUSION: Experiment on benchmark datasets proves the robustness of PredBLP. We demonstrate that lineage-specific models significantly outperform universal ones. We also test the generalization capability of PredBLP based on independent testing datasets as well as newly deposited BLPs in UniProt. PredBLP is proved to be able to exceed many state-of-art methods. A web server named PredBLP, which implements the proposed method, is free available for academic use. Jian Zhang 0020, Haiting Chai, Guifu Yang, Zhiqiang Ma 0003 |
BMC Bioinform. | 4 |
| 2016 | Identification of DNA-binding proteins using multi-features fusion and binary firefly optimization algorithmabstractBACKGROUND: DNA-binding proteins (DBPs) play fundamental roles in many biological processes. Therefore, the developing of effective computational tools for identifying DBPs is becoming highly desirable. RESULTS: In this study, we proposed an accurate method for the prediction of DBPs. Firstly, we focused on the challenge of improving DBP prediction accuracy with information solely from the sequence. Secondly, we used multiple informative features to encode the protein. These features included evolutionary conservation profile, secondary structure motifs, and physicochemical properties. Thirdly, we introduced a novel improved Binary Firefly Algorithm (BFA) to remove redundant or noisy features as well as select optimal parameters for the classifier. The experimental results of our predictor on two benchmark datasets outperformed many state-of-the-art predictors, which revealed the effectiveness of our method. The promising prediction performance on a new-compiled independent testing dataset from PDB and a large-scale dataset from UniProt proved the good generalization ability of our method. In addition, the BFA forged in this research would be of great potential in practical applications in optimization fields, especially in feature selection problems. CONCLUSIONS: A highly accurate method was proposed for the identification of DBPs. A user-friendly web-server named iDbP (identification of DNA-binding Proteins) was constructed and provided for academic use. Jian Zhang 0020, Haiting Chai, Zhiqiang Ma 0003, Guifu Yang |
BMC Bioinform. | 4 |
| 2011 | Content-Based Biometric Image Hiding ApproachabstractRecently, the use of information hiding techniques to protect biometric data has been an active topic. This paper proposes a novel image hiding approach based on correlation analysis to protect network-based transmitted biometric image for identification. Firstly, the correlation between the biometric image and the cover image is analyzed using principal component analysis (PCA) and genetic algorithm (GA). The purpose of correlation analysis is to enable the cover image to represent the secret image in content as much as possible, not just as a carrier of hidden information. Then, the unrepresented part of the biometric image, as the secret image, is encrypted and hidden into the middle-significant-bit plane (MSB) of the cover image redundantly. Extensive experimental results demonstrate that the proposed hiding approach not only gains good imperceptibility, but also resists some common attacks validated by the biometric identification accuracy. Miao Qi, Jun Kong 0004, Yinghua Lu, Ning Du, Zhiqiang Ma 0003 |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2011 | A structure-preserved local matching approach for face recognition
Jianzhong Wang 0003, Zhiqiang Ma 0003, Baoxue Zhang, Miao Qi, Jun Kong 0004 |
Pattern Recognit. Lett. | 2 |
| 2007 | A Novel Off-Line Signature Verification Based on Adaptive Multi-resolution Wavelet Zero-Crossing and One-Class-One-Network
Zhiqiang Ma 0003, Xiaoyun Zeng, Chunguang Zhou |
ISNN (3) | 1 |