EDBT 2026 Demo / reviewers in the wild / expert
Peng Chen 0001
dblp:27/7017-1
· DBLP profile ↗
75ranked-venue papers
13as first author
20since 2021 · last 2026
0000-0002-5810-8159ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 54 · 8 first-author · 16 since 2021Artificial intelligence and machine learning · 16 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A vision-language network for stored-grain pest counting
Rui Li 0027, Chengjun Xie, Peng Chen 0001, Jie Zhang 0033, Jianming Du, Runsheng Qi |
Expert Syst. Appl. | 4 |
| 2025 | Local global information aggregation graph convolution for skeleton-based action recognition
Shichong Xie, Shengze Li, Peng Chen 0001, Bing Wang 0004, Jun Zhang 0011 |
Neurocomputing | 3 |
| 2024 | Medical Tumor Image Classification Based on Few-Shot LearningabstractAs a high mortality disease, cancer seriously affects people's life and well-being. Reliance on pathologists to assess disease progression from pathological images is inaccurate and burdensome. Computer aided diagnosis (CAD) system can effectively assist diagnosis and make more credible decisions. However, a large number of labeled medical images that contribute to improve the accuracy of machine learning algorithm, especially for deep learning in CAD, are difficult to collect. Therefore, in this work, an improved few-shot learning method is proposed for medical image recognition. In addition, to make full use of the limited feature information in one or more samples, a feature fusion strategy is involved in our model. On the dataset of BreakHis and skin lesions, the experimental results show that our model achieved the classification accuracy of 91.22% and 71.20% respectively when only 10 labeled samples are given, which is superior to other state-of-the-art methods. Kun Lu 0007, Jun Zhang 0011, Peng Chen 0001, Ke Yan 0001, Bing Wang 0004 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2023 | Collaborative Encoder for Accurate Inversion of Real Face Image
YaTe Liu, Chun-Hou Zheng 0001, Jun Zhang 0011, Bing Wang 0004, Peng Chen 0001 |
ICIC (2) | 5 |
| 2023 | Multiple Classification Network of Concrete Defects Based on Improved EfficientNetV2
Jiawei Ni, Kun Lu 0007, Jun Zhang 0011, Peng Chen 0001, Lejun Pan, Chenlin Zhu, Bing Wang 0004 |
ICIC (2) | 4 |
| 2023 | Improved Deep Learning-Based Efficientpose Algorithm for Egocentric Marker-Less Tool and Hand Pose Estimation in Manual Assembly
Zihan Niu, Jun Zhang 0011, Bing Wang 0004, Peng Chen 0001 |
ICIC (5) | 5 |
| 2023 | Efficient and Precise Detection of Surface Defects on PCBs: A YOLO Based Approach
Lejun Pan, Kun Lu 0007, Jun Zhang 0011, Peng Chen 0001, Jiawei Ni, Chenlin Zhu, Bing Wang 0004 |
ICIC (2) | 5 |
| 2023 | Improved YOLOv5s Method for Nut Detection on Ultra High Voltage Power Towers
Jun Zhang 0011, Bing Wang 0004, Peng Chen 0001 |
ICIC (5) | 5 |
| 2023 | Corneal Ulcer Automatic Classification Network Based on Improved Mobile ViT
Chenlin Zhu, Kun Lu 0007, Jun Zhang 0011, Peng Chen 0001, Lejun Pan, Jiawei Ni, Bing Wang 0004 |
ICIC (2) | 5 |
| 2023 | An attention-based feature pyramid network for single-stage small object detection
Lin Jiao, Chenrui Kang, Shifeng Dong, Peng Chen 0001, Gaoqiang Li, Rujing Wang |
Multim. Tools Appl. | 4 |
| 2023 | SGNet: Sequence-Based Convolution and Ligand Graph Network for Protein Binding Affinity PredictionabstractProtein-ligand binding can play an important role in many fields. It is of great importance to accurately predict the binding affinity between molecules by computational methods. Most computational binding affinity methods require molecular structures. However, there are still a large number of protein molecules with known amino acid sequences whose structures have not yet been solved. To address this issue, this paper proposes a sequence-based convolution and ligand graph network, called SGNet, to fuse the molecular graph information and the amino acid sequence information. This method integrates Conjoint Triad (CT) encoding of amino acid sequence and one-dimensional convolutional neural network module to extract protein molecules, develops graph attention network to extract molecular features of ligand, and then fuses the two feature sets to predict the binding affinity between molecules from the fully connected layer. As a result, SGNet achieves good prediction performance on both KIKD andIC50data sets, with prediction error RMSEs of 1.287 and 1.58, and correlation Pearson Rs of 0.687 and 0.592, respectively. Comparative experimental results under the same conditions showed that SGNet outperformed Kdeep and GraphDTA in predicting binding affinities between protein-ligand molecules. Peng Chen 0001, Huimin Shen, Youzhi Zhang 0004, Bing Wang 0004, Pengying Gu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2022 | COVID-19 Classification from Chest X-rays Based on Attention and Knowledge Distillation
Jiaxing Lv, Fazhan Zhu, Kun Lu 0007, Jun Zhang 0011, Peng Chen 0001, Yuan Zhao 0012 |
ICIC (1) | 6 |
| 2022 | A Sub-network Aggregation Neural Network for Non-invasive Blood Pressure Prediction
Xinghui Zhang, Chun-Hou Zheng 0001, Peng Chen 0001, Jun Zhang 0011, Bing Wang 0004 |
ICIC (1) | 3 |
| 2022 | A 3D Medical Image Segmentation Framework Fusing Convolution and Transformer Features
Fazhan Zhu, Jiaxing Lv, Kun Lu 0007, Hongshou Cong, Jun Zhang 0011, Peng Chen 0001, Yuan Zhao 0012 |
ICIC (1) | 7 |
| 2022 | Protein-Protein Interaction Sites Prediction Based on an Under-Sampling Strategy and Random Forest AlgorithmabstractThe computational methods of protein-protein interaction sites prediction can effectively avoid the shortcomings of high cost and time in traditional experimental approaches. However, the serious class imbalance between interface and non-interface residues on the protein sequences limits the prediction performance of these methods. This work therefore proposed a new strategy, NearMiss-based under-sampling for unbalancing datasets and Random Forest classification (NM-RF), to predict protein interaction sites. Herein, the residues on protein sequences were represented by the PSSM-derived features, hydropathy index (HI) and relative solvent accessibility (RSA). In order to resolve the class imbalance problem, an under-sampling method based on NearMiss algorithm is adopted to remove some non-interface residues, and then the random forest algorithm is used to perform binary classification on the balanced feature datasets. Experiments show that the accuracy of NM-RF model reaches 87.6% and 84.3% on Dtestset72 and PDBtestset164 respectively, which demonstrate the effectiveness of the proposed NM-RF method in differentiating the interface or non-interface residues. Minjie Li, Kun Lu 0007, Jun Zhang 0011, Yuming Zhou, Zhaoquan Chen, Dan Li 0025, Shicheng Zheng, Peng Chen 0001, Bing Wang 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 10 |
| 2022 | Transformer Model for Functional Near-Infrared Spectroscopy ClassificationabstractFunctional near-infrared spectroscopy (fNIRS) is a promising neuroimaging technology. The fNIRS classification problem has always been the focus of the brain-computer interface (BCI). Inspired by the success of Transformer based on self-attention mechanism in the fields of natural language processing and computer vision, we propose an fNIRS classification network based on Transformer, named fNIRS-T. We explore the spatial-level and channel-level representation of fNIRS signals to improve data utilization and network representation capacity. Besides, a preprocessing module, which consists of one-dimensional average pooling and layer normalization, is designed to replace filtering and baseline correction of data preprocessing. It makes fNIRS-T an end-to-end network, called fNIRS-PreT. Compared with traditional machine learning classifiers, convolutional neural network (CNN), and long short-term memory (LSTM), the proposed models obtain the best accuracy on three open-access datasets. Specifically, in the most extensive ternary classification task (30 subjects) that includes three types of overt movements, fNIRS-T, CNN, and LSTM obtain 75.49%, 72.89%, and 61.94% on test sets, respectively. Compared to traditional classifiers, fNIRS-T is at least 27.41% higher than statistical features and 6.79% higher than well-designed features. In the individual subject experiment of the ternary classification task, fNIRS-T achieves an average subject accuracy of 78.22% and surpasses CNN and LSTM by a large margin of +4.75% and +11.33%. fNIRS-PreT using raw data also achieves competitive performance to fNIRS-T. Therefore, the proposed models improve the performance of fNIRS-based BCI significantly. Zenghui Wang 0009, Jun Zhang 0011, Xiaochu Zhang, Peng Chen 0001, Bing Wang 0004 |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | Recognition and counting of wheat mites in wheat fields by a three-step deep learning method
Peng Chen 0001, Weilu Li, Sijie Yao, Chun Ma, Jun Zhang 0011, Bing Wang 0004, Chun-Hou Zheng 0001, Chengjun Xie |
Neurocomputing | 1 |
| 2021 | A Convolutional Neural Network System to Discriminate Drug-Target InteractionsabstractBiological targets are most commonly proteins such as enzymes, ion channels, and receptors. They are anything within a living organism to bind with some other entities (like an endogenous ligand or a drug), resulting in change in their behaviors or functions. Exploring potential drug-target interactions (DTIs) are crucial for drug discovery and effective drug development. Computational methods were widely applied in drug-target interactions, since experimental methods are extremely time-consuming and resource-intensive. In this paper, we proposed a novel deep learning-based prediction system, with a new negative instance generation, to identify DTIs. As a result, our method achieved an accuracy of 0.9800 on our created dataset. Another dataset derived from DrugBank was used to further assess the generalization of the model, which yielded a good performance with accuracy of 0.8814 and AUC value of 0.9527 on the dataset. The outcome of our experimental results indicated that the proposed method, involving the credible negative generation, can be employed to discriminate the interactions between drugs and targets. Website: http://www.dlearningapp.com/web/DrugCNN.htm. DeNan Xia, Benyue Su, Peng Chen 0001, Bing Wang 0004, Jinyan Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2021 | Imbalance Data Processing Strategy for Protein Interaction Sites PredictionabstractProtein-protein interactions play essential roles in various biological progresses. Identifying protein interaction sites can facilitate researchers to understand life activities and therefore will be helpful for drug design. However, the number of experimental determined protein interaction sites is far less than that of protein sites in protein-protein interaction or protein complexes. Therefore, the negative and positive samples are usually imbalanced, which is common but bring result bias on the prediction of protein interaction sites by computational approaches. In this work, we presented three imbalance data processing strategies to reconstruct the original dataset, and then extracted protein features from the evolutionary conservation of amino acids to build a predictor for identification of protein interaction sites. On a dataset with 10,430 surface residues but only 2,299 interface residues, the imbalance dataset processing strategies can obviously reduce the prediction bias, and therefore improve the prediction performance of protein interaction sites. The experimental results show that our prediction models can achieve a better prediction performance, such as a prediction accuracy of 0.758, or a high F-measure of 0.737, which demonstrated the effectiveness of our method. Bing Wang 0004, Changqing Mei, Yuming Zhou, Mu-Tian Cheng, Chun-Hou Zheng 0001, Lei Wang 0069, Jun Zhang 0011, Peng Chen 0001, Yan Xiong 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 9 |
| 2021 | Potential Pathogenic Genes Prioritization Based on Protein Domain Interaction Network AnalysisabstractPathogenicity-related studies are of great importance in understanding the pathogenesis of complex diseases and improving the level of clinical medicine. This work proposed a bioinformatics scheme to analyze cancer-related gene mutations, and try to figure out potential genes associated with diseases from the protein domain-domain interaction network. Herein, five measures of the principle of centrality lethality had been adopted to implement potential correlation analysis, and prioritize the significance of genes. This method was further applied to KEGG pathway analysis by taking the malignant melanoma as an example. The experimental results show that 25 domains can be found, and 18 of them have high potential to be pathogenically important related to malignant melanoma. Finally, a web-based tool, named Human Cancer Related Domain Interaction Network Analyzer, is developed for potential pathogenic genes prioritization for 26 types of human cancers, and the analysis results can be visualized and downloaded online. Yuming Zhou, Mu-Tian Cheng, Chun-Hou Zheng 0001, Yan Xiong 0001, Peng Chen 0001, Zhiwei Ji, Bing Wang 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2020 | Discrete Haze Level Dehazing NetworkabstractIn contrast to traditional dehazing methods, deep learning based single image dehazing (SID) algorithms have achieved better performances by creating a mapping function from haze to haze-free images. Usually, the images taken from the natural scenes have different haze levels, but deep SID algorithms only process the hazy images as one group. It makes the deep SID algorithms difficult to deal with the image set with some images having specific haze density. In this paper, a Discrete Haze Level Dehazing network (DHL-Dehaze), a very effective method to dehaze multiple different haze level images, is proposed. The proposed approach considers a single image dehazing problem as a multi-domain image-to-image translation, instead of grouping all hazy images into the same domain. DHL-Dehaze provides computational derivation to describe the role of different haze levels for image translation. To verify the proposed approach, we synthesize two largescale datasets with multiple haze level images based on the NYU-Depth and DIML/CVL datasets. The experiments show that DHL-Dehaze can obtain excellent quantitative and qualitative dehazing results, especially when the haze concentration is high. Xiaofeng Cong, Jie Gui, Kai-Chao Miao, Jun Zhang 0011, Bing Wang 0004, Peng Chen 0001 |
ACM Multimedia | 6 |
| 2020 | Application of LSTM for short term fog forecasting based on meteorological elements
Kai-Chao Miao, Ting-Ting Han, Ye-Qing Yao, Peng Chen 0001, Bing Wang 0004, Jun Zhang 0011 |
Neurocomputing | 5 |
| 2020 | A Deep Learning-Based Chemical System for QSAR PredictionabstractResearch on quantitative structure-activity relationships (QSAR) provides an effective approach to determine new hits and promising lead compounds during drug discovery. In the past decades, various works have gained good performance for QSAR with the development of machine learning. The rise of deep learning, along with massive accessible chemical databases, made improvement on the QSAR performance. This article proposes a novel deep-learning-based method to implement QSAR prediction by the concatenation of end-to-end encoder-decoder model and convolutional neural network (CNN) architecture. The encoder-decoder model is mainly used to generate fixed-size latent features to represent chemical molecules; while these features are then input into CNN framework to train a robust and stable model and finally to predict active chemicals. Two models with different schemes are investigated to evaluate the validity of our proposed model on the same data sets. Experimental results showed that our proposed method outperforms other state-of-the-art methods in successful identification of chemical molecule whether it is active. Peng Chen 0001, Pengying Gu, Bing Wang 0004 |
IEEE J. Biomed. Health Informatics | 2 |
| 2019 | Identification of Apple Leaf Diseases Based on Convolutional Neural Network
Peng Chen 0001, Jun Zhang 0011, Bing Wang 0004 |
ICIC (1) | 2 |
| 2019 | Identification of Apple Tree Trunk Diseases Based on Improved Convolutional Neural Network with Fused Loss Functions
Jie Hang, Dexiang Zhang, Peng Chen 0001, Jun Zhang 0011, Bing Wang 0004 |
ICIC (1) | 3 |
| 2019 | Real-Time Pedestrian Detection in Monitoring Scene Based on Head Model
Panpan Lu, Kun Lu 0007, Jun Zhang 0011, Peng Chen 0001, Bing Wang 0004 |
ICIC (2) | 5 |
| 2019 | An Optimization Regression Model for Predicting Average Temperature of Core Dead Stock Column
Bing Dai, Hongming Long, Jun Zhang 0011, Peng Chen 0001, Bing Wang 0004 |
ICIC (3) | 6 |
| 2019 | Ranking Research Institutions Based on the Combination of Individual and Network Features
Jun Zhang 0011, Peng Chen 0001, Bing Wang 0004 |
ICIC (3) | 4 |
| 2019 | Urine Sediment Detection Based on Deep Learning
Xiao-Tao Xu, Jun Zhang 0011, Peng Chen 0001, Bing Wang 0004 |
ICIC (1) | 3 |
| 2019 | Predicting drug-target interactions from drug structure and protein sequence using novel convolutional neural networksabstractBACKGROUND: Accurate identification of potential interactions between drugs and protein targets is a critical step to accelerate drug discovery. Despite many relative experimental researches have been done in the past decades, detecting drug-target interactions (DTIs) remains to be extremely resource-intensive and time-consuming. Therefore, many computational approaches have been developed for predicting drug-target associations on a large scale. RESULTS: In this paper, we proposed an deep learning-based method to predict DTIs only using the information of drug structures and protein sequences. The final results showed that our method can achieve good performance with the accuracies up to 92.0%, 90.0%, 92.0% and 90.7% for the target families of enzymes, ion channels, GPCRs and nuclear receptors of our created dataset, respectively. Another dataset derived from DrugBank was used to further assess the generalization of the model, which yielded an accuracy of 0.9015 and an AUC value of 0.9557. CONCLUSION: It was elucidated that our model shows improved performance in comparison with other state-of-the-art computational methods on the common benchmark datasets. Experimental results demonstrated that our model successfully extracted more nuanced yet useful features, and therefore can be used as a practical tool to discover new drugs. AVAILABILITY: http://deeplearner.ahu.edu.cn/web/CnnDTI.htm. Peng Chen 0001, Pengying Gu, Jun Zhang 0011, Bing Wang 0004 |
BMC Bioinform. | 3 |
| 2019 | Semi-supervised prediction of protein interaction sites from unlabeled sample informationabstractBACKGROUND: The recognition of protein interaction sites is of great significance in many biological processes, signaling pathways and drug designs. However, most sites on protein sequences cannot be defined as interface or non-interface sites because only a small part of protein interactions had been identified, which will cause the lack of prediction accuracy and generalization ability of predictors in protein interaction sites prediction. Therefore, it is necessary to effectively improve prediction performance of protein interaction sites using large amounts of unlabeled data together with small amounts of labeled data and background knowledge today. RESULTS: In this work, three semi-supervised support vector machine-based methods are proposed to improve the performance in the protein interaction sites prediction, in which the information of unlabeled protein sites can be involved. Herein, five features related with the evolutionary conservation of amino acids are extracted from HSSP database and Consurf Sever, i.e., residue spatial sequence spectrum, residue sequence information entropy and relative entropy, residue sequence conserved weight and residual Base evolution rate, to represent the residues within the protein sequence. Then three predictors are built for identifying the interface residues from protein surface using three types of semi-supervised support vector machine algorithms. CONCLUSION: The experimental results demonstrated that the semi-supervised approaches can effectively improve prediction performance of protein interaction sites when unlabeled information is involved into the predictors and one of them can achieve the best prediction performance, i.e., the accuracy of 70.7%, the sensitivity of 62.67% and the specificity of 78.72%, respectively. With comparison to the existing studies, the semi-supervised models show the improvement of the predication performance. Changqing Mei, Yuming Zhou, Chun-Hou Zheng 0001, Xiao Zhen, Yan Xiong 0001, Peng Chen 0001, Jun Zhang 0011, Bing Wang 0004 |
BMC Bioinform. | 8 |
| 2019 | Occurrence prediction of pests and diseases in cotton on the basis of weather factors by long short term memory networkabstractBACKGROUND: The occurrence of cotton pests and diseases has always been an important factor affecting the total cotton production. Cotton has a great dependence on environmental factors during its growth, especially climate change. In recent years, machine learning and especially deep learning methods have been widely used in many fields and have achieved good results. METHODS: First, this papaer used the common Aprioro algorithm to find the association rules between weather factors and the occurrence of cotton pests. Then, in this paper, the problem of predicting the occurrence of pests and diseases is formulated as time series prediction, and an LSTM-based method was developed to solve the problem. RESULTS: The association analysis reveals that moderate temperature, humid air, low wind spreed and rain fall in autumn and winter are more likely to occur cotton pests and diseases. The discovery was then used to predict the occurrence of pests and diseases. Experimental results showed that LSTM performs well on the prediction of occurrence of pests and diseases in cotton fields, and yields the Area Under the Curve (AUC) of 0.97. CONCLUSION: Suitable temperature, humidity, low rainfall, low wind speed, suitable sunshine time and low evaporation are more likely to cause cotton pests and diseases. Based on these associations as well as historical weather and pest records, LSTM network is a good predictor for future pest and disease occurrences. Moreover, compared to the traditional machine learning models (i.e., SVM and Random Forest), the LSTM network performs the best. Qingxin Xiao, Weilu Li, Yuanzhong Kai, Peng Chen 0001, Jun Zhang 0011, Bing Wang 0004 |
BMC Bioinform. | 4 |
| 2019 | Deep spatial attention hashing network for image retrieval
Lin-Wei Ge, Jun Zhang 0011, Peng Chen 0001, Bing Wang 0004, Chun-Hou Zheng 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2018 | Similarity-Based Integrated Method for Predicting Drug-Disease Interactions
Yan-Zhe Di, Peng Chen 0001, Chun-Hou Zheng 0001 |
ICIC (2) | 2 |
| 2018 | Chinese Text Detection Using Deep Learning Model and Synthetic Data
Wei-wei Gao, Jun Zhang 0011, Peng Chen 0001, Bing Wang 0004 |
ICIC (1) | 3 |
| 2018 | Convolutional Neural Network for Short Term Fog Forecasting Based on Meteorological Elements
Ting-Ting Han, Kai-Chao Miao, Ye-Qing Yao, Cheng-Xiao Liu, Jian-Ping Zhou, Peng Chen 0001, Xia Yi, Bing Wang 0004, Jun Zhang 0011 |
ICIC (3) | 7 |
| 2018 | Using Novel Convolutional Neural Networks Architecture to Predict Drug-Target Interactions
DeNan Xia, Peng Chen 0001, Bing Wang 0004 |
ICIC (2) | 3 |
| 2018 | Prediction of Protein-Protein Interaction Sites Combing Sequence Profile and Hydrophobic Information
Lili Peng, Nian Zhou, Peng Chen 0001, Jun Zhang 0011, Bing Wang 0004 |
ICIC (1) | 4 |
| 2018 | Cells Counting with Convolutional Neural Network
Run-xu Tan, Jun Zhang 0011, Peng Chen 0001, Bing Wang 0004 |
ICIC (3) | 3 |
| 2018 | Prediction of Crop Pests and Diseases in Cotton by Long Short Term Memory Network
Qingxin Xiao, Weilu Li, Peng Chen 0001, Bing Wang 0004 |
ICIC (2) | 3 |
| 2018 | Deep Convolutional Neural Network for Fog Detection
Jun Zhang 0011, Ting-Ting Han, Kai-Chao Miao, Ye-Qing Yao, Cheng-Xiao Liu, Jian-Ping Zhou, Peng Chen 0001, Bing Wang 0004 |
ICIC (2) | 9 |
| 2018 | Verifying TCM Syndrome Hypothesis Based on Improved Latent Tree Model
Nian Zhou, Lingshan Zhou, Lili Peng, Bing Wang 0004, Peng Chen 0001, Jun Zhang 0011 |
ICIC (2) | 5 |
| 2018 | dbMPIKT: a database of kinetic and thermodynamic mutant protein interactionsabstractBACKGROUND: Protein-protein interactions (PPIs) play important roles in biological functions. Studies of the effects of mutants on protein interactions can provide further understanding of PPIs. Currently, many databases collect experimental mutants to assess protein interactions, but most of these databases are old and have not been updated for several years. RESULTS: To address this issue, we manually curated a kinetic and thermodynamic database of mutant protein interactions (dbMPIKT) that is freely accessible at our website. This database contains 5291 mutants in protein interactions collected from previous databases and the literature published within the last three years. Furthermore, some data analysis, such as mutation number, mutation type, protein pair source and network map construction, can be performed online. CONCLUSION: Our work can promote the study on PPIs, and novel information can be mined from the new database. Our database is available in http://DeepLearner.ahu.edu.cn/web/dbMPIKT/ for use by all, including both academics and non-academics. Quanya Liu, Peng Chen 0001, Bing Wang 0004, Jun Zhang 0011, Jinyan Li 0001 |
BMC Bioinform. | 2 |
| 2018 | Protein-protein interface hot spots prediction based on a hybrid feature selection strategyabstractBACKGROUND: Hot spots are interface residues that contribute most binding affinity to protein-protein interaction. A compact and relevant feature subset is important for building machine learning methods to predict hot spots on protein-protein interfaces. Although different methods have been used to detect the relevant feature subset from a variety of features related to interface residues, it is still a challenge to detect the optimal feature subset for building the final model. RESULTS: In this study, three different feature selection methods were compared to propose a new hybrid feature selection strategy. This new strategy was proved to effectively reduce the feature space when we were building the prediction models for identifying hotspot residues. It was tested on eighty-two features, both conventional and newly proposed. According to the strategy, combining the feature subsets selected by decision tree and mRMR (maximum Relevance Minimum Redundancy) individually, we were able to build a model with 6 features by using a PSFS (Pseudo Sequential Forward Selection) process. Compared with other state-of-art methods for the independent test set, our model had shown better or comparable predictive performances (with F-measure 0.622 and recall 0.821). Analysis of the 6 features confirmed that our newly proposed feature CNSV_REL1 was important for our model. The analysis also showed that the complementarity between features should be considered as an important aspect when conducting the feature selection. CONCLUSION: In this study, most important of all, a new strategy for feature selection was proposed and proved to be effective in selecting the optimal feature subset for building prediction models, which can be used to predict hot spot residues on protein-protein interfaces. Moreover, two aspects, the generalization of the single feature and the complementarity between features, were proved to be of great importance and should be considered in feature selection methods. Finally, our newly proposed feature CNSV_REL1 had been proved an alternative and effective feature in predicting hot spots by our study. Our model is available for users through a webserver: http://zhulab.ahu.edu.cn/iPPHOT/ . Yanhua Qiao, Yi Xiong 0002, Xiaolei Zhu 0001, Peng Chen 0001 |
BMC Bioinform. | 5 |
| 2017 | CAPTCHA Recognition Based on Faster R-CNN
Feng-Lin Du, Peng Chen 0001, Bing Wang 0004, Jun Zhang 0011 |
ICIC (2) | 4 |
| 2017 | A Machine Vision Method for Automatic Circular Parts Detection Based on Optimization Algorithm
Kun Lu 0007, Rui Hong, Peng Chen 0001, Jun Zhang 0011, Bing Wang 0004 |
ICIC (1) | 4 |
| 2017 | Utilization of rotation-invariant uniform LBP histogram distribution and statistics of connected regions in automatic image annotation based on multi-label learning
Sen Xia, Peng Chen 0001, Jun Zhang 0011, Bing Wang 0004 |
Neurocomputing | 2 |
| 2017 | DrugRPE: Random projection ensemble approach to drug-target interaction prediction
Jun Zhang 0011, Muchun Zhu, Peng Chen 0001, Bing Wang 0004 |
Neurocomputing | 3 |
| 2017 | Optimization enhanced genetic algorithm-support vector regression for the prediction of compound retention indices in gas chromatography
Jun Zhang 0011, Chun-Hou Zheng 0001, Bing Wang 0004, Peng Chen 0001 |
Neurocomputing | 5 |
| 2017 | Robust object tracking via multi-scale patch based sparse coding histogram
Zhongpei Wang, Hao Wang 0008, Jieqing Tan, Peng Chen 0001, Chengjun Xie |
Multim. Tools Appl. | 4 |
| 2016 | Prediction of Hot Spots Based on Physicochemical Features and Relative Accessible Surface Area of Amino Acid Sequence
Peng Chen 0001, Jun Zhang 0011, Bing Wang 0004 |
ICIC (1) | 2 |
| 2016 | Accurate Prediction of Protein Hot Spots Residues Based on Gentle AdaBoost Algorithm
Jun Zhang 0011, Chun-Hou Zheng 0001, Bing Wang 0004, Peng Chen 0001 |
ICIC (1) | 5 |
| 2016 | Inferring Disease-Related Domain Using Network-Based Method
Zhongwen Zhang, Peng Chen 0001, Jun Zhang 0011, Bing Wang 0004 |
ICIC (1) | 2 |
| 2016 | A Sequence-Based Dynamic Ensemble Learning System for Protein Ligand-Binding Site PredictionabstractBACKGROUND: Proteins have the fundamental ability to selectively bind to other molecules and perform specific functions through such interactions, such as protein-ligand binding. Accurate prediction of protein residues that physically bind to ligands is important for drug design and protein docking studies. Most of the successful protein-ligand binding predictions were based on known structures. However, structural information is not largely available in practice due to the huge gap between the number of known protein sequences and that of experimentally solved structures. RESULTS: This paper proposes a dynamic ensemble approach to identify protein-ligand binding residues by using sequence information only. To avoid problems resulting from highly imbalanced samples between the ligand-binding sites and non ligand-binding sites, we constructed several balanced data sets and we trained a random forest classifier for each of them. We dynamically selected a subset of classifiers according to the similarity between the target protein and the proteins in the training data set. The combination of the predictions of the classifier subset to each query protein target yielded the final predictions. The ensemble of these classifiers formed a sequence-based predictor to identify protein-ligand binding sites. CONCLUSIONS: Experimental results on two Critical Assessment of protein Structure Prediction datasets and the ccPDB dataset demonstrated that of our proposed method compared favorably with the state-of-the-art. AVAILABILITY: http://www2.ahu.edu.cn/pchen/web/LigandDSES.htm. Peng Chen 0001, Jun Zhang 0011, Xin Gao 0001, Jinyan Li 0001, Junfeng Xia, Bing Wang 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2015 | Compound Identification Using Random Projection for Gas Chromatography-Mass Spectrometry Data
Li-Li Cao, Zhi-Shui Zhang, Peng Chen 0001, Jun Zhang 0011 |
ICIC (3) | 3 |
| 2015 | Sequence-Based Random Projection Ensemble Approach to Identify Hotspot Residues from Whole Protein Sequence
Peng Chen 0001, Bing Wang 0004, Jun Zhang 0011 |
ICIC (2) | 1 |
| 2015 | A Random Projection Ensemble Approach to Drug-Target Interaction Prediction
Peng Chen 0001, Bing Wang 0004, Jun Zhang 0011 |
ICIC (3) | 1 |
| 2015 | A Multi-feature Fusion Method for Automatic Multi-label Image Annotation with Weighted Histogram Integral and Closure Regions Counting
Sen Xia, Peng Chen 0001, Jun Zhang 0011, Bing Wang 0004 |
ICIC (3) | 2 |
| 2015 | Prediction of Molecular Substructure Using Mass Spectral Data Based on Deep Learning
Zhi-Shui Zhang, Li-Li Cao, Jun Zhang 0011, Peng Chen 0001, Chun-Hou Zheng 0001 |
ICIC (2) | 4 |
| 2014 | LigandRFs: random forest ensemble to identify ligand-binding residues from sequence information aloneabstractBACKGROUND: Protein-ligand binding is important for some proteins to perform their functions. Protein-ligand binding sites are the residues of proteins that physically bind to ligands. Despite of the recent advances in computational prediction for protein-ligand binding sites, the state-of-the-art methods search for similar, known structures of the query and predict the binding sites based on the solved structures. However, such structural information is not commonly available. RESULTS: In this paper, we propose a sequence-based approach to identify protein-ligand binding residues. We propose a combination technique to reduce the effects of different sliding residue windows in the process of encoding input feature vectors. Moreover, due to the highly imbalanced samples between the ligand-binding sites and non ligand-binding sites, we construct several balanced data sets, for each of which a random forest (RF)-based classifier is trained. The ensemble of these RF classifiers forms a sequence-based protein-ligand binding site predictor. CONCLUSIONS: Experimental results on CASP9 and CASP8 data sets demonstrate that our method compares favorably with the state-of-the-art protein-ligand binding site prediction methods. Peng Chen 0001, Jianhua Z. Huang, Xin Gao 0001 |
BMC Bioinform. | 1 |
| 2014 | Sparse Representation-Based Approach for Unsupervised Feature SelectionabstractDimension reduction methods including feature selection and feature extraction have played an important role in data mining and pattern recognition. In this study, we propose a novel unsupervised feature selection approach based on sparse representation theory, namely Sparsity Score (SS). Due to the sparse representation procedure, SS not only owns the global property of Variance Score (VS) and the local property of Laplacian Score (LS), but also possesses the discriminating nature. Experimental results, based on three well-known face datasets (Yale, ORL and CMU PIE), reveal that SS performs well in the evaluation of the feature significance, and it significantly outperforms VS and LS. Yaru Su, Chuanxi Li, Rujing Wang, Peng Chen 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2014 | Collaborative object tracking model with local sparse representation
Chengjun Xie, Jieqing Tan, Peng Chen 0001, Jie Zhang 0033, Lei He 0002 |
J. Vis. Commun. Image Represent. | 3 |
| 2014 | Multi-scale patch-based sparse appearance model for robust object tracking
Chengjun Xie, Jieqing Tan, Peng Chen 0001, Jie Zhang 0033, Lei He 0002 |
Mach. Vis. Appl. | 3 |
| 2013 | Consensus of Sample-Balanced Classifiers for Identifying Ligand-Binding Residue by Co-evolutionary Physicochemical Characteristics of Amino Acids
Peng Chen 0001 |
ICIC (3) | 1 |
| 2013 | Prediction of peptide drift time in ion mobility mass spectrometry from sequence-based featuresabstractBACKGROUND: Ion mobility-mass spectrometry (IMMS), an analytical technique which combines the features of ion mobility spectrometry (IMS) and mass spectrometry (MS), can rapidly separates ions on a millisecond time-scale. IMMS becomes a powerful tool to analyzing complex mixtures, especially for the analysis of peptides in proteomics. The high-throughput nature of this technique provides a challenge for the identification of peptides in complex biological samples. As an important parameter, peptide drift time can be used for enhancing downstream data analysis in IMMS-based proteomics. RESULTS: In this paper, a model is presented based on least square support vectors regression (LS-SVR) method to predict peptide ion drift time in IMMS from the sequence-based features of peptide. Four descriptors were extracted from peptide sequence to represent peptide ions by a 34-component vector. The parameters of LS-SVR were selected by a grid searching strategy, and a 10-fold cross-validation approach was employed for the model training and testing. Our proposed method was tested on three datasets with different charge states. The high prediction performance achieve demonstrate the effectiveness and efficiency of the prediction model. CONCLUSIONS: Our proposed LS-SVR model can predict peptide drift time from sequence information in relative high prediction accuracy by a test on a dataset of 595 peptides. This work can enhance the confidence of protein identification by combining with current protein searching techniques. Bing Wang 0004, Jun Zhang 0011, Peng Chen 0001, Zhiwei Ji, Suping Deng |
BMC Bioinform. | 3 |
| 2013 | Multiple instance learning tracking method with local sparse representationabstractWhen objects undergo large pose change, illumination variation or partial occlusion, most existed visual tracking algorithms tend to drift away from targets and even fail in tracking them. To address this issue, in this study, the authors propose an online algorithm by combining multiple instance learning (MIL) and local sparse representation for tracking an object in a video system. The key idea in our method is to model the appearance of an object by local sparse codes that can be formed as training data for the MIL framework. First, local image patches of a target object are represented as sparse codes with an overcomplete dictionary, where the adaptive representation can be helpful in overcoming partial occlusion in object tracking. Then MIL learns the sparse codes by a classifier to discriminate the target from the background. Finally, results from the trained classifier are input into a particle filter framework to sequentially estimate the target state over time in visual tracking. In addition, to decrease the visual drift because of the accumulative errors when updating the dictionary and classifier, a two‐step object tracking method combining a static MIL classifier with a dynamical MIL classifier is proposed. Experiments on some publicly available benchmarks of video sequences show that our proposed tracker is more robust and effective than others. Chengjun Xie, Jieqing Tan, Peng Chen 0001, Jie Zhang 0033, Lei He 0002 |
IET Comput. Vis. | 3 |
| 2012 | Detection of Outlier Residues for Improving Interface Prediction in Protein HeterocomplexesabstractSequence-based understanding and identification of protein binding interfaces is a challenging research topic due to the complexity in protein systems and the imbalanced distribution between interface and noninterface residues. This paper presents an outlier detection idea to address the redundancy problem in protein interaction data. The cleaned training data are then used for improving the prediction performance. We use three novel measures to describe the extent a residue is considered as an outlier in comparison to the other residues: the distance of a residue instance from the center instance of all residue instances of the same class label (Dist), the probability of the class label of the residue instance (PCL), and the importance of within-class and between-class (IWB) residue instances. Outlier scores are computed by integrating the three factors; instances with a sufficiently large score are treated as outliers and removed. The data sets without outliers are taken as input for a support vector machine (SVM) ensemble. The proposed SVM ensemble trained on input data without outliers performs better than that with outliers. Our method is also more accurate than many literature methods on benchmark data sets. From our empirical studies, we found that some outlier interface residues are truly near to noninterface regions, and some outlier noninterface residues are close to interface regions. Peng Chen 0001, Limsoon Wong, Jinyan Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2011 | Protein Interface Residues Prediction Based on Amino Acid Properties Only
Bing Wang 0004, Peng Chen 0001, Jun Zhang 0011 |
ICIC (3) | 2 |
| 2010 | Sequence-based identification of interface residues by an integrative profile combining hydrophobic and evolutionary informationabstractBACKGROUND: Protein-protein interactions play essential roles in protein function determination and drug design. Numerous methods have been proposed to recognize their interaction sites, however, only a small proportion of protein complexes have been successfully resolved due to the high cost. Therefore, it is important to improve the performance for predicting protein interaction sites based on primary sequence alone. RESULTS: We propose a new idea to construct an integrative profile for each residue in a protein by combining its hydrophobic and evolutionary information. A support vector machine (SVM) ensemble is then developed, where SVMs train on different pairs of positive (interface sites) and negative (non-interface sites) subsets. The subsets having roughly the same sizes are grouped in the order of accessible surface area change before and after complexation. A self-organizing map (SOM) technique is applied to group similar input vectors to make more accurate the identification of interface residues. An ensemble of ten-SVMs achieves an MCC improvement by around 8% and F1 improvement by around 9% over that of three-SVMs. As expected, SVM ensembles constantly perform better than individual SVMs. In addition, the model by the integrative profiles outperforms that based on the sequence profile or the hydropathy scale alone. As our method uses a small number of features to encode the input vectors, our model is simpler, faster and more accurate than the existing methods. CONCLUSIONS: The integrative profile by combining hydrophobic and evolutionary information contributes most to the protein-protein interaction prediction. Results show that evolutionary context of residue with respect to hydrophobicity makes better the identification of protein interface residues. In addition, the ensemble of SVM classifiers improves the prediction performance. AVAILABILITY: Datasets and software are available at http://mail.ustc.edu.cn/~bigeagle/BMCBioinfo2010/index.htm. Peng Chen 0001, Jinyan Li 0001 |
BMC Bioinform. | 1 |
| 2008 | Prediction of Inter-residue Contact Clusters from Hydrophobic CoresabstractMotivation: Contact map is a key factor to represent a specific protein structure. As previous work reported, even a corrupted contact map can be used to reconstruct its corresponding protein structure. Thus we can predict the structure of a protein partially through the contact map prediction. To simplify the protein contact map prediction, we predict the inter-residue contact clusters centered at the groups of their neighboring contacts instead.Results: In this paper, we adopt a SVM predictor based approach to predict the inter-residue contact cluster centers. The input information of the SVM predictor includes sequence profile, evolutionary rate, and predicted secondary structure. The SVM predictor is based on hydrophobic cores that may be considered as locations of the groups of their neighboring inter-residue contacts. As a result, about 35% clustering centers of inter-residue contacts can accurately be predicted. Peng Chen 0001, Legand L. Burge III, Mahmood Mohammad, William M. Southerland, Clay S. Gloster Jr. |
ICMLA | 1 |
| 2007 | Prediction of Long-range Contacts from Sequence ProfileabstractTheoretic study in this paper shows that we can obtain exact long-range contacts by adopting one classifier if the centers of sequence profiles of residue pairs for long-range contacts and non-long-range contacts are known. The adopted classifier, referred to as multiple conditional probability mass function classifier (MCPMFC), can find an optimized transformation of the variables for each of the classes and therefore resulting in K separate classifiers. As a result, about 44.48% long-range contacts are around at the sequence profile (SP) centre for long-range contacts and about 20.9% long-range contacts are correctly predicted when considering the top L/5 (L is the protein sequence length) predicted contacts and the residue pair with 24 apart. The highest cluster result gives us a clue that SP center should be a sound pathway to investigate contact map in protein structures. Peng Chen 0001, Bing Wang 0004, Hau-San Wong, De-Shuang Huang |
IJCNN | 1 |
| 2006 | Long-Range Interaction Analysis using Principal Component AnalysisabstractThis paper analyzes the long-range interactions, which plays a fundamental and important role in many biologic fields, between residues in protein using principal component analysis (PCA). Firstly, one angular coordinate system of long-range interaction regions is constructed conveniently. Afterwards, a matrix of the angular values of residues can be analyzed by principal component analysis technique. Projecting the angular matrix onto its eigenvectors, it can be found that the projection is to satisfy Boltzmann distribution. By analyzing the thermodynamic environment of the interaction region and scaling the interaction regions, it can be concluded that the distribution of long-range interactions may also be obtained and as a result applied in prediction of contact map. Peng Chen 0001, Bing Wang 0004, Hau-San Wong, De-Shuang Huang |
IJCNN | 1 |
| 2006 | Predicting Protein-Protein Interaction Sites using Radial Basis Function Neural NetworksabstractIdentifying protein-protein interaction sites is crucial for understanding of the principles of biological systems and processes, as well as mutant design. This paper describes a novel method that can predict protein interaction sites in heterocomplexes using information of evolutionary conservation and spatial sequence profile. A predictor was generated to distinguish the interface residues from protein surface region by radial basis neural networks, which is trained by expectation maximization algorithm. Based on a non-redundant data set of heterodimers consisting of 75 protein chains, the efficiency and the effectiveness of our proposed approach can be validated by a better performance such as the accuracy of 0.60, the sensitivity of 58.3% and the specificity of 59.9%. Bing Wang 0004, Hau-San Wong, Peng Chen 0001, Hong-Qiang Wang, De-Shuang Huang |
IJCNN | 3 |
| 2005 | Prediction of contact map integrated PNN with conformational energyabstractThis paper presents a novel method to solve the protein's three-dimensional structure prediction problem. It is a machine learning approach by integrating probabilistic neural network (PNN) with conformational energy function (CEF) based on chemico-physical knowledge of amino acids. In this method, firstly, the principal components are extracted from selected protein structures with lower sequence identity, and an initial matrix of contact map is constructed by K-L expansion. Secondly, PNN is used for predicting the long-range interaction of amino acids in protein. In particular, this method uses the CEF and chemico-physical characteristics of amino acids to run the PNN predictor. Consequently, it was found that our proposed method is better than existing methods, such as the hybrid method of HMMSTR and the correlated mutation analysis method. As a result, this method can accurately predict 31% of contacts at a distance cutoff of 8/spl Aring/ for proteins whose sequence length is up to 200. Peng Chen 0001, De-Shuang Huang, Bing Wang 0004 |
IJCNN | 1 |
| 2005 | Predicting protein-protein interactions based on protein-domain relationshipsabstractThis paper proposes a new method that can predict the interactions between proteins intermediated by the protein-domain relations. We utilize the lazy expectation maximization (LEM) to compute an improved maximization likelihood estimation (MLE) model. The protein-domain relationships are extruded from Flam database and the combined data set of Uetz and Ito are used as the source of protein-protein interactions. Finally, the efficiency and the effectiveness of our proposed approach can be validated by a better performance such as the sensitivity of 80.1%, the specificity of 43.5%, and the lesser computational cost. Bing Wang 0004, De-Shuang Huang, Peng Chen 0001 |
IJCNN | 3 |