Haiou Li

dblp:146/4319 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
6since 2021 · last 2022
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2022 G Protein-Coupled Receptor Interaction Prediction Based on Deep Transfer Learning
abstract
G protein-coupled receptors (GPCRs) account for about 40% to 50% of drug targets. Many human diseases are related to G protein coupled receptors. Accurate prediction of GPCR interaction is not only essential to understand its structural role, but also helps design more effective drugs. At present, the prediction of GPCR interaction mainly uses machine learning methods. Machine learning methods generally require a large number of independent and identically distributed samples to achieve good results. However, the number of available GPCR samples that have been marked is scarce. Transfer learning has a strong advantage in dealing with such small sample problems. Therefore, this paper proposes a transfer learning method based on sample similarity, using XGBoost as a weak classifier and using the TrAdaBoost algorithm based on JS divergence for data weight initialization to transfer samples to construct a data set. After that, the deep neural network based on the attention mechanism is used for model training. The existing GPCR is used for prediction. In short-distance contact prediction, the accuracy of our method is 0.26 higher than similar methods.
Tengsheng Jiang, Yuhui Chen, Zhongtian Hu, Weizhong Lu, Qiming Fu 0001, Yijie Ding, Haiou Li, Hongjie Wu
IEEE ACM Trans. Comput. Biol. Bioinform.8
2021 DNA-Binding Protein Prediction Based on Deep Learning Feature Fusion
Tengsheng Jiang, Weizhong Lu, Qiming Fu 0001, Haiou Li, Hongjie Wu
ICIC (3)5
2021 Super-Large Medical Image Storage and Display Technology Based on Concentrated Points of Interest
Yuli Wang, Haiou Li, Weizhong Lu, Hongjie Wu
ICIC (1)3
2021 Research on RNA secondary structure predicting via bidirectional recurrent neural network
abstract
BACKGROUND: RNA secondary structure prediction is an important research content in the field of biological information. Predicting RNA secondary structure with pseudoknots has been proved to be an NP-hard problem. Traditional machine learning methods can not effectively apply protein sequence information with different sequence lengths to the prediction process due to the constraint of the self model when predicting the RNA secondary structure. In addition, there is a large difference between the number of paired bases and the number of unpaired bases in the RNA sequences, which means the problem of positive and negative sample imbalance is easy to make the model fall into a local optimum. To solve the above problems, this paper proposes a variable-length dynamic bidirectional Gated Recurrent Unit(VLDB GRU) model. The model can accept sequences with different lengths through the introduction of flag vector. The model can also make full use of the base information before and after the predicted base and can avoid losing part of the information due to truncation. Introducing a weight vector to predict the RNA training set by dynamically adjusting each base loss function solves the problem of balanced sample imbalance. RESULTS: The algorithm proposed in this paper is compared with the existing algorithms on five representative subsets of the data set RNA STRAND. The experimental results show that the accuracy and Matthews correlation coefficient of the method are improved by 4.7% and 11.4%, respectively. CONCLUSIONS: The flag vector introduced allows the model to effectively use the information before and after the protein sequence; the introduced weight vector solves the problem of unbalanced sample balance. Compared with other algorithms, the LVDB GRU algorithm proposed in this paper has the best detection results.
Weizhong Lu, Hongjie Wu, Yijie Ding, Zhengwei Song, Yu Zhang 0027, Qiming Fu 0001, Haiou Li
BMC Bioinform.8
2021 Band-gap tunable (GaxIn1-x)2O3 layer grown by magnetron sputtering
abstract
Multicomponent oxide (GaxIn1−x)2O3 films are prepared on (0001) sapphire substrates to realize a tunable band-gap by magnetron sputtering technology followed by thermal annealing. The optical properties and band structure evolution over the whole range of compositions in ternary compounds (GaxIn1−x)2O3 are investigated in detail. The X-ray diffraction spectra clearly indicate that (GaxIn1−x)2O3 films with Ga content varying from 0.11 to 0.55 have both cubic and monoclinic structures, and that for films with Ga content higher than 0.74, only the monoclinic structure appears. The transmittance of all films is greater than 86% in the visible range with sharp absorption edges and clear fringes. In addition, a blue shift of ultraviolet absorption edges from 380 to 250 nm is noted with increasing Ga content, indicating increasing band-gap energy from 3.61 to 4.64 eV. The experimental results lay a foundation for the application of transparent conductive compound (GaxIn1−x)2O3 thin films in photoelectric and photovoltaic industry, especially in display, light-emitting diode, and solar cell applications.
Fabi Zhang, Jinyu Sun, Haiou Li, Tangyou Sun, Gongli Xiao, Xingpeng Liu, Xiuyun Zhang, Daoyou Guo, Xianghu Wang, Zujun Qin
Frontiers Inf. Technol. Electron. Eng.3
2021 Empirical Potential Energy Function Toward ab Initio Folding G Protein-Coupled Receptors
abstract
Approximately 40-50 percent of all drugs targets are G protein-coupled receptors (GPCRs). Three-dimensional structure of GPCRs is important to probe their biophysical and biochemical functions and their pharmaceutical applications. Lacking reliable and high quality free function is one of the ugent problems of computational predicting the three-dimensional structure in this community. We proposed a GPCR-specified energy function composed of four novel empirical potential energy terms: a two-dimensional contact energy force field, knowledge-based helix pair connection distance energy term, knowledge-based helix pair angle restraint energy term and a disulfide bond energy term. To validate the energy function, we employed an ab initio GPCR three-dimensional structure predictor to test if the energy function improved the accuracy of prediction. We evaluated 28 solved GPCRs and found that 21(75 percent) targets were correctly folded (TM-score>0.5). Also, the average TM-score using the energy function was 0.54, which was improved 134 percent than the TM-score 0.23 for MODELLER energy function and 170 percent than the TM-score 0.20 for Rosetta membrane energy function. The results confirmed that our empirical potential energy function toward ab initio folding is competitive to state-of-the-art solutions for structural prediction of GPCRs.
Hongjie Wu, Huajing Ling, Qiming Fu 0001, Weizhong Lu, Yijie Ding, Min Jiang 0009, Haiou Li
IEEE ACM Trans. Comput. Biol. Bioinform.8
2020 Prediction of Membrane Protein Interaction Based on Deep Residual Learning
Tengsheng Jiang, Hongjie Wu, Yuhui Chen, Haiou Li, Jin Qiu, Weizhong Lu, Qiming Fu 0001
ICIC (2)4
2019 Predicting RNA secondary structure via adaptive deep recurrent neural networks with energy-based filter
abstract
BACKGROUND: RNA secondary structure prediction is an important issue in structural bioinformatics, and RNA pseudoknotted secondary structure prediction represents an NP-hard problem. Recently, many different machine-learning methods, Markov models, and neural networks have been employed for this problem, with encouraging results regarding their predictive accuracy; however, their performances are usually limited by the requirements of the learning model and over-fitting, which requires use of a fixed number of training features. Because most natural biological sequences have variable lengths, the sequences have to be truncated before the features are employed by the learning model, which not only leads to the loss of information but also destroys biological-sequence integrity. RESULTS: To address this problem, we propose an adaptive sequence length based on deep-learning model and integrate an energy-based filter to remove the over-fitting base pairs. CONCLUSIONS: Comparative experiments conducted on an authoritative dataset RNA STRAND (RNA secondary STRucture and statistical Analysis Database) revealed a 12% higher accuracy relative to three currently used methods.
Weizhong Lu, Hongjie Wu, Hongmei Huang, Qiming Fu 0001, Haiou Li
BMC Bioinform.7
2019 Ranking near-native candidate protein structures via random forest classification
abstract
BACKGROUND: In ab initio protein-structure predictions, a large set of structural decoys are often generated, with the requirement to select best five or three candidates from the decoys. The clustered central structures with the most number of neighbors are frequently regarded as the near-native protein structures with the lowest free energy; however, limitations in clustering methods and three-dimensional structural-distance assessments make identifying exact order of the best five or three near-native candidate structures difficult. RESULTS: To address this issue, we propose a method that re-ranks the candidate structures via random forest classification using intra- and inter-cluster features from the results of the clustering. Comparative analysis indicated that our method was better able to identify the order of the candidate structures as comparing with current methods SPICKR, Calibur, and Durandal. The results confirmed that the identification of the first model were closer to the native structure in 12 of 43 cases versus four for SPICKER, and the same as the native structure in up to 27 of 43 cases versus 14 for Calibur and up to eight of 43 cases versus two for Durandal. CONCLUSIONS: In this study, we presented an improved method based on random forest classification to transform the problem of re-ranking the candidate structures by an binary classification. Our results indicate that this method is a powerful method for the problem and the effect of this method is better than other methods.
Hongjie Wu, Hongmei Huang, Weizhong Lu, Qiming Fu 0001, Yijie Ding, Haiou Li
BMC Bioinform.7
2019 Research on predicting 2D-HP protein folding using reinforcement learning with full state space
abstract
BACKGROUND: Protein structure prediction has always been an important issue in bioinformatics. Prediction of the two-dimensional structure of proteins based on the hydrophobic polarity model is a typical non-deterministic polynomial hard problem. Currently reported hydrophobic polarity model optimization methods, greedy method, brute-force method, and genetic algorithm usually cannot converge robustly to the lowest energy conformations. Reinforcement learning with the advantages of continuous Markov optimal decision-making and maximizing global cumulative return is especially suitable for solving global optimization problems of biological sequences. RESULTS: In this study, we proposed a novel hydrophobic polarity model optimization method derived from reinforcement learning which structured the full state space, and designed an energy-based reward function and a rigid overlap detection rule. To validate the performance, sixteen sequences were selected from the classical data set. The results indicated that reinforcement learning with full states successfully converged to the lowest energy conformations against all sequences, while the reinforcement learning with partial states folded 50% sequences to the lowest energy conformations. Reinforcement learning with full states hits the lowest energy on an average 5 times, which is 40 and 100% higher than the three and zero hit by the greedy algorithm and reinforcement learning with partial states respectively in the last 100 episodes. CONCLUSIONS: Our results indicate that reinforcement learning with full states is a powerful method for predicting two-dimensional hydrophobic-polarity protein structure. It has obvious competitive advantages compared with greedy algorithm and reinforcement learning with partial states.
Hongjie Wu, Qiming Fu 0001, Weizhong Lu, Haiou Li
BMC Bioinform.6
2017 Deep learning methods for protein torsion angle prediction
abstract
BACKGROUND: Deep learning is one of the most powerful machine learning methods that has achieved the state-of-the-art performance in many domains. Since deep learning was introduced to the field of bioinformatics in 2012, it has achieved success in a number of areas such as protein residue-residue contact prediction, secondary structure prediction, and fold recognition. In this work, we developed deep learning methods to improve the prediction of torsion (dihedral) angles of proteins. RESULTS: We design four different deep learning architectures to predict protein torsion angles. The architectures including deep neural network (DNN) and deep restricted Boltzmann machine (DRBN), deep recurrent neural network (DRNN) and deep recurrent restricted Boltzmann machine (DReRBM) since the protein torsion angle prediction is a sequence related problem. In addition to existing protein features, two new features (predicted residue contact number and the error distribution of torsion angles extracted from sequence fragments) are used as input to each of the four deep learning architectures to predict phi and psi angles of protein backbone. The mean absolute error (MAE) of phi and psi angles predicted by DRNN, DReRBM, DRBM and DNN is about 20-21° and 29-30° on an independent dataset. The MAE of phi angle is comparable to the existing methods, but the MAE of psi angle is 29°, 2° lower than the existing methods. On the latest CASP12 targets, our methods also achieved the performance better than or comparable to a state-of-the art method. CONCLUSIONS: Our experiment demonstrates that deep learning is a valuable method for predicting protein torsion angles. The deep recurrent network architecture performs slightly better than deep feed-forward architecture, and the predicted residue contact number and the error distribution of torsion angles extracted from sequence fragments are useful features for improving prediction accuracy.
Haiou Li, Jie Hou 0001, Badri Adhikari, Qiang Lyu, Jianlin Cheng
BMC Bioinform.1
2014 Improved packing of protein side chains with parallel ant colonies
abstract
INTRODUCTION: The accurate packing of protein side chains is important for many computational biology problems, such as ab initio protein structure prediction, homology modelling, and protein design and ligand docking applications. Many of existing solutions are modelled as a computational optimisation problem. As well as the design of search algorithms, most solutions suffer from an inaccurate energy function for judging whether a prediction is good or bad. Even if the search has found the lowest energy, there is no certainty of obtaining the protein structures with correct side chains. METHODS: We present a side-chain modelling method, pacoPacker, which uses a parallel ant colony optimisation strategy based on sharing a single pheromone matrix. This parallel approach combines different sources of energy functions and generates protein side-chain conformations with the lowest energies jointly determined by the various energy functions. We further optimised the selected rotamers to construct subrotamer by rotamer minimisation, which reasonably improved the discreteness of the rotamer library. RESULTS: We focused on improving the accuracy of side-chain conformation prediction. For a testing set of 442 proteins, 87.19% of X1 and 77.11% of X12 angles were predicted correctly within 40° of the X-ray positions. We compared the accuracy of pacoPacker with state-of-the-art methods, such as CIS-RR and SCWRL4. We analysed the results from different perspectives, in terms of protein chain and individual residues. In this comprehensive benchmark testing, 51.5% of proteins within a length of 400 amino acids predicted by pacoPacker were superior to the results of CIS-RR and SCWRL4 simultaneously. Finally, we also showed the advantage of using the subrotamers strategy. All results confirmed that our parallel approach is competitive to state-of-the-art solutions for packing side chains. CONCLUSIONS: This parallel approach combines various sources of searching intelligence and energy functions to pack protein side chains. It provides a frame-work for combining different inaccuracy/usefulness objective functions by designing parallel heuristic search algorithms.
Lijun Quan, Haiou Li, Xiaoyan Xia, Hongjie Wu
BMC Bioinform.3
2013 A protein-peptide docking program with modeling receptor flexible areas
abstract
Predicting the structure of protein-peptide complexes using computational approaches is a difficult problem whose major challenges are properly dealing with molecular flexibility and conformational changes both of the receptor and ligand. Although significant improvements have been achieved in the modeling of side chains, methods for the backbone flexibility in docking still need improvement. In this study a new method is presented for docking peptide into receptor in a full flexible docking manner. It is a parallel approach that combines all the processes during the docking of a folding peptide with a flexible receptor.
Haiou Li, Lijun Quan, Xiaoyan Xia
BIBM1
2013 Packing protein side-chains by parallel ant colonies
abstract
Side-chains are crucial for proteins expressing their biochemical characteristics. Packing protein side-chains is then a necessary task for protein structure prediction, and critical to some descendant and important applications, such as protein design, docking and point mutation analysis. Given all possible candidate rotamers for each residue of protein backbone, packing protein side-chains can be modeled as a combinatorial optimization problem without an accurate energy function. This paper presents a parallel approach, pacoPacker, to pack protein side-chains by ant colony optimization. Each ant colony is used to pack side-chains with the guidance of an energy function. Different colonies use different energy functions. These multiple colonies are running in parallel and cooperate with each other by sharing the pheromone matrix whose role is to tune sampling the rotamer library. In this way, the intelligences embedded in different energy functions can be brought together to find out the best side-chains for the protein backbone. Experimental study has been conducted on two typical benchmarks, and the results show that pacoPacker is competitive to the state-of-art systems.
Lijun Quan, Haiou Li, Xiaoyan Xia
BIBM2