Jonathan D. Hirst

dblp:92/1356 · DBLP profile ↗
← Back
18ranked-venue papers
0as first author
3since 2021 · last 2026
0000-0002-2726-0983ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 3 since 2021Artificial intelligence and machine learning · 6Software engineering, systems software and programming languages · 3 · 3 since 2021
YearPublicationVenuePosition
2026 GDGraph: Geometry-Enhanced Dual-View Graph for Molecular Representation Learning
abstract
Learning effective molecular representations is crucial for accurate property prediction in AI-aided drug discovery. However, most existing molecular pre-training methods are still primarily based on 2D topological graphs, limiting their ability to exploit 3D geometric information. Moreover, methods that do incorporate 3D geometry often do not distinguish between the roles of atom-centered and bond-centered representations. To address these limitations, we propose GDGraph, a geometryenhanced dual-view framework for molecular representation learning. GDGraph models molecular geometry from two complementary structural perspectives: an atom view for capturing global spatial dependencies and a bond view for modeling local geometric patterns. To support this dual-view design, we introduce a multi-scale geometric feature encoding scheme and a view-specific geometry-aware learning strategy, enabling each view to focus on the geometric dependencies it is best suited to capture. Extensive experiments demonstrate that GDGraph achieves strong and stable performance on molecular property prediction benchmarks, and effectively predicts geometrysensitive quantum chemical properties on the QM9 dataset.
Yu Liu 0152, Jonathan D. Hirst, Jianfeng Ren, Bencan Tang, Dave Towey
COMPSAC2
2025 Chemically-aware Attention-based Multi-modal Fusion Framework for Molecular Representation Learning
abstract
Learning effective molecular representations is crucial for accurate property prediction in artificial intelligence (AI)-aided drug discovery. Graph and fingerprint representations have been widely used to encode molecular topological structures and chemical substructures. To enhance the feature embedding of each modality and leverage their complementary strengths, we propose a novel Chemically-aware Attention-based Multi-modal Fusion Framework (CAMFF) for molecular representation learning, which integrates molecular graphs and extended-connectivity fingerprints by exploiting various attention mechanisms. Specifically, the proposed CAMFF consists of three modules: 1) a graph embedding module incorporating multi-head attention to capture local heterogeneous interactions and all-pair self-attention to capture long-range atomic dependencies from molecular graph representations; 2) a fingerprint embedding module using a pre-trained Mol2Vec model to generate dense chemical substructure representations; and 3) a chemically-aware feature interaction and fusion module incorporating self-attention to enable interactions between various chemical substructures and cross-attention to ensure effective multi-modal alignment and fusion. To evaluate the effectiveness of CAMFF, we compare it with 14 state-of-the-art methods across 9 molecular property prediction benchmarks. CAMFF demonstrates competitive predictive performance and improves interpretability through attention-based visualization, showing its potential for real-world drug discovery.
Yu Liu 0152, Jonathan D. Hirst, Jianfeng Ren, Bencan Tang, Dave Towey
COMPSAC2
2024 Three-Branch Molecular Representation Learning Framework for Predicting Molecular Properties in Drug Discovery
abstract
Graph Neural Networks (GNNs) have been widely used to model molecules with a graph representation. However, GNNs face inherent challenges in accurately modeling long-range atomic interactions and identifying complex molecular substructures. This research proposes a novel Three-branch Molecular Representation Learning Framework (TMRLF) for predicting molecular properties: it integrates one branch of a GNN that extracts local molecular structural information with two branches of fully connected networks that capture the chemical substructure based on two fingerprints. Specifically, to better capture the long-range interactions, the GNN is designed with an attention mechanism to enhance the atomic interactions. As the Morgan fingerprint effectively captures functional groups of molecules and another well-used molecular fingerprint in the field of drug discovery, the Extended Reduced Graph (ErG) Fingerprint specifically targets molecular features with pharmacological relevance. These two fingerprints are both utilized to complement the chemical information and long-range information processing at the level of key structural features that GNNs lack. The proposed TMRLF extracts a robust feature representation of molecules, crucial for accurately predicting molecular properties and identifying potential drug candidates. Our proposed TMRLF is compared against six state-of-the-art models on eight benchmark datasets. It demonstrates superior capability in predicting molecular properties. Its effectiveness is further highlighted through proof-of-concept validation in identifying potential inhibitors for the Son of Sevenless Homolog 1 (SOSI) protein in real-world drug discovery scenarios.
Yu Liu 0152, Lihui Duo, Jonathan D. Hirst, Jianfeng Ren, Bencan Tang, Dave Towey
COMPSAC3
2010 Automatic structure classification of small proteins using random forest
abstract
BACKGROUND: Random forest, an ensemble based supervised machine learning algorithm, is used to predict the SCOP structural classification for a target structure, based on the similarity of its structural descriptors to those of a template structure with an equal number of secondary structure elements (SSEs). An initial assessment of random forest is carried out for domains consisting of three SSEs. The usability of random forest in classifying larger domains is demonstrated by applying it to domains consisting of four, five and six SSEs. RESULTS: Random forest, trained on SCOP version 1.69, achieves a predictive accuracy of up to 94% on an independent and non-overlapping test set derived from SCOP version 1.73. For classification to the SCOP Class, Fold, Super-family or Family levels, the predictive quality of the model in terms of Matthew's correlation coefficient (MCC) ranged from 0.61 to 0.83. As the number of constituent SSEs increases the MCC for classification to different structural levels decreases. CONCLUSIONS: The utility of random forest in classifying domains from the place-holder classes of SCOP to the true Class, Fold, Super-family or Family levels is demonstrated. Issues such as introduction of a new structural level in SCOP and the merger of singleton levels can also be addressed using random forest. A real-world scenario is mimicked by predicting the classification for those protein structures from the PDB, which are yet to be assigned to the SCOP classification hierarchy.
Pooja Jain, Jonathan D. Hirst
BMC Bioinform.2
2010 Predicting beta-turns and their types using predicted backbone dihedral angles and secondary structures
abstract
BACKGROUND: Beta-turns are secondary structure elements usually classified as coil. Their prediction is important, because of their role in protein folding and their frequent occurrence in protein chains. RESULTS: We have developed a novel method that predicts beta-turns and their types using information from multiple sequence alignments, predicted secondary structures and, for the first time, predicted dihedral angles. Our method uses support vector machines, a supervised classification technique, and is trained and tested on three established datasets of 426, 547 and 823 protein chains. We achieve a Matthews correlation coefficient of up to 0.49, when predicting the location of beta-turns, the highest reported value to date. Moreover, the additional dihedral information improves the prediction of beta-turn types I, II, IV, VIII and "non-specific", achieving correlation coefficients up to 0.39, 0.33, 0.27, 0.14 and 0.38, respectively. Our results are more accurate than other methods. CONCLUSIONS: We have created an accurate predictor of beta-turns and their types. Our method, called DEBT, is available online at http://comp.chem.nottingham.ac.uk/debt/.
Petros Kountouris, Jonathan D. Hirst
BMC Bioinform.2
2009 DichroCalc - circular and linear dichroism online
abstract
MOTIVATION: Circular dichroism (CD) is widely used in studies of protein folding. The CD spectrum of a protein can be estimated from its structure alone, using the well-established matrix method. In the last decade, a related spectroscopy, linear dichroism (LD), has been increasingly applied to study the orientation of proteins in solution. However, matrix method calculations of LD spectra have not been presented before. DichroCalc makes both CD and LD calculations available in an easy-to-use fashion. RESULTS: DichroCalc can be used without registration and calculates CD and LD spectra using a variety of matrix method parameters. PDB files can be uploaded as input or retrieved via their PDB code and a Perl-based parser is offered for easy handling of PDB files. AVAILABILITY: http://comp.chem.nottingham.ac.uk/dichrocalc and http://comp.chem.nottingham.ac.uk/parsepdb.
Benjamin M. Bulheller, Jonathan D. Hirst
Bioinform.2
2009 Automated Alphabet Reduction for Protein Datasets
abstract
BACKGROUND: We investigate automated and generic alphabet reduction techniques for protein structure prediction datasets. Reducing alphabet cardinality without losing key biochemical information opens the door to potentially faster machine learning, data mining and optimization applications in structural bioinformatics. Furthermore, reduced but informative alphabets often result in, e.g., more compact and human-friendly classification/clustering rules. In this paper we propose a robust and sophisticated alphabet reduction protocol based on mutual information and state-of-the-art optimization techniques. RESULTS: We applied this protocol to the prediction of two protein structural features: contact number and relative solvent accessibility. For both features we generated alphabets of two, three, four and five letters. The five-letter alphabets gave prediction accuracies statistically similar to that obtained using the full amino acid alphabet. Moreover, the automatically designed alphabets were compared against other reduced alphabets taken from the literature or human-designed, outperforming them. The differences between our alphabets and the alphabets taken from the literature were quantitatively analyzed. All the above process had been performed using a primary sequence representation of proteins. As a final experiment, we extrapolated the obtained five-letter alphabet to reduce a, much richer, protein representation based on evolutionary information for the prediction of the same two features. Again, the performance gap between the full representation and the reduced representation was small, showing that the results of our automated alphabet reduction protocol, even if they were obtained using a simple representation, are also able to capture the crucial information needed for state-of-the-art protein representations. CONCLUSION: Our automated alphabet reduction protocol generates competent reduced alphabets tailored specifically for a variety of protein datasets. This process is done without any domain knowledge, using information theory metrics instead. The reduced alphabets contain some unexpected (but sound) groups of amino acids, thus suggesting new ways of interpreting the data.
Jaume Bacardit, Michael Stout, Jonathan D. Hirst, Alfonso Valencia, Robert E. Smith 0001, Natalio Krasnogor
BMC Bioinform.3
2009 Prediction of backbone dihedral angles and protein secondary structure using support vector machines
abstract
BACKGROUND: The prediction of the secondary structure of a protein is a critical step in the prediction of its tertiary structure and, potentially, its function. Moreover, the backbone dihedral angles, highly correlated with secondary structures, provide crucial information about the local three-dimensional structure. RESULTS: We predict independently both the secondary structure and the backbone dihedral angles and combine the results in a loop to enhance each prediction reciprocally. Support vector machines, a state-of-the-art supervised classification technique, achieve secondary structure predictive accuracy of 80% on a non-redundant set of 513 proteins, significantly higher than other methods on the same dataset. The dihedral angle space is divided into a number of regions using two unsupervised clustering techniques in order to predict the region in which a new residue belongs. The performance of our method is comparable to, and in some cases more accurate than, other multi-class dihedral prediction methods. CONCLUSIONS: We have created an accurate predictor of backbone dihedral angles and secondary structure. Our method, called DISSPred, is available online at http://comp.chem.nottingham.ac.uk/disspred/.
Petros Kountouris, Jonathan D. Hirst
BMC Bioinform.2
2009 Prediction of topological contacts in proteins using learning classifier systems
Michael Stout, Jaume Bacardit, Jonathan D. Hirst, Robert E. Smith 0001, Natalio Krasnogor
Soft Comput.3
2008 Prediction of recursive convex hull class assignments for protein residues
abstract
MOTIVATION: We introduce a new method for designating the location of residues in folded protein structures based on the recursive convex hull (RCH) of a point set of atomic coordinates. The RCH can be calculated with an efficient and parameterless algorithm. RESULTS: We show that residue RCH class contains information complementary to widely studied measures such as solvent accessibility (SA), residue depth (RD) and to the distance of residues from the centroid of the chain, the residues' exposure (Exp). RCH is more conserved for related structures across folds and correlates better with changes in thermal stability of mutants than the other measures. Further, we assess the predictability of these measures using three types of machine-learning technique: decision trees (C4.5), Naive Bayes and Learning Classifier Systems (LCS) showing that RCH is more easily predicted than the other measures. As an exemplar application of predicted RCH class (in combination with other measures), we show that RCH is potentially helpful in improving prediction of residue contact numbers (CN).
Michael Stout, Jaume Bacardit, Jonathan D. Hirst, Natalio Krasnogor
Bioinform.3
2008 Prediction of glycosylation sites using random forests
abstract
BACKGROUND: Post translational modifications (PTMs) occur in the vast majority of proteins and are essential for function. Prediction of the sequence location of PTMs enhances the functional characterisation of proteins. Glycosylation is one type of PTM, and is implicated in protein folding, transport and function. RESULTS: We use the random forest algorithm and pairwise patterns to predict glycosylation sites. We identify pairwise patterns surrounding glycosylation sites and use an odds ratio to weight their propensity of association with modified residues. Our prediction program, GPP (glycosylation prediction program), predicts glycosylation sites with an accuracy of 90.8% for Ser sites, 92.0% for Thr sites and 92.8% for Asn sites. This is significantly better than current glycosylation predictors. We use the trepan algorithm to extract a set of comprehensible rules from GPP, which provide biological insight into all three major glycosylation types. CONCLUSION: We have created an accurate predictor of glycosylation sites and used this to extract comprehensible rules about the glycosylation process. GPP is available online at http://comp.chem.nottingham.ac.uk/glyco/.
Stephen E. Hamby, Jonathan D. Hirst
BMC Bioinform.2
2007 Automated alphabet reduction method with evolutionary algorithms for protein structure prediction
abstract
This paper focuses on automated procedures to reduce the dimensionality of protein structure prediction datasets by simplifying the way in which the primary sequence of a protein is represented. The potential benefits of this procedure are faster and easier learning process as well as the generation of more compact and human-readable classifiers. The dimensionality reduction procedure we propose consists on the reduction of the 20-letter amino acid (AA) alphabet, which is normally used to specify a protein sequence, into a lower cardinality alphabet. This reduction comes about by a clustering of AA types accordingly to their physical and chemical similarity. Our automated reduction procedure is guided by a fitness function based on the Mutual Information between the AA-based input attributes of the dataset and the protein structure feature that being predicted. To search for the optimal reduction, the Extended Compact Genetic Algorithm (ECGA) was used, and afterwards the results of this process were fed into (and validated by) BioHEL, a genetics-based machine learning technique. BioHEL used the reduced alphabet to induce rules for protein structure prediction features. BioHEL results are compared to two standard machine
Jaume Bacardit, Michael Stout, Jonathan D. Hirst, Kumara Sastry, Xavier Llorà, Natalio Krasnogor
GECCO3
2007 ProCKSI: a decision support system for Protein (Structure) Comparison, Knowledge, Similarity and Information
abstract
BACKGROUND: We introduce the decision support system for Protein (Structure) Comparison, Knowledge, Similarity and Information (ProCKSI). ProCKSI integrates various protein similarity measures through an easy to use interface that allows the comparison of multiple proteins simultaneously. It employs the Universal Similarity Metric (USM), the Maximum Contact Map Overlap (MaxCMO) of protein structures and other external methods such as the DaliLite and the TM-align methods, the Combinatorial Extension (CE) of the optimal path, and the FAST Align and Search Tool (FAST). Additionally, ProCKSI allows the user to upload a user-defined similarity matrix supplementing the methods mentioned, and computes a similarity consensus in order to provide a rich, integrated, multicriteria view of large datasets of protein structures. RESULTS: We present ProCKSI's architecture and workflow describing its intuitive user interface, and show its potential on three distinct test-cases. In the first case, ProCKSI is used to evaluate the results of a previous CASP competition, assessing the similarity of proposed models for given targets where the structures could have a large deviation from one another. To perform this type of comparison reliably, we introduce a new consensus method. The second study deals with the verification of a classification scheme for protein kinases, originally derived by sequence comparison by Hanks and Hunter, but here we use a consensus similarity measure based on structures. In the third experiment using the Rost and Sander dataset (RS126), we investigate how a combination of different sets of similarity measures influences the quality and performance of ProCKSI's new consensus measure. ProCKSI performs well with all three datasets, showing its potential for complex, simultaneous multi-method assessment of structural similarity in large protein datasets. Furthermore, combining different similarity measures is usually more robust than relying on one single, unique measure. CONCLUSION: Based on a diverse set of similarity measures, ProCKSI computes a consensus similarity profile for the entire protein set. All results can be clustered, visualised, analysed and easily compared with each other through a simple and intuitive interface.ProCKSI is publicly available at http://www.procksi.net for academic and non-commercial use.
Daniel Barthel, Jonathan D. Hirst, Jacek Blazewicz, Edmund K. Burke, Natalio Krasnogor
BMC Bioinform.2
2006 Coordination number prediction using learning classifier systems: performance and interpretability
abstract
The prediction of the coordination number (CN) of an amino acid in a protein structure has recently received renewed attention. In a recent paper, Kinjo et al. proposed a real-valued definition of CN and a criterion to map it onto a finite set of classes, in order to predict it using classification approaches. The literature reports several kinds of input information used for CN prediction. The aim of this paper is to assess the performance of a state-of-the-art learning method, Learning Classifier Systems (LCS) on this CN definition, with various degrees of precision, based on several combinations of input attributes. Moreover, we will compare the LCS performance to other well-known learning techniques. Our experiments are also intended to determinethe minimum set of input information needed to achieve good predictive performance, so as to generate competent yet simple and interpretable classification rules. Thus, the generated predictors (rule sets) are analyzed for their interpretability.
Jaume Bacardit, Michael Stout, Natalio Krasnogor, Jonathan D. Hirst, Jacek Blazewicz
GECCO4
2005 A fuzzy sets based generalization of contact maps for the overlap of protein structures
David A. Pelta, Natalio Krasnogor, Carlos Bousoño-Calzón, José L. Verdegay, Jonathan D. Hirst, Edmund K. Burke
Fuzzy Sets Syst.5
2004 Predicting protein secondary structure by cascade-correlation neural networks
abstract
Abstract Summary: The back-propagation neural network algorithm is a commonly used method for predicting the secondary structure of proteins. Whilst popular, this method can be slow to learn and here we compare it with an alternative: the cascade-correlation architecture. Using a constructive algorithm, cascade-correlation achieves predictive accuracies comparable to those obtained by back-propagation, in shorter time. Availability: A web server is available to a trained cascade-correlation neural network, with links to the source code for both algorithms.
Matthew J. Wood, Jonathan D. Hirst
Bioinform.2
2002 Alignment Of Protein Structures With A Memetic Evolutionary Algorithm
Robert D. Carr, William E. Hart, Natalio Krasnogor, Jonathan D. Hirst, Edmund K. Burke
GECCO4
2002 Multimeme Algorithms for Protein Structure Prediction
Natalio Krasnogor, B. P. Blackburne, Edmund K. Burke, Jonathan D. Hirst
PPSN4