John Parkinson

dblp:48/3354 · DBLP profile ↗
← Back
19ranked-venue papers
4as first author
4since 2021 · last 2026
0000-0001-9815-1189ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 18 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Correction: ToxoNet: A high confidence map of protein-protein interactions in Toxoplasma gondii
abstract
[This corrects the article DOI: 10.1371/journal.pcbi.1012208.].
Lakshmipuram S. Swapna, Grant C. Stevens, Aline Sardinha-Silva, Lucas Zhongming Hu, Verena Brand, Daniel D. Fusca, Cuihong Wan, Xuejian Xiong, Jon P. Boyle, Michael E. Grigg, Andrew Emili, John Parkinson
PLoS Comput. Biol.12
2024 Cell4D: a general purpose spatial stochastic simulator for cellular pathways
abstract
Abstract Background With the generation of vast compendia of biological datasets, the challenge is how best to interpret ‘omics data alongside biochemical and other small-scale experiments to gain meaningful biological insights. Key to this challenge are computational methods that enable domain-users to generate novel hypotheses that can be used to guide future experiments. Of particular interest are flexible modeling platforms, capable of simulating a diverse range of biological systems with low barriers of adoption to those with limited computational expertise. Results We introduce Cell4D, a spatial-temporal modeling platform combining a robust simulation engine with integrated graphics visualization, a model design editor, and an underlying XML data model capable of capturing a variety of cellular functions. Cell4D provides an interactive visualization mode, allowing intuitive feedback on model behavior and exploration of novel hypotheses, together with a non-graphics mode, compatible with high performance cloud compute solutions, to facilitate generation of statistical data. To demonstrate the flexibility and effectiveness of Cell4D, we investigate the dynamics of CEACAM1 localization in T-cell activation. We confirm the importance of Ca 2+ microdomains in activating calmodulin and highlight a key role of activated calmodulin on the surface expression of CEACAM1. We further show how lymphocyte-specific protein tyrosine kinase can help regulate this cell surface expression and exploit spatial modeling features of Cell4D to test the hypothesis that lipid rafts regulate clustering of CEACAM1 to promote trans-binding to neighbouring cells. Conclusions Through demonstrating its ability to test and generate hypotheses, Cell4D represents an effective tool to help integrate knowledge across diverse, large and small-scale datasets.
Donny Chan, Graham L. Cromar, Billy Taj, John Parkinson
BMC Bioinform.4
2024 ToxoNet: A high confidence map of protein-protein interactions in Toxoplasma gondii
abstract
The apicomplexan intracellular parasite Toxoplasma gondii is a major food borne pathogen that is highly prevalent in the global population. The majority of the T. gondii proteome remains uncharacterized and the organization of proteins into complexes is unclear. To overcome this knowledge gap, we used a biochemical fractionation strategy to predict interactions by correlation profiling. To overcome the deficit of high-quality training data in non-model organisms, we complemented a supervised machine learning strategy, with an unsupervised approach, based on similarity network fusion. The resulting combined high confidence network, ToxoNet, comprises 2,063 interactions connecting 652 proteins. Clustering identifies 93 protein complexes. We identified clusters enriched in mitochondrial machinery that include previously uncharacterized proteins that likely represent novel adaptations to oxidative phosphorylation. Furthermore, complexes enriched in proteins localized to secretory organelles and the inner membrane complex, predict additional novel components representing novel targets for detailed functional characterization. We present ToxoNet as a publicly available resource with the expectation that it will help drive future hypotheses within the research community.
Lakshmipuram S. Swapna, Grant C. Stevens, Aline Sardinha-Silva, Lucas Zhongming Hu, Verena Brand, Daniel D. Fusca, Cuihong Wan, Xuejian Xiong, Jon P. Boyle, Michael E. Grigg, Andrew Emili, John Parkinson
PLoS Comput. Biol.12
2022 Architect: A tool for aiding the reconstruction of high-quality metabolic models through improved enzyme annotation
abstract
Constraint-based modeling is a powerful framework for studying cellular metabolism, with applications ranging from predicting growth rates and optimizing production of high value metabolites to identifying enzymes in pathogens that may be targeted for therapeutic interventions. Results from modeling experiments can be affected at least in part by the quality of the metabolic models used. Reconstructing a metabolic network manually can produce a high-quality metabolic model but is a time-consuming task. At the same time, current methods for automating the process typically transfer metabolic function based on sequence similarity, a process known to produce many false positives. We created Architect, a pipeline for automatic metabolic model reconstruction from protein sequences. First, it performs enzyme annotation through an ensemble approach, whereby a likelihood score is computed for an EC prediction based on predictions from existing tools; for this step, our method shows both increased precision and recall compared to individual tools. Next, Architect uses these annotations to construct a high-quality metabolic network which is then gap-filled based on likelihood scores from the ensemble approach. The resulting metabolic model is output in SBML format, suitable for constraints-based analyses. Through comparisons of enzyme annotations and curated metabolic models, we demonstrate improved performance of Architect over other state-of-the-art tools, notably with higher precision and recall on the eukaryote C. elegans and when compared to UniProt annotations in two bacterial species. Code for Architect is available at https://github.com/ParkinsonLab/Architect. For ease-of-use, Architect can be readily set up and utilized using its Docker image, maintained on Docker Hub.
Nirvana Nursimulu, Alan M. Moses, John Parkinson
PLoS Comput. Biol.3
2020 A systematic pipeline for classifying bacterial operons reveals the evolutionary landscape of biofilm machineries
abstract
In bacteria functionally related genes comprising metabolic pathways and protein complexes are frequently encoded in operons and are widely conserved across phylogenetically diverse species. The evolution of these operon-encoded processes is affected by diverse mechanisms such as gene duplication, loss, rearrangement, and horizontal transfer. These mechanisms can result in functional diversification, increasing the potential evolution of novel biological pathways, and enabling pre-existing pathways to adapt to the requirements of particular environments. Despite the fundamental importance that these mechanisms play in bacterial environmental adaptation, a systematic approach for studying the evolution of operon organization is lacking. Herein, we present a novel method to study the evolution of operons based on phylogenetic clustering of operon-encoded protein families and genomic-proximity network visualizations of operon architectures. We applied this approach to study the evolution of the synthase dependent exopolysaccharide (EPS) biosynthetic systems: cellulose, acetylated cellulose, poly-β-1,6-N-acetyl-D-glucosamine (PNAG), Pel, and alginate. These polymers have important roles in biofilm formation, antibiotic tolerance, and as virulence factors in opportunistic pathogens. Our approach revealed the complex evolutionary landscape of EPS machineries, and enabled operons to be classified into evolutionarily distinct lineages. Cellulose operons show phyla-specific operon lineages resulting from gene loss, rearrangement, and the acquisition of accessory loci, and the occurrence of whole-operon duplications arising through horizonal gene transfer. Our evolution-based classification also distinguishes between PNAG production from Gram-negative and Gram-positive bacteria on the basis of structural and functional evolution of the acetylation modification domains shared by PgaB and IcaB loci, respectively. We also predict several pel-like operon lineages in Gram-positive bacteria and demonstrate in our companion paper (Whitfield et al PLoS Pathogens, in press) that Bacillus cereus produces a Pel-dependent biofilm that is regulated by cyclic-3',5'-dimeric guanosine monophosphate (c-di-GMP).
Cedoljub Bundalovic-Torma, Gregory B. Whitfield, Lindsey S. Marmont, P. Lynne Howell, John Parkinson
PLoS Comput. Biol.5
2018 Improved enzyme annotation with EC-specific cutoffs using DETECT v2
abstract
Summary: We present DETECT v2-an enzyme annotation tool which considers the effect of sequence diversity when assigning enzymatic function [as an Enzyme Commission (EC) number] to a protein sequence. In addition to capturing more enzyme classes than the previous version, we now provide EC-specific cutoffs that greatly increase precision and recall of assignments and show its performance in the context of pathways. Availability and implementation: https://github.com/ParkinsonLab/DETECT-v2. Supplementary information: Supplementary data are available at Bioinformatics online.
Nirvana Nursimulu, Leon L. Xu, James Wasmuth, Ivan Krukov, John Parkinson
Bioinform.5
2015 Hyperscape: visualization for complex biological networks
abstract
MOTIVATION: Network biology has emerged as a powerful tool to uncover the organizational properties of living systems through the application of graph theoretic approaches. However, due to limitations in underlying data models and visualization software, knowledge relating to large molecular assemblies and biologically active fragments is poorly represented. RESULTS: Here, we demonstrate a novel hypergraph implementation that better captures hierarchical structures, using components of elastic fibers and chromatin modification as models. These reveal unprecedented views of the biology of these systems, demonstrating the unique capacity of hypergraphs to resolve overlaps and uncover new insights into the subfunctionalization of variant complexes. AVAILABILITY AND IMPLEMENTATION: Hyperscape is available as a web application at http://www.compsysbio.org/hyperscape. Source code, examples and a tutorial are freely available under a GNU license. CONTACTS: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Graham L. Cromar, Anthony Zhao, Alex Yang, John Parkinson
Bioinform.4
2013 Identification of a Functional Connectome for Long-Term Fear Memory in Mice
abstract
Long-term memories are thought to depend upon the coordinated activation of a broad network of cortical and subcortical brain regions. However, the distributed nature of this representation has made it challenging to define the neural elements of the memory trace, and lesion and electrophysiological approaches provide only a narrow window into what is appreciated a much more global network. Here we used a global mapping approach to identify networks of brain regions activated following recall of long-term fear memories in mice. Analysis of Fos expression across 84 brain regions allowed us to identify regions that were co-active following memory recall. These analyses revealed that the functional organization of long-term fear memories depends on memory age and is altered in mutant mice that exhibit premature forgetting. Most importantly, these analyses indicate that long-term memory recall engages a network that has a distinct thalamic-hippocampal-cortical signature. This network is concurrently integrated and segregated and therefore has small-world properties, and contains hub-like regions in the prefrontal cortex and thalamus that may play privileged roles in memory expression.
Anne L. Wheeler, Cátia M. Teixeira, Afra H. Wang, Xuejian Xiong, Natasa Kovacevic, Jason P. Lerch, Anthony R. McIntosh, John Parkinson, Paul W. Frankland
PLoS Comput. Biol.8
2012 Modelling the Self-Assembly of Elastomeric Proteins Provides Insights into the Evolution of Their Domain Architectures
abstract
Elastomeric proteins have evolved independently multiple times through evolution. Produced as monomers, they self-assemble into polymeric structures that impart properties of stretch and recoil. They are composed of an alternating domain architecture of elastomeric domains interspersed with cross-linking elements. While the former provide the elasticity as well as help drive the assembly process, the latter serve to stabilise the polymer. Changes in the number and arrangement of the elastomeric and cross-linking regions have been shown to significantly impact their assembly and mechanical properties. However, to date, such studies are relatively limited. Here we present a theoretical study that examines the impact of domain architecture on polymer assembly and integrity. At the core of this study is a novel simulation environment that uses a model of diffusion limited aggregation to simulate the self-assembly of rod-like particles with alternating domain architectures. Applying the model to different domain architectures, we generate a variety of aggregates which are subsequently analysed by graph-theoretic metrics to predict their structural integrity. Our results show that the relative length and number of elastomeric and cross-linking domains can significantly impact the morphology and structural integrity of the resultant polymeric structure. For example, the most highly connected polymers were those constructed from asymmetric rods consisting of relatively large cross-linking elements interspersed with smaller elastomeric domains. In addition to providing insights into the evolution of elastomeric proteins, simulations such as those presented here may prove valuable for the tuneable design of new molecules that may be exploited as useful biomaterials.
Hongyan Song, John Parkinson
PLoS Comput. Biol.2
2011 PhyloPro: a web-based tool for the generation and visualization of phylogenetic profiles across Eukarya
abstract
SUMMARY: With increasing numbers of eukaryotic genome sequences, phylogenetic profiles of eukaryotic genes are becoming increasingly informative. Here, we introduce a new web-tool Phylopro (http://compsysbio.org/phylopro/), which uses the 120 available eukaryotic genome sequences to visualize the evolutionary trajectories of user-defined subsets of model organism genes. Applied to pathways or complexes, PhyloPro allows the user to rapidly identify core conserved elements of biological processes together with those that may represent lineage-specific innovations. PhyloPro thus provides a valuable resource for the evolutionary and comparative studies of biological systems.
Xuejian Xiong, Hongyan Song, Tuan On, Lucas Lochovsky, Nicholas J. Provart, John Parkinson
Bioinform.6
2010 DETECT - a Density Estimation Tool for Enzyme ClassificaTion and its application to Plasmodium falciparum
abstract
MOTIVATION: A major challenge in genomics is the accurate annotation of component genes. Enzymes are typically predicted using homology-based search methods, where the membership of a protein to an enzyme family is based on single-sequence comparisons. As such, these methods are often error-prone and lack useful measures of reliability for the prediction. RESULTS: Here, we present DETECT, a probabilistic method for enzyme prediction that accounts for the sequence diversity across enzyme families. By comparing the global alignment scores of an unknown protein to those of all known enzymes, an integrated likelihood score can be readily calculated, ranking the reaction classes relevant for that protein. Comparisons to BLAST reveal significant improvements in enzyme annotation accuracy. Applied to Plasmodium falciparum, we identify potential annotation errors and predict novel enzymes of therapeutic interest. AVAILABILITY: A standalone application is available from the website: http://www.compsysbio.org/projects/DETECT/
Stacy S. Hung, James Wasmuth, Christopher Sanford, John Parkinson
Bioinform.4
2009 The Modular Organization of Protein Interactions in Escherichia coli
abstract
Escherichia coli serves as an excellent model for the study of fundamental cellular processes such as metabolism, signalling and gene expression. Understanding the function and organization of proteins within these processes is an important step towards a 'systems' view of E. coli. Integrating experimental and computational interaction data, we present a reliable network of 3,989 functional interactions between 1,941 E. coli proteins ( approximately 45% of its proteome). These were combined with a recently generated set of 3,888 high-quality physical interactions between 918 proteins and clustered to reveal 316 discrete modules. In addition to known protein complexes (e.g., RNA and DNA polymerases), we identified modules that represent biochemical pathways (e.g., nitrate regulation and cell wall biosynthesis) as well as batteries of functionally and evolutionarily related processes. To aid the interpretation of modular relationships, several case examples are presented, including both well characterized and novel biochemical systems. Together these data provide a global view of the modular organization of the E. coli proteome and yield unique insights into structural and evolutionary relationships in bacterial networks.
José M. Peregrín-Alvarez, Xuejian Xiong, Chong Su, John Parkinson
PLoS Comput. Biol.4
2008 SubSeqer: a graph-based approach for the detection and identification of repetitive elements in low-complexity sequences
abstract
Low-complexity, repetitive protein sequences with a limited amino acid palette are abundant in nature, and many of them play an important role in the structure and function of certain types of proteins. However, such repetitive sequences often do not have rigidly defined motifs. Consequently, the identification of these low-complexity repetitive elements has proven challenging for existing pattern-matching algorithms. Here we introduce a new web-tool SubSeqer (http://compsysbio.org/subseqer/) which uses graphical visualization methods borrowed from protein interaction studies to identify and characterize repetitive elements in low-complexity sequences. Given their abundance, we suggest that SubSeqer represents a valuable resource for the study of typically neglected low-complexity sequences.
David He, John Parkinson
Bioinform.2
2006 Cell++ - simulating biochemical pathways
abstract
MOTIVATION: With the generation of a wealth of information, detailing cellular components, their functions and interactions, there is a growing need for the development of new computational tools capable of interpreting these data within spatial and dynamic contexts. Here, we introduce Cell++, a novel stochastic simulation environment with the capacity to study a wide variety of biochemical processes within a spatial context. RESULTS: Focusing on three case studies, we highlight the potential impact of spatial organization in the evolution and engineering of signaling and metabolic pathways. In addition to altering signaling and metabolic efficiency, simulations also demonstrated features consistent with the phenomenon of metabolic channeling. AVAILABILITY: Cell++ is licensed under the GNU general public license (GPL) and has been successfully implemented under Linux and IRIX operating systems. Source code together with a simple tutorial is available at http://www.compsysbio.org/CellSim/.
Christopher Sanford, Matthew L. K. Yip, Carl White, John Parkinson
Bioinform.4
2005 Exploring Parasite Gene Space
James Wasmuth, Ralf Schmid, Alasdair Anthony, John Parkinson, Mark L. Blaxter
BMC Bioinform.4
2004 PartiGene-constructing partial genomes
abstract
UNLABELLED: Expressed sequence tags (ESTs) offer a low-cost approach to gene discovery and are being used by an increasing number of laboratories to obtain sequence information for a wide variety of organisms. The challenge lies in processing and organizing this data within a genomic context to facilitate large scale analyses. Here we present PartiGene, an integrated sequence analysis suite that uses freely available public domain software to (1) process raw trace chromatograms into sequence objects suitable for submission to dbEST; (2) place these sequences within a genomic context; (3) perform customizable first-pass annotation of the data; and (4) present the data as HTML tables and an SQL database resource. PartiGene has been used to create a number of non-model organism database resources including NEMBASE (http://www.nematodes.org) and LumbriBase (http://www.earthworms.org/). The packages are readily portable, freely available and can be run on simple Linux-based workstations. AVAILABILITY: PartiGene is available from http://www.nematodes.org/PartiGene and also forms part of the EST analysis software, associated with the Natural Environmental Research Council (UK) Bio-Linux project (http://envgen.nox.ac.uk/biolinux.html).
John Parkinson, Alasdair Anthony, James Wasmuth, Ralf Schmid, Ann Hedley, Mark L. Blaxter
Bioinform.1
2003 SimiTri-visualizing similarity relationships for groups of sequences
abstract
Global sequence comparisons between large datasets, such as those arising from genome projects, can be problematic to display and analyze. We have developed SimiTri, a Java/Perl-based application, which allows simultaneous display and analysis of relative similarity relationships of the dataset of interest to three different databases. We illustrate its utility in identifying Caenorhabditis elegans genes that have distinct patterns of phylogenetic affinity suggestive of horizontal gene transfer. SimiTri is freely downloadable from http://www.nematodes.org/SimiTri/ and the source code is freely available from the authors.
John Parkinson, Mark L. Blaxter
Bioinform.1
2002 Making sense of EST sequences by CLOBBing them
abstract
BACKGROUND: Expressed sequence tags (ESTs) are single pass reads from randomly selected cDNA clones. They provide a highly cost-effective method to access and identify expressed genes. However, they are often prone to sequencing errors and typically define incomplete transcripts. To increase the amount of information obtainable from ESTs and reduce sequencing errors, it is necessary to cluster ESTs into groups sharing significant sequence similarity. RESULTS: As part of our ongoing EST programs investigating 'orphan' genomes, we have developed a clustering algorithm, CLOBB (Cluster on the basis of BLAST similarity) to identify and cluster ESTs. CLOBB may be used incrementally, preserving original cluster designations. It tracks cluster-specific events such as merging, identifies 'superclusters' of related clusters and avoids the expansion of chimeric clusters. Based on the Perl scripting language, CLOBB is highly portable relying only on a local installation of NCBI's freely available BLAST executable and can be usefully applied to > 95 % of the current EST datasets. Analysis of the Danio rerio EST dataset demonstrates that CLOBB compares favourably with two less portable systems, UniGene and TIGR Gene Indices. CONCLUSIONS: CLOBB provides a highly portable EST clustering solution and is freely downloaded from: http://www.nematodes.org/CLOBB
John Parkinson, David B. Guiliano, Mark L. Blaxter
BMC Bioinform.1
1990 Making CASE Work
John Parkinson
CAiSE1