VLDB 2026 Research / reviewers in the wild / expert
James R. Green
dblp:20/1472 · also James Green 0001, James Robert Green
· DBLP profile ↗
25ranked-venue papers
1as first author
14since 2021 · last 2025
0000-0002-6039-2355ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Galileo: Learning Global & Local Features of Many Remote Sensing ModalitiesabstractWe introduce a highly multimodal transformer to represent many remote sensing modalities - multispectral optical, synthetic aperture radar, elevation, weather, pseudo-labels, and more - across space and time. These inputs are useful for diverse remote sensing tasks, such as crop mapping and flood detection. However, learning shared representations of remote sensing data is challenging, given the diversity of relevant data modalities, and because objects of interest vary massively in scale, from small boats (1-2 pixels and fast) to glaciers (thousands of pixels and slow). We present a novel self-supervised learning algorithm that extracts multi-scale features across a flexible set of input modalities through masked modeling. Our dual global and local contrastive losses differ in their targets (deep representations vs. shallow input projections) and masking strategies (structured vs. not). Our Galileo is a single generalist model that outperforms SoTA specialist models for satellite images and pixel time series across eleven benchmarks and multiple tasks. Gabriel Tseng, Anthony Fuller, Marlena Reil, Henry Herzog, Patrick Beukema, Favyen Bastani, James R. Green, Evan Shelhamer, Hannah Kerner, David Rolnick |
ICML | 7 |
| 2025 | LineShield - A Generalized LiDAR Pipeline for Automated Vegetation Encroachment Detection on PowerlinesabstractVegetation encroachment on the powerline poses significant risks to the reliability and safety of the power infrastructure. Many LiDAR-Based methods have tried to address this problem, yet these methods lack scalability and generality across different LiDAR collection methods with varying resolutions and data sizes. This paper presents LineShield, a generalized LiDAR-based pipeline for automated vegetation encroachment detection on powerlines. The pipeline addresses these challenges on both airborne and mobile LiDAR datasets. Key components: clustering with DBSCAN for accurate powerline segmentation, PCA-based alignment for standardizing orientation, sliding window traversal for efficient processing of large datasets, voxel downsampling for reducing data complexity, and proximity-based severity classification for prioritizing interventions. Experimental results demonstrate the pipeline’s adaptability across three datasets, namely, ECLAIR, DALES, and Toronto-3D, achieving up to 97.4% detection accuracy and maintaining performance in both urban and rural environments. The approach also provides detailed reporting to help maintenance teams prioritize interventions. Aziz Al-Najjar, Marzieh Amini, James R. Green, Felix Kwamena |
ISCAS | 3 |
| 2025 | LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-SupervisionabstractVision transformers are ever larger, more accurate, and more expensive to compute.
At high resolution, the expense is even more extreme as the number of tokens grows quadratically in the image size.
We turn to adaptive computation to cope with this cost by learning to predict where to compute.
Our LookWhere method divides the computation between a low-resolution selector and a high-resolution extractor without ever processing the full high-resolution input.
We jointly pretrain the selector and extractor without task supervision by distillation from a self-supervised teacher, in effect learning where and what to compute at the same time.
Unlike prior token reduction methods, which pay to save by pruning already-computed tokens, and prior token selection methods, which
require complex and expensive per-task optimization, LookWhere economically and accurately selects and extracts transferrable representations of images.
We show that LookWhere excels at sparse recognition on high-resolution inputs (Traffic Signs), maintaining accuracy while reducing FLOPs by 17x and time by 4x, and standard recognition tasks that are global (ImageNet classification) and local (ADE20K segmentation), improving accuracy while reducing time by 1.36x. Anthony Fuller, Yousef Yassin, Junfeng Wen, Tarek Ibrahim, Daniel G. Kyrollos, James R. Green, Evan Shelhamer |
NeurIPS | 6 |
| 2025 | Soft Contrastive Representation Learning for Cloud-Particle Images Captured In-Flight by the New HVPS-4 Airborne ProbeabstractCloud properties underpin accurate climate modeling and are often derived from the individual particles comprising a cloud. Studying these cloud particles is challenging due to their intricate shapes, called “habits,” and manual classification via probe-generated images is time-consuming and subjective. We propose a novel method for habit representation learning that uses minimal labeled data by leveraging self-supervised learning (SSL) with Vision Transformers (ViTs) on a newly acquired dataset of 124000 images captured by the novel high-volume precipitation spectrometer ver. 4 (HVPS-4) probe. Our approach significantly outperforms ImageNet pretraining by 48% on a 293-sample annotated dataset. Notably, we present the first SSL scheme for learning habit representations, leveraging data collected in flight from the probe. Our results demonstrate that self-supervised pretraining significantly improves habit classification even when using single-channel HVPS-4 data. We achieve further gains using sequential views and a soft contrastive objective tailored for sequential, in-flight measurements. Our work paves the way for applying SSL to multiview and multiscale data from advanced cloud-particle imaging probes, enabling comprehensive characterization of the flight environment. We publicly release data, code, and models associated with this study. Yousef Yassin, Anthony Fuller, Keyvan Ranjbar, Kenny Bala, Leonid Nichman, James R. Green |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2025 | Function approximations valid in both time and frequency domains using legendre moments
Hamid Reza Aghamiri, James R. Green, B. John Oommen |
Pattern Anal. Appl. | 2 |
| 2024 | LookHere: Vision Transformers with Directed Attention Generalize and ExtrapolateabstractHigh-resolution images offer more information about scenes that can improve model accuracy. However, the dominant model architecture in computer vision, the vision transformer (ViT), cannot effectively leverage larger images without finetuning — ViTs poorly extrapolate to more patches at test time, although transformers offer sequence length flexibility. We attribute this shortcoming to the current patch position encoding methods, which create a distribution shift when extrapolating.
We propose a drop-in replacement for the position encoding of plain ViTs that restricts attention heads to fixed fields of view, pointed in different directions, using 2D attention masks. Our novel method, called LookHere, provides translation-equivariance, ensures attention head diversity, and limits the distribution shift that attention heads face when extrapolating. We demonstrate that LookHere improves performance on classification (avg. 1.6%), against adversarial attack (avg. 5.4%), and decreases calibration error (avg. 1.5%) — on ImageNet without extrapolation. With extrapolation, LookHere outperforms the current SoTA position encoding method, 2D-RoPE, by 21.7% on ImageNet when trained at $224^2$ px and tested at $1024^2$ px. Additionally, we release a high-resolution test set to improve the evaluation of high-resolution image classifiers, called ImageNet-HR. Anthony Fuller, Daniel G. Kyrollos, Yousef Yassin, James R. Green |
NeurIPS | 4 |
| 2023 | Domain Adaptation Applied to microRNA Target PredictionabstractMicroRNA target prediction attempts to determine which genes are regulated by which miRNA. When developing such predictors, we are often faced with data scarcity in the species of interest. In this paper we explore the use of domain adaptation as a means to intelligently increase the available training data by effective pooling of data from multiple training species. Through a variety of experiments, we discovered that, on average, domain adaptation improves the performance of miRNA target predictors. Significant performance gains can be found in inter-kingdom experiments where an animal training set is augmented with plant data to perform target prediction on a plant species, or vice versa. Additionally, the inclusion of position data in cases where the target species is animal significantly improved performance overall. Taken together, these advancements can lead to improved species-specific miRNA target prediction. Victoria Ajila, James R. Green |
BIBE | 2 |
| 2023 | Data Augmentation and Deep Learning in Audio Classification Problems: Alignment Between Training and Test EnvironmentsabstractThe global COVID-19 pandemic increased the interest in automatic analysis and classification of cough sounds. However, developing such a system requires a large dataset of expert-labelled cough sounds, which remains elusive. Data augmentation techniques are often employed to train machine learning models with limited data, while ensuring model robustness to real-world variations. We have recently proposed the use of “natural” spectral data augmentation methods, including noise and reverberation [1]. In this paper, we further investigate the method by studying the alignment between training and testing environments. We augment training sound data with varying levels of reverberation and Gaussian noise, and evaluate augmented convolutional neural network models across a range of test environments on two audio classification tasks: speech command recognition and cough sound classification. Results demonstrate the broad robustness of the data-augmented deep learning model across various test environments. Furthermore, the augmented model's performance on cough classification is compared to 4 expert annotations on a large cough dataset and is seen to be near-human-level accuracy. This augmented model is recommended for classifying cough sounds in natural settings beyond laboratory conditions. Saiful Huq, Pengcheng Xi, Rafik A. Goubran, Frank Knoefel, James R. Green |
BIBE | 5 |
| 2023 | CROMA: Remote Sensing Representations with Contrastive Radar-Optical Masked AutoencodersabstractA vital and rapidly growing application, remote sensing offers vast yet sparsely labeled, spatially aligned multimodal data; this makes self-supervised learning algorithms invaluable. We present CROMA: a framework that combines contrastive and reconstruction self-supervised objectives to learn rich unimodal and multimodal representations. Our method separately encodes masked-out multispectral optical and synthetic aperture radar samples—aligned in space and time—and performs cross-modal contrastive learning. Another encoder fuses these sensors, producing joint multimodal encodings that are used to predict the masked patches via a lightweight decoder. We show that these objectives are complementary when leveraged on spatially aligned multimodal data. We also introduce X- and 2D-ALiBi, which spatially biases our cross- and self-attention matrices. These strategies improve representations and allow our models to effectively extrapolate to images up to $17.6\times$ larger at test-time. CROMA outperforms the current SoTA multispectral model, evaluated on: four classification benchmarks—finetuning (avg.$\uparrow$ 1.8%), linear (avg.$\uparrow$ 2.4%) and nonlinear (avg.$\uparrow$ 1.4%) probing, $k$NN classification (avg.$\uparrow$ 3.5%), and $K$-means clustering (avg.$\uparrow$ 8.4%); and three segmentation benchmarks (avg.$\uparrow$ 6.4%). CROMA’s rich, optionally multimodal representations can be widely leveraged across remote sensing applications. Anthony Fuller, Koreen Millard, James R. Green |
NeurIPS | 3 |
| 2022 | Depth Encoding for Neonatal Patient SegmentationabstractPatient segmentation is an important step in neonatal monitoring for subsequent applications including vital sign monitoring, jaundice detection, or clonic seizure detection. Many studies have faced difficulty in obtaining clear delineation between the patient and the background, instead settling for segmentation of specific body parts, segmentation of visible skin only, or manual selection of regions of interest. The outline of the full body, however, provides a more holistic detection of the entire patient which can be useful for various applications. This study investigates whole-body semantic segmentation of patients under varying levels of coverage (unclothed, clothed, and partially covered with blankets), in complex scenes from the neonatal intensive care unit, from patients placed across all bed types (crib, incubator, and overhead warmer), and in different subject poses (supine, prone, and fetal position). To improve the patient-background delineation, several depth encoding and RGB-D fusion techniques are investigated with a Mask R-CNN model. This study demonstrates how depth information can effectively enhance the RGB image for neonatal patient segmentation, especially when the delineation is unclear due to patient coverage. While RGB images provided suitable predictions for optimal “clear” contours, RGB-D fusion images were required to achieve accurate patient segmentation in challenging “unclear” contours with over 71% IOU and over 82% Dice metrics. Yasmina Souley Dosso, Kim Greenwood, JoAnn Harrold, James R. Green |
BIBM | 4 |
| 2022 | Emergence of an Autonomous Vehicle Secondary Data Market for Breakthrough ApplicationsabstractThe prophesied circulation of fleets of autonomous vehicles (AVs) in urban and rural environments promises unprecedented opportunities to remotely sense streetscapes at fine-grain spatial and temporal resolution. AVs employ a variety of on-board sensors to capture information about the local environs for the primary purpose of vehicular navigation. However, we propose that these data may find further secondary use in a broad array of breakthrough applications: technologies and use cases that are enabled through the fine-grain spatio-temporal sensing of the lived environment. Consequently, a market for the secondary use of AV-collected data is emergent and a cloud-based architecture to manage the collection, processing, and communication of AV-derived data is required. Excitingly, the application of machine learning models to extract desirable secondary information from these fine-grain spatio-temporal data will enable unprecedented global-scale and time-series studies. Herein, we outline our vision for the utility of a Remote sensing AV-based Informatics Layer (RAIL) and the breakthrough applications it would enable. We define our vision based on recent and relevant trends in AV technology, discuss anticipated applications, discuss key technical considerations, and explore theoretical economic models for the exposed API. We conclude with discussion of the socio-technical ramifications of this system. Kevin Dick, James R. Green |
IEEE Big Data | 2 |
| 2022 | SatViT: Pretraining Transformers for Earth ObservationabstractDespite the enormous success of the ’pre-training and fine-tuning’ paradigm, widespread across machine learning, it has yet to pervade remote sensing (RS). To help rectify this, we pre-train a vision transformer (ViT) on 1.3 million satellite-derived RS images. We pre-train SatViT using a state-of-the-art self-supervised learning algorithm called masked autoencoding (MAE), which learns general representations by reconstructing held-out image patches. Crucially, this approach does not require annotated data, allowing us to pre-train on unlabeled images acquired from Sentinel-1 & 2. After fine-tuning, SatViT outperforms state-of-the-art ImageNet and RS-specific pre-trained models on both of our downstream tasks. We further improve overall accuracy (by 3.2% and 0.21%) by continuing to pre-train SatViT—still using MAE—on the unlabelled target datasets. Most importantly, we release our code, pre-trained model weights, and tutorials aimed at helping researchers fine-tune our models. (https://github.com/antofuller/SatViT). Anthony Fuller, Koreen Millard, James R. Green |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2021 | A novel Greedy approach for Sequence based Computational prediction of Binding-Sites in Protein-Protein InteractionabstractComputational prediction of protein-protein interaction (PPI) from protein sequence is important as many cellular functions are made possible through PPI. The Protein Interaction Prediction Engine (PIPE) software suite was developed for such predictions. The specific location of interaction is predicted by the PIPE-Sites predictor, which depends on PIPE engine. This PIPE-Sites predictor is here updated through the use of a large high-quality dataset of known PPI sites. Additionally, a similarity-weighted score had been recently developed in PIPE4 and has been proven to be more accurate for the likelihood of PPI prediction. However, PIPE-Sites are shown to be ineffective when applied to similarity-weighted score data. Thus, we here propose and evaluate a new sequence-based PPI site prediction method, named Panorama. This new method leverages similarity-weighted score data to further increase performance over two different performance metrics when evaluated on both$\boldsymbol{H}$. sapiens and$\boldsymbol{S}$, cerevisiae PPI site data. Aishwarya Purohit, Shrinivas Acharya, James R. Green |
BIBE | 3 |
| 2021 | MetaHate: A Meta-Model for Hate Speech DetectionabstractWe present MetaHate, a NLP meta-model for detecting hatefulness in tweets by combining predictors for hate, emotion, sentiment, and offensiveness. We evaluate this model with the TweetEval benchmark for hate speech detection. MetaHate improves the baseline TweetEval RoBERTa based model on the TweetEval benchmark. Optimizing the decision threshold for the macro-averaged F1-score, MetaHate achieves a F1-score of 0.70, while the TweetEval RoBERTa-Twitter Retrained Hate model achieves a F1-score of 0.63. This improvement on one of the most difficult tasks on the TweetEval benchmark was achieved with no additional training data and negligible computational time and cost. MetaHate demonstrates the utility of leveraging predictions from language models trained for various tasks to improve performance on a single task. Daniel G. Kyrollos, James R. Green |
IEEE BigData | 2 |
| 2020 | Chaos Game Representations & Deep Learning for Proteome-Wide Protein PredictionabstractChaos Game Representation (CGR) is an emerging means of visualising and representing genomic and proteomic sequences. There exist many open questions related to its effective application to various computational tasks. In this work, we begin to address some of these questions by comparing four variants of the Chaos Game to generate CGR imagery as part of a multi-class classification task to identify the source organism for a given protein. We propose a novel nodal configuration for icosagon and 20-flake CGRs. Using two datasets, we performed fine-tuning using seven deep convolutional neural network (CNN) architectures and report modest performance over random among the 56 test conditions, highlighting certain shortcomings in effectively leveraging CGR in conjunction with deep CNN architectures. Many of the insights from this work will serve to orient subsequent protein-related studies involving CGR-based encoding and be generally applicable to disparate domains seeking to leverage CGR for sequence-type data. Kevin Dick, James R. Green |
BIBE | 2 |
| 2018 | Active Learning for microRNA Prediction
Mohsen Sheikh Hassani, James R. Green |
BIBM | 2 |
| 2018 | A review of network-based approaches to drug repositioningabstractExperimental drug development is time-consuming, expensive and limited to a relatively small number of targets. However, recent studies show that repositioning of existing drugs can function more efficiently than de novo experimental drug development to minimize costs and risks. Previous studies have proven that network analysis is a versatile platform for this purpose, as the biological networks are used to model interactions between many different biological concepts. The present study is an attempt to review network-based methods in predicting drug targets for drug repositioning. For each method, the preferred type of data set is described, and their advantages and limitations are discussed. For each method, we seek to provide a brief description, as well as an evaluation based on its performance metrics.We conclude that integrating distinct and complementary data should be used because each type of data set reveals a unique aspect of information about an organism. We also suggest that applying a standard set of evaluation metrics and data sets would be essential in this fast-growing research domain. Maryam Lotfi Shahreza, Nasser Ghadiri, Seyed Rasoul Mousavi, Jaleh Varshosaz, James R. Green |
Briefings Bioinform. | 5 |
| 2017 | Positome: A method for improving protein-protein interaction quality and prediction accuracyabstractThe progressive elucidation of positive protein-protein interactions (PPIs) as wet-lab techniques continue to improve in both throughput and precision has increased the number and quality of known PPIs across the spectrum of life. Creating high quality datasets of positive PPIs is critical for training PPI prediction algorithms and for assessing the performance of PPI detection efforts. We present the Positome, a web service to acquire sets of positive PPIs based on user-defined criteria pertaining to data provenance including interaction type, throughput level, and detection method selection in addition to filtration by multiple lines of evidence (i.e. PPIs reported by independent research groups). The Positome provides a tunable interface to obtain a specified subset of interacting PPIs from the BioGRlD database. Both intra- and inter-species PPIs are supported. Using a number of model organisms, we demonstrate the trade-off between data quality and quantity, and the benefit of higher data quality on PPI prediction precision and recall. A web interface and REST web service are available at http://bioinf.sce.carleton.ca/POSITOME/. Kevin Dick, Frank Dehne, Ashkan Golshani, James R. Green |
CIBCB | 4 |
| 2017 | Exploring general-purpose protein features for distinguishing enzymes and non-enzymes within the twilight zoneabstractComputational prediction of protein function constitutes one of the more complex problems in Bioinformatics, because of the diversity of functions and mechanisms in that proteins exert in nature. This issue is reinforced especially for proteins that share very low primary or tertiary structure similarity to existing annotated proteomes. In this sense, new alignment-free (AF) tools are needed to overcome the inherent limitations of classic alignment-based approaches to this issue. We have recently introduced AF protein-numerical-encoding programs (TI2BioP and ProtDCal), whose sequence-based features have been successfully applied to detect remote protein homologs, post-translational modifications and antibacterial peptides. Here we aim to demonstrate the applicability of 4 AF protein descriptor families, implemented in our programs, for the identification enzyme-like proteins. At the same time, the use of our novel family of 3D–structure-based descriptors is introduced for the first time. The Dobson & Doig (D&D) benchmark dataset is used for the evaluation of our AF protein descriptors, because of its proven structural diversity that permits one to emulate an experiment within the twilight zone of alignment-based methods (pair-wise identity <30%). The performance of our sequence-based predictor was further assessed using a subset of formerly uncharacterized proteins which currently represent a benchmark annotation dataset. Four protein descriptor families (sequence-composition-based (0D), linear-topology-based (1D), pseudo-fold-topology-based (2D) and 3D–structure features (3D), were assessed using the D&D benchmark dataset. We show that only the families of ProtDCal’s descriptors (0D, 1D and 3D) encode significant information for enzymes and non-enzymes discrimination. The obtained 3D–structure-based classifier ranked first among several other SVM-based methods assessed in this dataset. Furthermore, the model leveraging 1D descriptors, showed a higher success rate than EzyPred on a benchmark annotation dataset from the Shewanella oneidensis proteome. The applicability of ProtDCal as a general-purpose-AF protein modelling method is illustrated through the discrimination between two comprehensive protein functional classes. The observed performances using the highly diverse D&D dataset, and the set of formerly uncharacterized (hard-to-annotate) proteins of Shewanella oneidensis , places our methodology on the top range of methods to model and predict protein function using alignment-free approaches. Yasser B. Ruiz-Blanco, Guillermín Agüero-Chapín, Enrique García-Hernández, Orlando Álvarez, Agostinho Antunes, James R. Green |
BMC Bioinform. | 6 |
| 2017 | Heter-LP: A heterogeneous label propagation algorithm and its application in drug repositioning
Maryam Lotfi Shahreza, Nasser Ghadiri, Seyed Rasoul Mousavi, Jaleh Varshosaz, James R. Green |
J. Biomed. Informatics | 5 |
| 2015 | ProtDCal: A program to compute general-purpose-numerical descriptors for sequences and 3D-structures of proteinsabstractBACKGROUND: The exponential growth of protein structural and sequence databases is enabling multifaceted approaches to understanding the long sought sequence-structure-function relationship. Advances in computation now make it possible to apply well-established data mining and pattern recognition techniques to these data to learn models that effectively relate structure and function. However, extracting meaningful numerical descriptors of protein sequence and structure is a key issue that requires an efficient and widely available solution. RESULTS: We here introduce ProtDCal, a new computational software suite capable of generating tens of thousands of features considering both sequence-based and 3D-structural descriptors. We demonstrate, by means of principle component analysis and Shannon entropy tests, how ProtDCal's sequence-based descriptors provide new and more relevant information not encoded by currently available servers for sequence-based protein feature generation. The wide diversity of the 3D-structure-based features generated by ProtDCal is shown to provide additional complementary information and effectively completes its general protein encoding capability. As demonstration of the utility of ProtDCal's features, prediction models of N-linked glycosylation sites are trained and evaluated. Classification performance compares favourably with that of contemporary predictors of N-linked glycosylation sites, in spite of not using domain-specific features as input information. CONCLUSIONS: ProtDCal provides a friendly and cross-platform graphical user interface, developed in the Java programming language and is freely available at: http://bioinf.sce.carleton.ca/ProtDCal/ . ProtDCal introduces local and group-based encoding which enhances the diversity of the information captured by the computed features. Furthermore, we have shown that adding structure-based descriptors contributes non-redundant additional information to the features-based characterization of polypeptide systems. This software is intended to provide a useful tool for general-purpose encoding of protein sequences and structures for applications is protein classification, similarity analyses and function prediction. Yasser B. Ruiz-Blanco, Waldo Paz, James R. Green, Yovani Marrero-Ponce |
BMC Bioinform. | 3 |
| 2014 | Efficient prediction of human protein-protein interactions at a global scaleabstractBACKGROUND: Our knowledge of global protein-protein interaction (PPI) networks in complex organisms such as humans is hindered by technical limitations of current methods. RESULTS: On the basis of short co-occurring polypeptide regions, we developed a tool called MP-PIPE capable of predicting a global human PPI network within 3 months. With a recall of 23% at a precision of 82.1%, we predicted 172,132 putative PPIs. We demonstrate the usefulness of these predictions through a range of experiments. CONCLUSIONS: The speed and accuracy associated with MP-PIPE can make this a potential tool to study individual human PPI networks (from genomic sequences alone) for personalized medicine. Andrew Schoenrock, Bahram Samanfar, Sylvain Pitre, Mohsen Hooshyar, Charles A. Phillips, Sadhna Phanse, Katayoun Omidi, Yuan Gui, Md Alamgir, Alex Wong 0003, Fredrik Barrenäs, Mohan Babu, Mikael Benson, Michael A. Langston, James R. Green, Frank Dehne, Ashkan Golshani |
BMC Bioinform. | 17 |
| 2011 | MP-PIPE: a massively parallel protein-protein interaction prediction engineabstractInteractions among proteins are essential to many biological functions in living cells but experimentally detected interactions represent only a small fraction of the real interaction network. Computational protein interaction prediction methods have become important to augment the experimental methods; in particular sequence based prediction methods that do not require additional data such as homologous sequences or 3D structure information which are often not available. Our Protein Interaction Prediction Engine (PIPE) method falls into this category. Park has recently compared PIPE with the other competing methods and concluded that our method "significantly outperforms the others in terms of recall-precision across both the yeast and human data". Here, we present MP-PIPE, a new massively parallel PIPE implementation for large scale, high throughput protein interaction prediction. MP-PIPE enabled us to perform the first ever complete scan of the entire human protein interaction network; a massively parallel computational experiment which took three months of full time 24/7 computation on a dedicated SUN UltraSparc T2+ based cluster with 50 nodes, 800 processor cores and 6,400 hardware supported threads. The implications for the understanding of human cell function will be significant as biologists are starting to analyze the 130,470 new protein interactions and possible new pathways in Human cells predicted by MP-PIPE. Andrew Schoenrock, Frank Dehne, James R. Green, Ashkan Golshani, Sylvain Pitre |
ICS | 3 |
| 2011 | Binding Site Prediction for Protein-Protein Interactions and Novel Motif Discovery using Re-occurring Polypeptide SequencesabstractBACKGROUND: While there are many methods for predicting protein-protein interaction, very few can determine the specific site of interaction on each protein. Characterization of the specific sequence regions mediating interaction (binding sites) is crucial for an understanding of cellular pathways. Experimental methods often report false binding sites due to experimental limitations, while computational methods tend to require data which is not available at the proteome-scale. Here we present PIPE-Sites, a novel method of protein specific binding site prediction based on pairs of re-occurring polypeptide sequences, which have been previously shown to accurately predict protein-protein interactions. PIPE-Sites operates at high specificity and requires only the sequences of query proteins and a database of known binary interactions with no binding site data, making it applicable to binding site prediction at the proteome-scale. RESULTS: PIPE-Sites was evaluated using a dataset of 265 yeast and 423 human interacting proteins pairs with experimentally-determined binding sites. We found that PIPE-Sites predictions were closer to the confirmed binding site than those of two existing binding site prediction methods based on domain-domain interactions, when applied to the same dataset. Finally, we applied PIPE-Sites to two datasets of 2347 yeast and 14,438 human novel interacting protein pairs predicted to interact with high confidence. An analysis of the predicted interaction sites revealed a number of protein subsequences which are highly re-occurring in binding sites and which may represent novel binding motifs. CONCLUSIONS: PIPE-Sites is an accurate method for predicting protein binding sites and is applicable to the proteome-scale. Thus, PIPE-Sites could be useful for exhaustive analysis of protein binding patterns in whole proteomes as well as discovery of novel binding motifs. PIPE-Sites is available online at http://pipe-sites.cgmlab.org/. Adam Amos-Binks, Catalin Patulea, Sylvain Pitre, Andrew Schoenrock, Yuan Gui, James R. Green, Ashkan Golshani, Frank Dehne |
BMC Bioinform. | 6 |
| 2009 | PCI-SS: MISO dynamic nonlinear protein secondary structure predictionabstractBACKGROUND: Since the function of a protein is largely dictated by its three dimensional configuration, determining a protein's structure is of fundamental importance to biology. Here we report on a novel approach to determining the one dimensional secondary structure of proteins (distinguishing alpha-helices, beta-strands, and non-regular structures) from primary sequence data which makes use of Parallel Cascade Identification (PCI), a powerful technique from the field of nonlinear system identification. RESULTS: Using PSI-BLAST divergent evolutionary profiles as input data, dynamic nonlinear systems are built through a black-box approach to model the process of protein folding. Genetic algorithms (GAs) are applied in order to optimize the architectural parameters of the PCI models. The three-state prediction problem is broken down into a combination of three binary sub-problems and protein structure classifiers are built using 2 layers of PCI classifiers. Careful construction of the optimization, training, and test datasets ensures that no homology exists between any training and testing data. A detailed comparison between PCI and 9 contemporary methods is provided over a set of 125 new protein chains guaranteed to be dissimilar to all training data. Unlike other secondary structure prediction methods, here a web service is developed to provide both human- and machine-readable interfaces to PCI-based protein secondary structure prediction. This server, called PCI-SS, is available at http://bioinf.sce.carleton.ca/PCISS. In addition to a dynamic PHP-generated web interface for humans, a Simple Object Access Protocol (SOAP) interface is added to permit invocation of the PCI-SS service remotely. This machine-readable interface facilitates incorporation of PCI-SS into multi-faceted systems biology analysis pipelines requiring protein secondary structure information, and greatly simplifies high-throughput analyses. XML is used to represent the input protein sequence data and also to encode the resulting structure prediction in a machine-readable format. To our knowledge, this represents the only publicly available SOAP-interface for a protein secondary structure prediction service with published WSDL interface definition. CONCLUSION: Relative to the 9 contemporary methods included in the comparison cascaded PCI classifiers perform well, however PCI finds greatest application as a consensus classifier. When PCI is used to combine a sequence-to-structure PCI-based classifier with the current leading ANN-based method, PSIPRED, the overall error rate (Q3) is maintained while the rate of occurrence of a particularly detrimental error is reduced by up to 25%. This improvement in BAD score, combined with the machine-readable SOAP web service interface makes PCI-SS particularly useful for inclusion in a tertiary structure prediction pipeline. James R. Green, Michael J. Korenberg, Mohammed O. Aboul-Magd |
BMC Bioinform. | 1 |