Igor Jurisica

dblp:27/4777 · DBLP profile ↗
← Back
39ranked-venue papers
6as first author
4since 2021 · last 2023
0000-0002-2507-946XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 24 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 11 · 4 first-authorDatabases, data management, data science and information retrieval · 5 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2023 USNAP: fast unique dense region detection and its application to lung cancer
abstract
MOTIVATION: Many real-world problems can be modeled as annotated graphs. Scalable graph algorithms that extract actionable information from such data are in demand since these graphs are large, varying in topology, and have diverse node/edge annotations. When these graphs change over time they create dynamic graphs, and open the possibility to find patterns across different time points. In this article, we introduce a scalable algorithm that finds unique dense regions across time points in dynamic graphs. Such algorithms have applications in many different areas, including the biological, financial, and social domains. RESULTS: There are three important contributions to this manuscript. First, we designed a scalable algorithm, USNAP, to effectively identify dense subgraphs that are unique to a time stamp given a dynamic graph. Importantly, USNAP provides a lower bound of the density measure in each step of the greedy algorithm. Second, insights and understanding obtained from validating USNAP on real data show its effectiveness. While USNAP is domain independent, we applied it to four non-small cell lung cancer gene expression datasets. Stages in non-small cell lung cancer were modeled as dynamic graphs, and input to USNAP. Pathway enrichment analyses and comprehensive interpretations from literature show that USNAP identified biologically relevant mechanisms for different stages of cancer progression. Third, USNAP is scalable, and has a time complexity of O(m+mc log nc+nc log nc), where m is the number of edges, and n is the number of vertices in the dynamic graph; mc is the number of edges, and nc is the number of vertices in the collapsed graph. AVAILABILITY AND IMPLEMENTATION: The code of USNAP is available at https://www.cs.utoronto.ca/~juris/data/USNAP22.
Serene Wong, Chiara Pastrello, Max Kotlyar, Christos Faloutsos, Igor Jurisica
Bioinform.5
2022 Pathway integration and annotation: building a puzzle with non-matching pieces and no reference picture
abstract
Biological pathways are a broadly used formalism for representing and interpreting the cascade of biochemical reactions underlying cellular and biological mechanisms. Pathway representation provides an ontological link among biomolecules such as RNA, DNA, small molecules, proteins, protein complexes, hormones and genes. Frequently, pathway annotations are used to identify mechanisms linked to genes within affected biological contexts. This important role and the simplicity and elegance in representing complex interactions led to an explosion of pathway representations and databases. Unfortunately, the lack of overlap across databases results in inconsistent enrichment analysis results, unless databases are integrated. However, due to absence of consensus, guidelines or gold standards in pathway definition and representation, integration of data across pathway databases is not straightforward. Despite multiple attempts to provide consolidated pathways, highly related, redundant, poorly overlapping or ambiguous pathways continue to render pathways analysis inconsistent and hard to interpret. Ontology-based integration will promote unbiased, comprehensive yet streamlined analysis of experiments, and will reduce the number of enriched pathways when performing pathway enrichment analysis. Moreover, appropriate and consolidated pathways provide better training data for pathway prediction algorithms. In this manuscript, we describe the current methods for pathway consolidation, their strengths and pitfalls, and highlight directions for future improvements to this research area.
Giuseppe Agapito, Chiara Pastrello, Yun Niu, Igor Jurisica
Briefings Bioinform.4
2022 miRAnno - network-based functional microRNA annotation
abstract
MOTIVATION: Functional annotation is a common part of microRNA (miRNA)-related research, typically carried as pathway enrichment analysis of the selected miRNA targets. Here, we propose miRAnno, a fast and easy-to-use web application for miRNA annotation. RESULTS: miRAnno uses comprehensive molecular interaction network and random walks with restart to measure the association between miRNAs and individual pathways. Independent validation shows that miRAnno achieves higher signal-to-noise ratio compared to the standard enrichment analysis. AVAILABILITY AND IMPLEMENTATION: miRAnno is freely available at https://ophid.utoronto.ca/miRAnno/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tomás Tokár, Chiara Pastrello, Mark Abovsky, Sara Rahmati, Igor Jurisica
Bioinform.5
2021 Comprehensive pathway enrichment analysis workflows: COVID-19 case study
abstract
Abstract The coronavirus disease 2019 (COVID-19) outbreak due to the novel coronavirus named severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) has been classified as a pandemic disease by the World Health Organization on the 12th March 2020. This world-wide crisis created an urgent need to identify effective countermeasures against SARS-CoV-2. In silico methods, artificial intelligence and bioinformatics analysis pipelines provide effective and useful infrastructure for comprehensive interrogation and interpretation of available data, helping to find biomarkers, explainable models and eventually cures. One class of such tools, pathway enrichment analysis (PEA) methods, helps researchers to find possible key targets present in biological pathways of host cells that are targeted by SARS-CoV-2. Since many software tools are available, it is not easy for non-computational users to choose the best one for their needs. In this paper, we highlight how to choose the most suitable PEA method based on the type of COVID-19 data to analyze. We aim to provide a comprehensive overview of PEA techniques and the tools that implement them.
Giuseppe Agapito, Chiara Pastrello, Igor Jurisica
Briefings Bioinform.3
2020 BioPAX-Parser: parsing and enrichment analysis of BioPAX pathways
abstract
SUMMARY: Biological pathways are fundamental for learning about healthy and disease states. Many existing formats support automatic software analysis of biological pathways, e.g. BioPAX (Biological Pathway Exchange). Although some algorithms are available as web application or stand-alone tools, no general graphical application for the parsing of BioPAX pathway data exists. Also, very few tools can perform pathway enrichment analysis (PEA) using pathway encoded in the BioPAX format. To fill this gap, we introduce BiP (BioPAX-Parser), an automatic and graphical software tool aimed at performing the parsing and accessing of BioPAX pathway data, along with PEA by using information coming from pathways encoded in BioPAX. AVAILABILITY AND IMPLEMENTATION: BiP is freely available for academic and non-profit organizations at https://gitlab.com/giuseppeagapito/bip under the LGPL 2.1, the GNU Lesser General Public License. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Giuseppe Agapito, Chiara Pastrello, Pietro H. Guzzi, Igor Jurisica, Mario Cannataro
Bioinform.4
2020 GSOAP: a tool for visualization of gene set over-representation analysis
abstract
MOTIVATION: Gene sets over-representation analysis (GSOA) is a common technique of enrichment analysis that measures the overlap between a gene set and selected instances (e.g. pathways). Despite its popularity, there is currently no established standard for visualization of GSOA results. RESULTS: Here, we propose a visual exploration of the GSOA results by showing the relationships among the enriched instances, while highlighting important instance attributes, such as significance, closeness (centrality) and clustering. AVAILABILITY AND IMPLEMENTATION: GSOAP is implemented as an R package and is available at https://github.com/tomastokar/gsoap.
Tomás Tokár, Chiara Pastrello, Igor Jurisica
Bioinform.3
2018 SDREGION: Fast Spotting of Changing Communities in Biological Networks
abstract
Given a large, dynamic graph, how can we trace the activities of groups of vertices over time? Given a dynamic biological graph modeling a given disease progression, which genes interact closely at the early stage of the disease, and their interactions are being disrupted in the latter stage of the disease? Which genes interact sparsely at the early stage of the disease, and their interactions increase as the disease progresses? Knowing the answers to these questions is important as they give insights to the underlying molecular mechanism to disease progression, and potential treatments that target these mechanisms can be developed. There are three main contributions to this paper. First, we designed a novel algorithm, SDREGION, that identifies subgraphs that decrease or increase in density monotonically over time, referred to as d-regions or i-regions, respectively. We introduced the objective function, -density, for identifying d-(i-)regions. Second, SDREGION is a generic algorithm, applicable across several real datasets. In this manuscript, we showed its effectiveness, and made observations in the modeling of the progression of lung cancer. In particular, we observed that SDREGION identified d-(i-)regions that capture mechanisms that align with literature. Importantly, findings that were identified but were not retrospectively validated by literature may provide novel mechanisms in tumor progression that will guide future biological experiments. Third, SDREGION is scalable with a time complexity of O(mlogn + nlogn) where m is the number of edges, and n is the number of vertices in a given dynamic graph.
Serene Wong, Chiara Pastrello, Max Kotlyar, Christos Faloutsos, Igor Jurisica
KDD5
2018 Encompassing new use cases - level 3.0 of the HUPO-PSI format for molecular interactions
abstract
BACKGROUND: Systems biologists study interaction data to understand the behaviour of whole cell systems, and their environment, at a molecular level. In order to effectively achieve this goal, it is critical that researchers have high quality interaction datasets available to them, in a standard data format, and also a suite of tools with which to analyse such data and form experimentally testable hypotheses from them. The PSI-MI XML standard interchange format was initially published in 2004, and expanded in 2007 to enable the download and interchange of molecular interaction data. PSI-XML2.5 was designed to describe experimental data and to date has fulfilled this basic requirement. However, new use cases have arisen that the format cannot properly accommodate. These include data abstracted from more than one publication such as allosteric/cooperative interactions and protein complexes, dynamic interactions and the need to link kinetic and affinity data to specific mutational changes. RESULTS: The Molecular Interaction workgroup of the HUPO-PSI has extended the existing, well-used XML interchange format for molecular interaction data to meet new use cases and enable the capture of new data types, following extensive community consultation. PSI-MI XML3.0 expands the capabilities of the format beyond simple experimental data, with a concomitant update of the tool suite which serves this format. The format has been implemented by key data producers such as the International Molecular Exchange (IMEx) Consortium of protein interaction databases and the Complex Portal. CONCLUSIONS: PSI-MI XML3.0 has been developed by the data producers, data users, tool developers and database providers who constitute the PSI-MI workgroup. This group now actively supports PSI-MI XML2.5 as the main interchange format for experimental data, PSI-MI XML3.0 which additionally handles more complex data types, and the simpler, tab-delimited MITAB2.5, 2.6 and 2.7 for rapid parsing and download.
M. Sivade Dumousseau, Diego Alonso-López, Mais G. Ammari, Glyn Bradley, Nancy H. Campbell, Arnaud Céol, Gianni Cesareni, Colin W. Combe, Javier De Las Rivas, Noemi del-Toro, Joshua Heimbach, Henning Hermjakob, Igor Jurisica, Luana Licata, Ruth C. Lovering, David J. Lynn, Birgit Meldal, Gos Micklem, Simona Panni, Pablo Porras, Sylvie Ricard-Blum, Bernd Roechert, Lukasz Salwínski, Anjali Shrivastava, Julie M. Sullivan, Nicolas Thierry-Mieg, Yo Yehudi, Kim Van Roey, Sandra E. Orchard
BMC Bioinform.13
2016 A Simulated Annealing Algorithm for Maximum Common Edge Subgraph Detection in Biological Networks
abstract
Network alignment is a challenging computational problem that identifies node or edge mappings between two or more networks, with the aim to unravel common patterns among them. Pairwise network alignment is already intractable, making multiple network comparison even more difficult. Here, we introduce a heuristic algorithm for the multiple maximum common edge subgraph problem that is able to detect large common substructures shared across multiple, real-world size networks efficiently. Our algorithm uses a combination of iterated local search, simulated annealing and a pheromone-based perturbation strategy. We implemented multiple local search strategies and annealing schedules, that were evaluated on a range of synthetic networks and real protein-protein interaction networks. Our method is parallelized and well-suited to exploit current multi-core CPU architectures. While it is generic, we apply it to unravel a biochemical backbone inherent in different species, modeled as multiple maximum common subgraphs.
Simon J. Larsen, Frederik G. Alkærsig, Henrik J. Ditzel, Igor Jurisica, Nicolas Alcaraz, Jan Baumbach
GECCO4
2016 Robust quantitative scratch assay
abstract
UNLABELLED: The wound healing assay (or scratch assay) is a technique frequently used to quantify the dependence of cell motility-a central process in tissue repair and evolution of disease-subject to various treatments conditions. However processing the resulting data is a laborious task due its high throughput and variability across images. This Robust Quantitative Scratch Assay algorithm introduced statistical outputs where migration rates are estimated, cellular behaviour is distinguished and outliers are identified among groups of unique experimental conditions. Furthermore, the RQSA decreased measurement errors and increased accuracy in the wound boundary at comparable processing times compared to previously developed method (TScratch). AVAILABILITY AND IMPLEMENTATION: The RQSA is freely available at: http://ophid.utoronto.ca/RQSA/RQSA_Scripts.zip The image sets used for training and validation and results are available at: (http://ophid.utoronto.ca/RQSA/trainingSet.zip, http://ophid.utoronto.ca/RQSA/validationSet.zip, http://ophid.utoronto.ca/RQSA/ValidationSetResults.zip, http://ophid.utoronto.ca/RQSA/ValidationSet_H1975.zip, http://ophid.utoronto.ca/RQSA/ValidationSet_H1975Results.zip, http://ophid.utoronto.ca/RQSA/RobustnessSet.zip, http://ophid.utoronto.ca/RQSA/RobustnessSet.zip). Supplementary Material is provided for detailed description of the development of the RQSA. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Andrea Vargas, Marc Angeli, Chiara Pastrello, Rosanne McQuaid, Andrea Jurisicova, Igor Jurisica
Bioinform.7
2015 Where are we at regarding species translation? A review of the sbv IMPROVER challenge
abstract
Abstract Contact: [email protected]
Julia Hoeng, Manuel C. Peitsch, Pablo Meyer 0001, Igor Jurisica
Bioinform.4
2015 Prioritizing Therapeutics for Lung Cancer: An Integrative Meta-analysis of Cancer Gene Signatures and Chemogenomic Data
abstract
Repurposing FDA-approved drugs with the aid of gene signatures of disease can accelerate the development of new therapeutics. A major challenge to developing reliable drug predictions is heterogeneity. Different gene signatures of the same disease or drug treatment often show poor overlap across studies, as a consequence of both biological and technical variability, and this can affect the quality and reproducibility of computational drug predictions. Existing algorithms for signature-based drug repurposing use only individual signatures as input. But for many diseases, there are dozens of signatures in the public domain. Methods that exploit all available transcriptional knowledge on a disease should produce improved drug predictions. Here, we adapt an established meta-analysis framework to address the problem of drug repurposing using an ensemble of disease signatures. Our computational pipeline takes as input a collection of disease signatures, and outputs a list of drugs predicted to consistently reverse pathological gene changes. We apply our method to conduct the largest and most systematic repurposing study on lung cancer transcriptomes, using 21 signatures. We show that scaling up transcriptional knowledge significantly increases the reproducibility of top drug hits, from 44% to 78%. We extensively characterize drug hits in silico, demonstrating that they slow growth significantly in nine lung cancer cell lines from the NCI-60 collection, and identify CALM1 and PLA2G4A as promising drug targets for lung cancer. Our meta-analysis pipeline is general, and applicable to any disease context; it can be applied to improve the results of signature-based drug repurposing by leveraging the large number of disease signatures in the public domain.
Kristen Fortney, Joshua Griesman, Max Kotlyar, Chiara Pastrello, Marc Angeli, Ming-Sound Tsao, Igor Jurisica
PLoS Comput. Biol.7
2014 Knowledge Discovery and interactive Data Mining in Bioinformatics - State-of-the-Art, future challenges and research directions
abstract
Computers are incredibly fast, accurate, and stupid.Human beings are incredibly slow, inaccurate, and brilliant.Together they are powerful beyond imagination (Einstein never said that [1]).
Andreas Holzinger, Matthias Dehmer, Igor Jurisica
BMC Bioinform.3
2013 Visual Data Mining of Biological Networks: One Size Does Not Fit All
abstract
High-throughput technologies produce massive amounts of data. However, individual methods yield data specific to the technique used and biological setup. The integration of such diverse data is necessary for the qualitative analysis of information relevant to hypotheses or discoveries. It is often useful to integrate these datasets using pathways and protein interaction networks to get a broader view of the experiment. The resulting network needs to be able to focus on either the large-scale picture or on the more detailed small-scale subsets, depending on the research question and goals. In this tutorial, we illustrate a workflow useful to integrate, analyze, and visualize data from different sources, and highlight important features of tools to support such analyses.
Chiara Pastrello, David Otasek, Kristen Fortney, Giuseppe Agapito, Mario Cannataro, Elize Shirdel, Igor Jurisica
PLoS Comput. Biol.7
2012 Construction of New Medicines via Game Proof Search
abstract
The production of any new medicine requires solutions to many planning problems. The most fundamental of these is determining the sequence of chemical reactions necessary to physically create the drug. Surprisingly, these organic syntheses can be modeled as branching paths in a discrete, fully-observable state space, making the construction of new medicines an application of heuristic search. We describe a model of organic chemistry that is amenable to traditional AI techniques from game tree search, regression, and automatic assembly sequencing. We demonstrate the applicability of AND/OR graph search by developing the first chemistry solver to use proof-number search. Finally, we construct a benchmark suite of organic synthesis problems collected from undergraduate organic chemistry exams, and we analyze our solvers performance both on this suite and in recreating the synthetic plan for a multibillion dollar drug.
Abraham Heifets, Igor Jurisica
AAAI2
2011 Optimized application of penalized regression methods to diverse genomic data
abstract
MOTIVATION: Penalized regression methods have been adopted widely for high-dimensional feature selection and prediction in many bioinformatic and biostatistical contexts. While their theoretical properties are well-understood, specific methodology for their optimal application to genomic data has not been determined. RESULTS: Through simulation of contrasting scenarios of correlated high-dimensional survival data, we compared the LASSO, Ridge and Elastic Net penalties for prediction and variable selection. We found that a 2D tuning of the Elastic Net penalties was necessary to avoid mimicking the performance of LASSO or Ridge regression. Furthermore, we found that in a simulated scenario favoring the LASSO penalty, a univariate pre-filter made the Elastic Net behave more like Ridge regression, which was detrimental to prediction performance. We demonstrate the real-life application of these methods to predicting the survival of cancer patients from microarray data, and to classification of obese and lean individuals from metagenomic data. Based on these results, we provide an optimized set of guidelines for the application of penalized regression for reproducible class comparison and prediction with genomic data. AVAILABILITY AND IMPLEMENTATION: A parallelized implementation of the methods presented for regression and for simulation of synthetic data is provided as the pensim R package, available at http://cran.r-project.org/web/packages/pensim/index.html. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Levi Waldron, Melania Pintilie, Ming-Sound Tsao, Frances A. Shepherd, Curtis Huttenhower, Igor Jurisica
Bioinform.6
2011 A tree-based approach for motif discovery and sequence classification
abstract
MOTIVATION: Pattern discovery algorithms are widely used for the analysis of DNA and protein sequences. Most algorithms have been designed to find overrepresented motifs in sparse datasets of long sequences, and ignore most positional information. We introduce an algorithm optimized to exploit spatial information in sparse-but-populous datasets. RESULTS: Our algorithm Tree-based Weighted-Position Pattern Discovery and Classification (T-WPPDC) supports both unsupervised pattern discovery and supervised sequence classification. It identifies positionally enriched patterns using the Kullback-Leibler distance between foreground and background sequences at each position. This spatial information is used to discover positionally important patterns. T-WPPDC then uses a scoring function to discriminate different biological classes. We validated T-WPPDC on an important biological problem: prediction of single nucleotide polymorphisms (SNPs) from flanking sequence. We evaluated 672 separate experiments on 120 datasets derived from multiple species. T-WPPDC outperformed other pattern discovery methods and was comparable to the supervised machine learning algorithms. The algorithm is computationally efficient and largely insensitive to dataset size. It allows arbitrary parameterization and is embarrassingly parallelizable. CONCLUSIONS: T-WPPDC is a minimally parameterized algorithm for both pattern discovery and sequence classification that directly incorporates positional information. We use it to confirm the predictability of SNPs from flanking sequence, and show that positional information is a key to this biological problem. AVAILABILITY: The algorithm, code and data are available at: http://www.cs.utoronto.ca/~juris/data/TWPPDC
Rui Yan 0008, Paul C. Boutros, Igor Jurisica
Bioinform.3
2011 Cancer Computational Biology
abstract
EditorialIntroduction of high-throughput measurement technolo-gies combined with the increase of the scientific knowl-edge base, with respect to our understanding of cellularand biological processes, resulted in establishing compu-ter and information science as an important and funda-mental component of modern biology. High-throughputmeasurement technologies, such as microarray-basedprofiling, mass spectrometry screens, and high-through-put sequencing, give rise to several computational chal-lenges. On one hand, they require a rigorous approachto assay design. Scientists and technology developerswork on optimizing assay components so as to maxi-mize the information obtained through the measure-ment. On the other hand, the use of high-throughputmeasurement gives rise to large quantities of data thatneeds to be pre-processed and analyzed to obtain mean-ingful knowledge. This processing and analysis is per-formed on various levels - from pre-processing the rawdata, such as images from microarrays or raw sequencereads - to analyzing the data and to the discovery of bio-markers or other biologically meaningful characteristics.Measurement technology addresses several aspects ofcellular processes such as DNA, RNA, proteomics,metabolomics, epigenetics and pathways. This increasein the scientific knowledge base also leads to a centralrole played by data analysis and modeling, stronglygrounded in computational methods. Systems biology orintegrative biology approaches and network analysis areof specific importance in this context.The above is even further emphasized in the contextof cancer research. Samples are complex and heteroge-neous, and cancer related mechanisms involve manylayers of the process that leads from the genome to cel-lular function. One example of a specific need of canceris the study of large scale aberrations in the genome.CNVs (copy number variations) were recentlyrecognized as abundant in normal cell populations andas related to many other disease types but they are stilla hallmark of cancer [1,2]. Genomes in cancer cellsoften have a structure that allows them to bypassgrowth control cellular processes. Regions coding fortumor suppressor genes are often deleted and regionsharboring oncogenes may be amplified. This is the case,for example, for p16 and myc, respectively [3-5]. Rear-rangements, such as inversions and translocations, giverise to tumor-driving fusion products as in the case ofBCR-Abl and the Philadelphia Chromosome as well asin more recent findings implicating fusion structures insolid tumors. Cancer research therefore makes use ofdata analysis methods and tools that address interpreta-tion of copy number data and the understanding of theeffect of genome changes on transcriptome level as wellas proteome level profiles of tumors. Other specificcomputational needs of cancer research are related toepigenetic changes, somatic evolution, definition of genesets in the context of specific cancer types, and to drugsand data that measures the effects of drugs.Computational biologists focusing on cancer developmethods for the genome scale characterization of tumors,on various levels of the molecular process. Data analysismethods often rely on the analysis of high-throughputmeasurement data and they provide understanding of therelationship between various molecular characteristics ofcells. For example - how do genome structural aberra-tions and changes in copy number, a result of increasedgenome instability in cancer, affect the expression ofgenes and other functional elements such as miRNA, andhow do the latter changes affect the function of relatedproteins. Understanding of the association of genomiccharacteristics and clinical properties of primary tumorsamples, xenografts or cell lines contributes to persona-lized cancer medicine through the development of pre-dictive biomarkers of drug efficacy. Many researchprojects therefore aim to discover biomarkers, at eithergenome, transcriptome or proteome level that are prog-nostic of cancer progression or predictive of response tospecific therapeutic agents [6,7]. Cancer computationalbiology also focuses on analyzing molecules and
Zohar Yakhini, Igor Jurisica
BMC Bioinform.2
2010 Runtime Estimation Using the Case-Based Reasoning Approach for Scheduling in a Grid Environment
Edward Xia, Igor Jurisica, Julie Waterhouse, Valerie Sloan
ICCBR2
2010 Evaluation of linguistic features useful in extraction of interactions from PubMed; Application to annotating known, high-throughput and predicted interactions in I2D
abstract
MOTIVATION: Identification and characterization of protein-protein interactions (PPIs) is one of the key aims in biological research. While previous research in text mining has made substantial progress in automatic PPI detection from literature, the need to improve the precision and recall of the process remains. More accurate PPI detection will also improve the ability to extract experimental data related to PPIs and provide multiple evidence for each interaction. RESULTS: We developed an interaction detection method and explored the usefulness of various features in automatically identifying PPIs in text. The results show that our approach outperforms other systems using the AImed dataset. In the tests where our system achieves better precision with reduced recall, we discuss possible approaches for improvement. In addition to test datasets, we evaluated the performance on interactions from five human-curated databases-BIND, DIP, HPRD, IntAct and MINT-where our system consistently identified evidence for approximately 60% of interactions when both proteins appear in at least one sentence in the PubMed abstract. We then applied the system to extract articles from PubMed to annotate known, high-throughput and interologous interactions in I(2)D. AVAILABILITY: The data and software are available at: http://www.cs.utoronto.ca/ approximately juris/data/BI09/.
Yun Niu, David Otasek, Igor Jurisica
Bioinform.3
2010 The FlowVizMenu and Parallel Scatterplot Matrix: Hybrid Multidimensional Visualizations for Network Exploration
abstract
A standard approach for visualizing multivariate networks is to use one or more multidimensional views (for example, scatterplots) for selecting nodes by various metrics, possibly coordinated with a node-link view of the network. In this paper, we present three novel approaches for achieving a tighter integration of these views through hybrid techniques for multidimensional visualization, graph selection and layout. First, we present the FlowVizMenu, a radial menu containing a scatterplot that can be popped up transiently and manipulated with rapid, fluid gestures to select and modify the axes of its scatterplot. Second, the FlowVizMenu can be used to steer an attribute-driven layout of the network, causing certain nodes of a node-link diagram to move toward their corresponding positions in a scatterplot while others can be positioned manually or by force-directed layout. Third, we describe a novel hybrid approach that combines a scatterplot matrix (SPLOM) and parallel coordinates called the Parallel Scatterplot Matrix (P-SPLOM), which can be used to visualize and select features within the network. We also describe a novel arrangement of scatterplots called the Scatterplot Staircase (SPLOS) that requires less space than a traditional scatterplot matrix. Initial user feedback is reported.
Christophe Viau, Michael J. McGuffin, Yves Chiricota, Igor Jurisica
IEEE Trans. Vis. Comput. Graph.4
2009 NAViGaTOR: Network Analysis, Visualization and Graphing Toronto
abstract
SUMMARY: NAViGaTOR is a powerful graphing application for the 2D and 3D visualization of biological networks. NAViGaTOR includes a rich suite of visual mark-up tools for manual and automated annotation, fast and scalable layout algorithms and OpenGL hardware acceleration to facilitate the visualization of large graphs. Publication-quality images can be rendered through SVG graphics export. NAViGaTOR supports community-developed data formats (PSI-XML, BioPax and GML), is platform-independent and is extensible through a plug-in architecture. AVAILABILITY: NAViGaTOR is freely available to the research community from http://ophid.utoronto.ca/navigator/. Installers and documentation are provided for 32- and 64-bit Windows, Mac, Linux and Unix. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Kevin R. Brown, David Otasek, Michael J. McGuffin, Wing Xie, Baiju Devani, Ian Lawson van Toch, Igor Jurisica
Bioinform.8
2009 Interaction Techniques for Selecting and Manipulating Subgraphs in Network Visualizations
abstract
We present a novel and extensible set of interaction techniques for manipulating visualizations of networks by selecting subgraphs and then applying various commands to modify their layout or graphical properties. Our techniques integrate traditional rectangle and lasso selection, and also support selecting a node's neighbourhood by dragging out its radius (in edges) using a novel kind of radial menu. Commands for translation, rotation, scaling, or modifying graphical properties (such as opacity) and layout patterns can be performed by using a hotbox (a transiently popped-up, semi-transparent set of widgets) that has been extended in novel ways to integrate specification of commands with 1D or 2D arguments. Our techniques require only one mouse button and one keyboard key, and are designed for fast, gestural, in-place interaction. We present the design and integration of these interaction techniques, and illustrate their use in interactive graph visualization. Our techniques are implemented in NAViGaTOR, a software package for visualizing and analyzing biological networks. An initial usability study is also reported.
Michael J. McGuffin, Igor Jurisica
IEEE Trans. Vis. Comput. Graph.2
2008 Detecting Protein-Protein Interaction Sentences Using a Mixture Model
Yun Niu, Igor Jurisica
NLDB2
2006 Efficient estimation of graphlet frequency distributions in protein-protein interaction networks
abstract
MOTIVATION: Algorithmic and modeling advances in the area of protein-protein interaction (PPI) network analysis could contribute to the understanding of biological processes. Local structure of networks can be measured by the frequency distribution of graphlets, small connected non-isomorphic induced subgraphs. This measure of local structure has been used to show that high-confidence PPI networks have local structure of geometric random graphs. Finding graphlets exhaustively in a large network is computationally intensive. More complete PPI networks, as well as PPI networks of higher organisms, will thus require efficient heuristic approaches. RESULTS: We propose two efficient and scalable heuristics for finding graphlets in high-confidence PPI networks. We show that both PPI and their model geometric random networks, have defined boundaries that are sparser than the 'inner parts' of the networks. In addition, these networks exhibit 'uniformity' of local structure inside the networks. Our first heuristic exploits these two structural properties of PPI and geometric random networks to find good estimates of graphlet frequency distributions in these networks up to 690 times faster than the exhaustive searches. Our second heuristic is a variant of a more standard sampling technique and it produces accurate approximate results up to 377 times faster than the exhaustive searches. We indicate how the combination of these approaches may result in an even better heuristic. AVAILABILITY: Supplementary information is available at http://www.cs.toronto.edu/~natasha/BIOINF-2005-0946/Supplementary.pdf. Software implementing the algorithms is available at http://www.cs.toronto.edu/~natasha/BIOINF-2005-0946/estimate_grap-hlets.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Natasa Przulj, Derek G. Corneil, Igor Jurisica
Bioinform.3
2005 An Ensemble of Case-Based Classifiers for High-Dimensional Biological Domains
Niloofar Arshadi, Igor Jurisica
ICCBR2
2005 Online Predicted Human Interaction Database
abstract
MOTIVATION: High-throughput experiments are being performed at an ever-increasing rate to systematically elucidate protein-protein interaction (PPI) networks for model organisms, while the complexities of higher eukaryotes have prevented these experiments for humans. RESULTS: The Online Predicted Human Interaction Database (OPHID) is a web-based database of predicted interactions between human proteins. It combines the literature-derived human PPI from BIND, HPRD and MINT, with predictions made from Saccharomyces cerevisiae, Caenorhabditis elegans, Drosophila melanogaster and Mus musculus. The 23,889 predicted interactions currently listed in OPHID are evaluated using protein domains, gene co-expression and Gene Ontology terms. OPHID can be queried using single or multiple IDs and results can be visualized using our custom graph visualization program. AVAILABILITY: Freely available to academic users at http://ophid.utoronto.ca, both in tab-delimited and PSI-MI formats. Commercial users, please contact I.J. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: http://ophid.utoronto.ca/supplInfo.pdf.
Kevin R. Brown, Igor Jurisica
Bioinform.2
2005 Data Mining for Case-Based Reasoning in High-Dimensional Biological Domains
abstract
Case-based reasoning (CBR) is a suitable paradigm for class discovery in molecular biology, where the rules that define the domain knowledge are difficult to obtain and the number and the complexity of the rules affecting the problem are too large for formal knowledge representation. To extend the capabilities of CBR, we propose the mixture of experts for case-based reasoning (MOE4CBR), a method that combines an ensemble of CBR classifiers with spectral clustering and logistic regression. Our approach not only achieves higher prediction accuracy, but also leads to the selection of a subset of features that have meaningful relationships with their class labels. We evaluate MOE4CBR by applying the method to a CBR system called TA3 - a computational framework for CBR systems. For two ovarian mass spectrometry data sets, the prediction accuracy improves from 80 percent to 93 percent and from 90 percent to 98.4 percent, respectively. We also apply the method to leukemia and lung microarray data sets with prediction accuracy improving from 65 percent to 74 percent and from 60 percent to 70 percent, respectively. Finally, we compare our list of discovered biomarkers with the lists of selected biomarkers from other studies for the mass spectrometry data sets.
Niloofar Arshadi, Igor Jurisica
IEEE Trans. Knowl. Data Eng.2
2004 Protein complex prediction via cost-based clustering
abstract
MOTIVATION: Understanding principles of cellular organization and function can be enhanced if we detect known and predict still undiscovered protein complexes within the cell's protein-protein interaction (PPI) network. Such predictions may be used as an inexpensive tool to direct biological experiments. The increasing amount of available PPI data necessitates an accurate and scalable approach to protein complex identification. RESULTS: We have developed the Restricted Neighborhood Search Clustering Algorithm (RNSC) to efficiently partition networks into clusters using a cost function. We applied this cost-based clustering algorithm to PPI networks of Saccharomyces cerevisiae, Drosophila melanogaster and Caenorhabditis elegans to identify and predict protein complexes. We have determined functional and graph-theoretic properties of true protein complexes from the MIPS database. Based on these properties, we defined filters to distinguish between identified network clusters and true protein complexes. CONCLUSIONS: Our application of the cost-based clustering algorithm provides an accurate and scalable method of detecting and predicting protein complexes within a PPI network.
Andrew D. King, Natasa Przulj, Igor Jurisica
Bioinform.3
2004 Modeling interactome: scale-free or geometric?
abstract
MOTIVATION: Networks have been used to model many real-world phenomena to better understand the phenomena and to guide experiments in order to predict their behavior. Since incorrect models lead to incorrect predictions, it is vital to have as accurate a model as possible. As a result, new techniques and models for analyzing and modeling real-world networks have recently been introduced. RESULTS: One example of large and complex networks involves protein-protein interaction (PPI) networks. We analyze PPI networks of yeast Saccharomyces cerevisiae and fruitfly Drosophila melanogaster using a newly introduced measure of local network structure as well as the standardly used measures of global network structure. We examine the fit of four different network models, including Erdos-Renyi, scale-free and geometric random network models, to these PPI networks with respect to the measures of local and global network structure. We demonstrate that the currently accepted scale-free model of PPI networks fails to fit the data in several respects and show that a random geometric model provides a much more accurate model of the PPI data. We hypothesize that only the noise in these networks is scale-free. CONCLUSIONS: We systematically evaluate how well-different network models fit the PPI networks. We show that the structure of PPI networks is better modeled by a geometric random graph than by a scale-free model. SUPPLEMENTARY INFORMATION: Supplementary information is available at http://www.cs.utoronto.ca/~juris/data/data/ppiGRG04/
Natasa Przulj, Derek G. Corneil, Igor Jurisica
Bioinform.3
2004 Functional topology in a network of protein interactions
abstract
MOTIVATION: The building blocks of biological networks are individual protein-protein interactions (PPIs). The cumulative PPI data set in Saccharomyces cerevisiae now exceeds 78 000. Studying the network of these interactions will provide valuable insight into the inner workings of cells. RESULTS: We performed a systematic graph theory-based analysis of this PPI network to construct computational models for describing and predicting the properties of lethal mutations and proteins participating in genetic interactions, functional groups, protein complexes and signaling pathways. Our analysis suggests that lethal mutations are not only highly connected within the network, but they also satisfy an additional property: their removal causes a disruption in network structure. We also provide evidence for the existence of alternate paths that bypass viable proteins in PPI networks, while such paths do not exist for lethal mutations. In addition, we show that distinct functional classes of proteins have differing network properties. We also demonstrate a way to extract and iteratively predict protein complexes and signaling pathways. We evaluate the power of predictions by comparing them with a random model, and assess accuracy of predictions by analyzing their overlap with MIPS database. CONCLUSIONS: Our models provide a means for understanding the complex wiring underlying cellular function, and enable us to predict essentiality, genetic interaction, function, protein complexes and cellular pathways. This analysis uncovers structure-function relationships observable in a large PPI network.
Natasa Przulj, Dennis A. Wigle, Igor Jurisica
Bioinform.3
2004 Ontologies for Knowledge Management: An Information Systems Perspective
Igor Jurisica, John Mylopoulos, Eric S. K. Yu
Knowl. Inf. Syst.1
2002 Binary tree-structured vector quantization approach to clustering and visualizing microarray data
abstract
Abstract Motivation: With the increasing number of gene expression databases, the need for more powerful analysis and visualization tools is growing. Many techniques have successfully been applied to unravel latent similarities among genes and/or experiments. Most of the current systems for microarray data analysis use statistical methods, hierarchical clustering, self-organizing maps, support vector machines, or k-means clustering to organize genes or experiments into ‘meaningful’ groups. Without prior explicit bias almost all of these clustering methods applied to gene expression data not only produce different results, but may also produce clusters with little or no biological relevance. Of these methods, agglomerative hierarchical clustering has been the most widely applied, although many limitations have been identified. Results: Starting with a systematic comparison of the underlying theories behind clustering approaches, we have devised a technique that combines tree-structured vector quantization and partitive k-means clustering (BTSVQ). This hybrid technique has revealed clinically relevant clusters in three large publicly available data sets. In contrast to existing systems, our approach is less sensitive to data preprocessing and data normalization. In addition, the clustering results produced by the technique have strong similarities to those of self-organizing maps (SOMs). We discuss the advantages and the mathematical reasoning behind our approach. Availability: The BTSVQ system is implemented in Matlab R12 using the SOM toolbox for the visualization and preprocessing of the data http://www.cis.hut.fi/projects/somtoolbox/ BTSVQ is available for non-commercial use http://www.uhnres.utoronto.ca/ta3/BTSVQ Contact: [email protected] Keywords: microarray data clustering and visulization; self-organizing maps, partitive k-means clustering; lung cancer.
Mujahid Sultan, Dennis A. Wigle, Christian A. Cumbaa, Marlena Maziarz, Janice I. Glasgow, Ming-Sound Tsao, Igor Jurisica
ISMB7
2001 Image-Feature Extraction for Protein Crystallization: Integrating Image Analysis and Case-Based Reasoning
Igor Jurisica, Phil Rogers, Janice I. Glasgow, Suzanne Fortier, Joseph R. Luft, Melissa A. Bianca, Robert J. Collins, George T. DeTitta
IAAI1
2000 Incremental Iterative Retrieval and Browsing for Efficient Conversational CBR Systems
Igor Jurisica, Janice I. Glasgow, John Mylopoulos
Appl. Intell.1
1998 Building Quality into Case-Based Reasoning Systems
Igor Jurisica, Brian A. Nixon
CAiSE1
1998 Case-based reasoning in IVF: prediction and knowledge mining
Igor Jurisica, John Mylopoulos, Janice I. Glasgow, Heather Shapiro, Robert F. Casper
Artif. Intell. Medicine1
1996 Case-Based Classification Using Similarity-Based Retrieval
abstract
Classification involves associating instances with particular classes by maximizing intra-class similarities and minimizing inter-class similarities. The paper presents a novel approach to case-based classification. The algorithm is based on a notion of similarity assessment and was developed for supporting flexible retrieval of relevant information. Validity of the proposed approach is tested on real world domains, and the system's performance is compared to that of other machine learning algorithms.
Igor Jurisica, Janice I. Glasgow
ICTAI1
1992 A Statistical Approach to Solving the EBL Utility Problem
Russell Greiner, Igor Jurisica
AAAI2