Liisa Holm

dblp:23/3677 · DBLP profile ↗
← Back
32ranked-venue papers
6as first author
1since 2021 · last 2022
0000-0002-7807-2966ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 32 · 6 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
21 papers
Bioinformatics and computational biology · 100% Computational science and engineering · 0% Medical and health informatics · 0%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 100%

Topics — the 30 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
protein structure analysis
0.542019
Benchmarking fold detection by DaliLite v.5 · Bioinform. 2019
Searching protein structure databases with DaliLite v.3 · Bioinform. 2008
DaliLite workbench for protein structure comparison · Bioinform. 2000
Bioinformatics and computational biology
protein function prediction
0.442015
PANNZER: high-throughput functional annotation of uncharacterized proteins in an error-prone environment · Bioinform. 2015
A novel method for assigning functional linkages to proteins using enhanced phylogenetic trees · Bioinform. 2011
SANS: high-throughput retrieval of protein sequences allowing 50% mismatches · Bioinform. 2012
Bioinformatics and computational biology › protein structure prediction › template-based modeling
fold recognition
0.412019
Benchmarking fold detection by DaliLite v.5 · Bioinform. 2019
Bioinformatics and computational biology
protein sequence analysis
0.222012
SANS: high-throughput retrieval of protein sequences allowing 50% mismatches · Bioinform. 2012
Bayesian search of functionally divergent protein subgroups and their function specific residues · Bioinform. 2006
Bioinformatics and computational biology › functional genomics › functional enrichment analysis
gene set enrichment analysis
0.212014
Gene set analysis: limitations in popular existing methods and proposed improvements · Bioinform. 2014
Bioinformatics and computational biology › sequence analysis
sequence similarity search
0.122012
SANS: high-throughput retrieval of protein sequences allowing 50% mismatches · Bioinform. 2012
Removing near-neighbour redundancy from large protein sequence collections · Bioinform. 1998
Bioinformatics and computational biology
metabolomics
0.112011
MPEA - metabolite pathway enrichment analysis · Bioinform. 2011
Bioinformatics and computational biology › genomics › microbial genomics
bacterial genomics
0.112009
LOCP - locating pilus operons in Gram-positive bacteria · Bioinform. 2009
Bioinformatics and computational biology › protein sequence analysis
protein sequence database search
0.112007
The global trace graph, a novel paradigm for searching protein sequence databases · Bioinform. 2007
Bioinformatics and computational biology › sequence analysis › homology detection
remote homology detection
0.112007
The global trace graph, a novel paradigm for searching protein sequence databases · Bioinform. 2007
Bioinformatics and computational biology › biostatistics › statistical bioinformatics
statistical genomics
0.112014
Gene set analysis: limitations in popular existing methods and proposed improvements · Bioinform. 2014
Bioinformatics and computational biology › protein analysis
protein domain interaction
0.112005
PSIbase: a database of Protein Structural Interactome map (PSIMAP) · Bioinform. 2005
Bioinformatics and computational biology › structural biology
protein structure and function
0.112005
PSIbase: a database of Protein Structural Interactome map (PSIMAP) · Bioinform. 2005
Bioinformatics and computational biology
sequence alignment
0.012003
Accurate detection of very sparse sequence motifs · RECOMB 2003
Bioinformatics and computational biology
phylogenetics
0.012011
A novel method for assigning functional linkages to proteins using enhanced phylogenetic trees · Bioinform. 2011
Bioinformatics and computational biology › phylogenetics
phylogenetic profiling
0.012011
A novel method for assigning functional linkages to proteins using enhanced phylogenetic trees · Bioinform. 2011
Bioinformatics and computational biology › protein sequence analysis
protein homology search
0.022000
Sequence search algorithm assessment and testing toolkit (SAT) · Bioinform. 2000
RSDB: representative protein sequence databases have high information content · Bioinform. 2000
Bioinformatics and computational biology › biological database
protein sequence database
0.022000
RSDB: representative protein sequence databases have high information content · Bioinform. 2000
Sequence search algorithm assessment and testing toolkit (SAT) · Bioinform. 2000
Bioinformatics and computational biology › protein function prediction › protein classification
protein family classification
0.012001
Picasso: generating a covering set of protein family profiles · Bioinform. 2001
Bioinformatics and computational biology
protein structure prediction
0.012000
Estimating the significance of sequence order in protein secondary structure and prediction · Bioinform. 2000
Bioinformatics and computational biology › protein structure prediction
secondary structure prediction
0.012000
Estimating the significance of sequence order in protein secondary structure and prediction · Bioinform. 2000
Parallel and multicore computing
parallel computing
0.012008
Searching protein structure databases with DaliLite v.3 · Bioinform. 2008
Bioinformatics and computational biology
multiple sequence alignment
0.011998
COFFEE: an objective function for multiple sequence alignments · Bioinform. 1998
Bioinformatics and computational biology
sequence analysis
0.011998
COFFEE: an objective function for multiple sequence alignments · Bioinform. 1998
Bioinformatics and computational biology › sequence analysis
sequence clustering
0.011998
Removing near-neighbour redundancy from large protein sequence collections · Bioinform. 1998
Bioinformatics and computational biology › biological database
sequence database curation
0.011998
Removing near-neighbour redundancy from large protein sequence collections · Bioinform. 1998
Bioinformatics and computational biology › sequence analysis
sequence database redundancy reduction
0.011998
Removing near-neighbour redundancy from large protein sequence collections · Bioinform. 1998
Bioinformatics and computational biology › structural bioinformatics
protein structure classification
0.011997
Decision Support System for the Evolutionary Classification of Protein Structures · ISMB 1997
Bioinformatics and computational biology › protein structure analysis
protein structure search
0.011995
3-D Lookup: Fast Protein Structure Database Searches at 90% Reliability · ISMB 1995
Bioinformatics and computational biology › structural bioinformatics › structural similarity
protein structure similarity search
0.011995
3-D Lookup: Fast Protein Structure Database Searches at 90% Reliability · ISMB 1995

Methods — techniques the papers use, named apart from their topics

knowledge-based search · 0.4hierarchical search · 0.4fmax benchmarking · 0.4weighted k-nearest neighbor · 0.2statistical testing · 0.2BLAST · 0.2rotation testing · 0.2permutation testing · 0.2asymptotic distribution modeling · 0.2k-nearest neighbor classification · 0.1knowledge-based search space pruning · 0.1
YearPublicationVenuePosition
2022 Correction: Novel comparison of evaluation metrics for gene ontology classifiers reveals drastic performance differences
abstract
Correction: Novel comparison of evaluation metrics for gene ontology classifiers reveals drastic performance differences Ilya Plyusnin, Liisa Holm, Petri To ¨ro ¨nenThere are several errors in Table 1.The values for the column rec in rows ic SimGIC2, ic2 Sim-GIC2, AJacc E, and ic2 Smin1 are incorrect.Please see the correct Table 1 here.
Ilya Plyusnin, Liisa Holm, Petri Törönen
PLoS Comput. Biol.2
2019 Benchmarking fold detection by DaliLite v.5
abstract
MOTIVATION: Protein structure comparison plays a fundamental role in understanding the evolutionary relationships between proteins. Here, we release a new version of the DaliLite standalone software. The novelties are hierarchical search of the structure database organized into sequence based clusters, and remote access to our knowledge base of structural neighbors. The detection of fold, superfamily and family level similarities by DaliLite and state-of-the-art competitors was benchmarked against a manually curated structural classification. RESULTS: Database search strategies were evaluated using Fmax with query-specific thresholds. DaliLite and DeepAlign outperformed TM-score based methods at all levels of the benchmark, and DaliLite outperformed DeepAlign at fold level. Hierarchical and knowledge-based searches got close to the performance of systematic pairwise comparison. The knowledge-based search was four times as efficient as the hierarchical search. The knowledge-based search dynamically adjusts the depth of the search, enabling a trade-off between speed and recall. AVAILABILITY AND IMPLEMENTATION: http://ekhidna2.biocenter.helsinki.fi/dali/README.v5.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Liisa Holm
Bioinform.1
2019 Novel comparison of evaluation metrics for gene ontology classifiers reveals drastic performance differences
abstract
Automated protein annotation using the Gene Ontology (GO) plays an important role in the biosciences. Evaluation has always been considered central to developing novel annotation methods, but little attention has been paid to the evaluation metrics themselves. Evaluation metrics define how well an annotation method performs and allows for them to be ranked against one another. Unfortunately, most of these metrics were adopted from the machine learning literature without establishing whether they were appropriate for GO annotations. We propose a novel approach for comparing GO evaluation metrics called Artificial Dilution Series (ADS). Our approach uses existing annotation data to generate a series of annotation sets with different levels of correctness (referred to as their signal level). We calculate the evaluation metric being tested for each annotation set in the series, allowing us to identify whether it can separate different signal levels. Finally, we contrast these results with several false positive annotation sets, which are designed to expose systematic weaknesses in GO assessment. We compared 37 evaluation metrics for GO annotation using ADS and identified drastic differences between metrics. We show that some metrics struggle to differentiate between different signal levels, while others give erroneously high scores to the false positive data sets. Based on our findings, we provide guidelines on which evaluation metrics perform well with the Gene Ontology and propose improvements to several well-known evaluation metrics. In general, we argue that evaluation metrics should be tested for their performance and we provide software for this purpose (https://bitbucket.org/plyusnin/ads/). ADS is applicable to other areas of science where the evaluation of prediction results is non-trivial.
Ilya Plyusnin, Liisa Holm, Petri Törönen
PLoS Comput. Biol.2
2018 TOPAZ: asymmetric suffix array neighbourhood search for massive protein databases
abstract
BACKGROUND: Protein homology search is an important, yet time-consuming, step in everything from protein annotation to metagenomics. Its application, however, has become increasingly challenging, due to the exponential growth of protein databases. In order to perform homology search at the required scale, many methods have been proposed as alternatives to BLAST that make an explicit trade-off between sensitivity and speed. One such method, SANSparallel, uses a parallel implementation of the suffix array neighbourhood search (SANS) technique to achieve high speed and provides several modes to allow for greater sensitivity at the expense of performance. RESULTS: We present a new approach called asymmetric SANS together with scored seeds and an alternative suffix array ordering scheme called optimal substitution ordering. These techniques dramatically improve both the sensitivity and speed of the SANS approach. Our implementation, TOPAZ, is one of the top performing methods in terms of speed, sensitivity and scalability. In our benchmark, searching UniProtKB for homologous proteins to the Dickeya solani proteome, TOPAZ took less than 3 minutes to achieve a sensitivity of 0.84 compared to BLAST. CONCLUSIONS: Despite the trade-off homology search methods have to make between sensitivity and speed, TOPAZ stands out as one of the most sensitive and highest performance methods currently available.
Alan Medlar, Liisa Holm
BMC Bioinform.2
2018 BARCOSEL: a tool for selecting an optimal barcode set for high-throughput sequencing
abstract
BACKGROUND: Current high-throughput sequencing platforms provide capacity to sequence multiple samples in parallel. Different samples are labeled by attaching a short sample specific nucleotide sequence, barcode, to each DNA molecule prior pooling them into a mix containing a number of libraries to be sequenced simultaneously. After sequencing, the samples are binned by identifying the barcode sequence within each sequence read. In order to tolerate sequencing errors, barcodes should be sufficiently apart from each other in sequence space. An additional constraint due to both nucleotide usage and basecalling accuracy is that the proportion of different nucleotides should be in balance in each barcode position. The number of samples to be mixed in each sequencing run may vary and this introduces a problem how to select the best subset of available barcodes at sequencing core facility for each sequencing run. There are plenty of tools available for de novo barcode design, but they are not suitable for subset selection. RESULTS: We have developed a tool which can be used for three different tasks: 1) selecting an optimal barcode set from a larger set of candidates, 2) checking the compatibility of user-defined set of barcodes, e.g. whether two or more libraries with existing barcodes can be combined in a single sequencing pool, and 3) augmenting an existing set of barcodes. In our approach the selection process is formulated as a minimization problem. We define the cost function and a set of constraints and use integer programming to solve the resulting combinatorial problem. Based on the desired number of barcodes to be selected and the set of candidate sequences given by user, the necessary constraints are automatically generated and the optimal solution can be found. The method is implemented in C programming language and web interface is available at http://ekhidna2.biocenter.helsinki.fi/barcosel . CONCLUSIONS: Increasing capacity of sequencing platforms raises the challenge of mixing barcodes. Our method allows the user to select a given number of barcodes among the larger existing barcode set so that both sequencing errors are tolerated and the nucleotide balance is optimized. The tool is easy to access via web browser.
Panu Somervuo, Patrik Koskinen, Liisa Holm, Petri Auvinen, Lars Paulin
BMC Bioinform.4
2016 Robust multi-group gene set analysis with few replicates
abstract
BACKGROUND: Competitive gene set analysis is a standard exploratory tool for gene expression data. Permutation-based competitive gene set analysis methods are preferable to parametric ones because the latter make strong statistical assumptions which are not always met. For permutation-based methods, we permute samples, as opposed to genes, as doing so preserves the inter-gene correlation structure. Unfortunately, up until now, sample permutation-based methods have required a minimum of six replicates per sample group. RESULTS: We propose a new permutation-based competitive gene set analysis method for multi-group gene expression data with as few as three replicates per group. The method is based on advanced sample permutation technique that utilizes all groups within a data set for pairwise comparisons. We present a comprehensive evaluation of different permutation techniques, using multiple data sets and contrast the performance of our method, mGSZm, with other state of the art methods. We show that mGSZm is robust, and that, despite only using less than six replicates, we are able to consistently identify a high proportion of the top ranked gene sets from the analysis of a substantially larger data set. Further, we highlight other methods where performance is highly variable and appears dependent on the underlying data set being analyzed. CONCLUSIONS: Our results demonstrate that robust gene set analysis of multi-group gene expression data is permissible with as few as three replicates. In doing so, we have extended the applicability of such approaches to resource constrained experiments where additional data generation is prohibitively difficult or expensive. An R package implementing the proposed method and supplementary materials are available from the website http://ekhidna.biocenter.helsinki.fi/downloads/pashupati/mGSZm.html .
Pashupati P. Mishra, Alan Medlar, Liisa Holm, Petri Törönen
BMC Bioinform.3
2015 PANNZER: high-throughput functional annotation of uncharacterized proteins in an error-prone environment
abstract
Abstract Motivation: The last decade has seen a remarkable growth in protein databases. This growth comes at a price: a growing number of submitted protein sequences lack functional annotation. Approximately 32% of sequences submitted to the most comprehensive protein database UniProtKB are labelled as ‘Unknown protein’ or alike. Also the functionally annotated parts are reported to contain 30–40% of errors. Here, we introduce a high-throughput tool for more reliable functional annotation called Protein ANNotation with Z-score (PANNZER). PANNZER predicts Gene Ontology (GO) classes and free text descriptions about protein functionality. PANNZER uses weighted k-nearest neighbour methods with statistical testing to maximize the reliability of a functional annotation. Results: Our results in free text description line prediction show that we outperformed all competing methods with a clear margin. In GO prediction we show clear improvement to our older method that performed well in CAFA 2011 challenge. Availability and implementation: The PANNZER program was developed using the Python programming language (Version 2.6). The stand-alone installation of the PANNZER requires MySQL database for data storage and the BLAST (BLASTALL v.2.2.21) tools for the sequence similarity search. The tutorial, evaluation test sets and results are available on the PANNZER web site. PANNZER is freely available at http://ekhidna.biocenter.helsinki.fi/pannzer. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Patrik Koskinen, Petri Törönen, Jussi Nokso-Koivisto, Liisa Holm
Bioinform.4
2014 Gene set analysis: limitations in popular existing methods and proposed improvements
abstract
MOTIVATION: Gene set analysis is the analysis of a set of genes that collectively contribute to a biological process. Most popular gene set analysis methods are based on empirical P-value that requires large number of permutations. Despite numerous gene set analysis methods developed in the past decade, the most popular methods still suffer from serious limitations. RESULTS: We present a gene set analysis method (mGSZ) based on Gene Set Z-scoring function (GSZ) and asymptotic P-values. Asymptotic P-value calculation requires fewer permutations, and thus speeds up the gene set analysis process. We compare the GSZ-scoring function with seven popular gene set scoring functions and show that GSZ stands out as the best scoring function. In addition, we show improved performance of the GSA method when the max-mean statistics is replaced by the GSZ scoring function. We demonstrate the importance of both gene and sample permutations by showing the consequences in the absence of one or the other. A comparison of asymptotic and empirical methods of P-value estimation demonstrates a clear advantage of asymptotic P-value over empirical P-value. We show that mGSZ outperforms the state-of-the-art methods based on two different evaluations. We compared mGSZ results with permutation and rotation tests and show that rotation does not improve our asymptotic P-values. We also propose well-known asymptotic distribution models for three of the compared methods. AVAILABILITY AND IMPLEMENTATION: mGSZ is available as R package from cran.r-project.org.
Pashupati P. Mishra, Petri Törönen, Yrjö Leino, Liisa Holm
Bioinform.4
2014 Comparative Genome-Scale Reconstruction of Gapless Metabolic Networks for Present and Ancestral Species
abstract
We introduce a novel computational approach, CoReCo, for comparative metabolic reconstruction and provide genome-scale metabolic network models for 49 important fungal species. Leveraging on the exponential growth in sequenced genome availability, our method reconstructs genome-scale gapless metabolic networks simultaneously for a large number of species by integrating sequence data in a probabilistic framework. High reconstruction accuracy is demonstrated by comparisons to the well-curated Saccharomyces cerevisiae consensus model and large-scale knock-out experiments. Our comparative approach is particularly useful in scenarios where the quality of available sequence data is lacking, and when reconstructing evolutionary distant species. Moreover, the reconstructed networks are fully carbon mapped, allowing their use in 13C flux analysis. We demonstrate the functionality and usability of the reconstructed fungal models with computational steady-state biomass production experiment, as these fungi include some of the most important production organisms in industrial biotechnology. In contrast to many existing reconstruction techniques, only minimal manual effort is required before the reconstructed models are usable in flux balance experiments. CoReCo is available at http://esaskar.github.io/CoReCo/.
Esa Pitkänen, Paula Jouhten, Jian Hou 0007, Muhammad Fahad Syed, Peter Blomberg, Jana Kludas, Merja Oja, Liisa Holm, Merja Penttilä, Juho Rousu, Mikko Arvas
PLoS Comput. Biol.8
2013 GOParGenPy: a high throughput method to generate Gene Ontology data matrices
abstract
BACKGROUND: Gene Ontology (GO) is a popular standard in the annotation of gene products and provides information related to genes across all species. The structure of GO is dynamic and is updated on a daily basis. However, the popular existing methods use outdated versions of GO. Moreover, these tools are slow to process large datasets consisting of more than 20,000 genes. RESULTS: We have developed GOParGenPy, a platform independent software tool to generate the binary data matrix showing the GO class membership, including parental classes, of a set of GO annotated genes. GOParGenPy is at least an order of magnitude faster than popular tools for Gene Ontology analysis and it can handle larger datasets than the existing tools. It can use any available version of the GO structure and allows the user to select the source of GO annotation. GO structure selection is critical for analysis, as we show that GO classes have rapid turnover between different GO structure releases. CONCLUSIONS: GOParGenPy is an easy to use software tool which can generate sparse or full binary matrices from GO annotated gene sets. The obtained binary matrix can then be used with any analysis environment and with any analysis methods.
Ajay Kumar 0020, Liisa Holm, Petri Törönen
BMC Bioinform.2
2012 SANS: high-throughput retrieval of protein sequences allowing 50% mismatches
abstract
MOTIVATION: The genomic era in molecular biology has brought on a rapidly widening gap between the amount of sequence data and first-hand experimental characterization of proteins. Fortunately, the theory of evolution provides a simple solution: functional and structural information can be transferred between homologous proteins. Sequence similarity searching followed by k-nearest neighbor classification is the most widely used tool to predict the function or structure of anonymous gene products that come out of genome sequencing projects. RESULTS: We present a novel word filter, suffix array neighborhood search (SANS), to identify protein sequence similarities in the range of 50-100% identity with sensitivity comparable to BLAST and 10 times the speed of USEARCH. In contrast to these previous approaches, the complexity of the search is proportional only to the length of the query sequence and independent of database size, enabling fast searching and functional annotation into the future despite rapidly expanding databases. AVAILABILITY AND IMPLEMENTATION: The software is freely available to non-commercial users from our website http://ekhidna.biocenter.helsinki.fi/downloads/sans. CONTACT: [email protected].
Patrik Koskinen, Liisa Holm
Bioinform.2
2012 BLANNOTATOR: enhanced homology-based function prediction of bacterial proteins
abstract
BACKGROUND: Automated function prediction has played a central role in determining the biological functions of bacterial proteins. Typically, protein function annotation relies on homology, and function is inferred from other proteins with similar sequences. This approach has become popular in bacterial genomics because it is one of the few methods that is practical for large datasets and because it does not require additional functional genomics experiments. However, the existing solutions produce erroneous predictions in many cases, especially when query sequences have low levels of identity with the annotated source protein. This problem has created a pressing need for improvements in homology-based annotation. RESULTS: We present an automated method for the functional annotation of bacterial protein sequences. Based on sequence similarity searches, BLANNOTATOR accurately annotates query sequences with one-line summary descriptions of protein function. It groups sequences identified by BLAST into subsets according to their annotation and bases its prediction on a set of sequences with consistent functional information. We show the results of BLANNOTATOR's performance in sets of bacterial proteins with known functions. We simulated the annotation process for 3090 SWISS-PROT proteins using a database in its state preceding the functional characterisation of the query protein. For this dataset, our method outperformed the five others that we tested, and the improved performance was maintained even in the absence of highly related sequence hits. We further demonstrate the value of our tool by analysing the putative proteome of Lactobacillus crispatus strain ST1. CONCLUSIONS: BLANNOTATOR is an accurate method for bacterial protein function prediction. It is practical for genome-scale data and does not require pre-existing sequence clustering; thus, this method suits the needs of bacterial genome and metagenome researchers. The method and a web-server are available at http://ekhidna.biocenter.helsinki.fi/poxo/blannotator/.
Matti Kankainen, Teija Ojala, Liisa Holm
BMC Bioinform.3
2012 Comprehensive comparison of graph based multiple protein sequence alignment strategies
abstract
BACKGROUND: Alignment of protein sequences (MPSA) is the starting point for a multitude of applications in molecular biology. Here, we present a novel MPSA program based on the SeqAn sequence alignment library. Our implementation has a strict modular structure, which allows to swap different components of the alignment process and, thus, to investigate their contribution to the alignment quality and computation time. We systematically varied information sources, guiding trees, score transformations and iterative refinement options, and evaluated the resulting alignments on BAliBASE and SABmark. RESULTS: Our results indicate the optimal alignment strategy based on the choices compared. First, we show that pairwise global and local alignments contain sufficient information to construct a high quality multiple alignment. Second, single linkage clustering is almost invariably the best algorithm to build a guiding tree for progressive alignment. Third, triplet library extension, with introduction of new edges, is the most efficient consistency transformation of those compared. Alternatively, one can apply tree dependent partitioning as a post processing step, which was shown to be comparable with the best consistency transformation in both time and accuracy. Finally, propagating information beyond four transitive links introduces more noise than signal. CONCLUSIONS: This is the first time multiple protein alignment strategies are comprehensively and clearly compared using a single implementation platform. In particular, we showed which of the existing consistency transformations and iterative refinement techniques are the most valid. Our implementation is freely available at http://ekhidna.biocenter.helsinki.fi/MMSA and as a supplementary file attached to this article (see Additional file 1).
Ilya Plyusnin, Liisa Holm
BMC Bioinform.2
2011 MPEA - metabolite pathway enrichment analysis
abstract
UNLABELLED: We present metabolite pathway enrichment analysis (MPEA) for the visualization and biological interpretation of metabolite data at the system level. Our tool follows the concept of gene set enrichment analysis (GSEA) and tests whether metabolites involved in some predefined pathway occur towards the top (or bottom) of a ranked query compound list. In particular, MPEA is designed to handle many-to-many relationships that may occur between the query compounds and metabolite annotations. For a demonstration, we analysed metabolite profiles of 14 twin pairs with differing body weights. MPEA found significant pathways from data that had no significant individual query compounds, its results were congruent with those discovered from transcriptomics data and it detected more pathways than the competing metabolic pathway method did. AVAILABILITY: The web server and source code of MPEA are available at http://ekhidna.biocenter.helsinki.fi/poxo/mpea/.
Matti Kankainen, Gopal Peddinti, Liisa Holm, Matej Oresic
Bioinform.3
2011 A novel method for assigning functional linkages to proteins using enhanced phylogenetic trees
abstract
MOTIVATION: Functional linkages implicate pairwise relationships between proteins that work together to implement biological tasks. During evolution, functionally linked proteins are likely to be preserved or eliminated across a range of genomes in a correlated fashion. Based on this hypothesis, phylogenetic profiling-based approaches try to detect pairs of protein families that show similar evolutionary patterns. Traditionally, the evolutionary pattern of a protein is encoded by either a binary profile of presence and absence of this protein across species or an occurrence profile that indicates the distribution of copies of this protein across species. RESULTS: In our study, we characterize each protein by its enhanced phylogenetic tree, a novel graphical model of the evolution of a protein family with explicitly marked by speciation and duplication events. By topological comparison between enhanced phylogenetic trees, we are able to detect the functionally associated protein pairs. Because the enhanced phylogenetic trees contain more evolutionary information of proteins, our method shows greater performance and discovers functional linkages among proteins more reliably compared with the conventional approaches.
Hung Xuan Ta, Patrik Koskinen, Liisa Holm
Bioinform.3
2009 LOCP - locating pilus operons in Gram-positive bacteria
abstract
UNLABELLED: Pilus operons encode a pivotal host-microbe interaction structure that is vital for pathogenicity, colonization and adhesion. LOCP is a computational tool to quickly test whether or not the genome of a gram-positive bacterium of interest or some DNA-contig in a metagenomic sample hold pilus operons. Predictions are made based on distinctive motifs of pilus-related protein sequences and on the tendency of these protein sequences to occur in dense clusters. The tool showed a phenomenal accuracy and revealed that various novel, and even unexpected, gram-positive bacteria do possess pilus operons. Thus, the tool helps us to focus the laboratory research on genes behind this important and indicative feature, and to screen for strains containing them. AVAILABILITY: Software is available at http://ekhidna.biocenter.helsinki.fi/locp/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ilya Plyusnin, Liisa Holm, Matti Kankainen
Bioinform.2
2009 Robust extraction of functional signals from gene set analysis using a generalized threshold free scoring function
abstract
BACKGROUND: A central task in contemporary biosciences is the identification of biological processes showing response in genome-wide differential gene expression experiments. Two types of analysis are common. Either, one generates an ordered list based on the differential expression values of the probed genes and examines the tail areas of the list for over-representation of various functional classes. Alternatively, one monitors the average differential expression level of genes belonging to a given functional class. So far these two types of method have not been combined. RESULTS: We introduce a scoring function, Gene Set Z-score (GSZ), for the analysis of functional class over-representation that combines two previous analysis methods. GSZ encompasses popular functions such as correlation, hypergeometric test, Max-Mean and Random Sets as limiting cases. GSZ is stable against changes in class size as well as across different positions of the analysed gene list in tests with randomized data. GSZ shows the best overall performance in a detailed comparison to popular functions using artificial data. Likewise, GSZ stands out in a cross-validation of methods using split real data. A comparison of empirical p-values further shows a strong difference in favour of GSZ, which clearly reports better p-values for top classes than the other methods. Furthermore, GSZ detects relevant biological themes that are missed by the other methods. These observations also hold when comparing GSZ with popular program packages. CONCLUSION: GSZ and improved versions of earlier methods are a useful contribution to the analysis of differential gene expression. The methods and supplementary material are available from the website http://ekhidna.biocenter.helsinki.fi/users/petri/public/GSZ/GSZscore.html.
Petri Törönen, Pauli J. Ojala, Pekka Marttinen, Liisa Holm
BMC Bioinform.4
2009 Generation of Gene Ontology benchmark datasets with various types of positive signal
abstract
BACKGROUND: The analysis of over-represented functional classes in a list of genes is one of the most essential bioinformatics research topics. Typical examples of such lists are the differentially expressed genes from transcriptional analysis which need to be linked to functional information represented in the Gene Ontology (GO). Despite the importance of this procedure, there is a little work on consistent evaluation of various GO analysis methods. Especially, there is no literature on creating benchmark datasets for GO analysis tools. RESULTS: We propose a methodology for the evaluation of GO analysis tools, which consists of creating gene lists with a selected signal level and a selected number of independent over-represented classes. The methodology starts with a real life GO data matrix, and therefore the generated datasets have similar features to real positive datasets. The user can select the signal level for over-representation, the number of independent positive classes in the dataset, and the size of the final gene list. We present the use of the effective number and various normalizations while embedding the signal to a selected class or classes and the use of binary correlation to ensure that the selected signal classes are independent with each other. The usefulness of generated datasets is demonstrated by comparing different GO class ranking and GO clustering methods. CONCLUSION: The presented methods aid the development and evaluation of GO analysis methods as they enable thorough testing with different signal types and different signal levels. As an example, our comparisons reveal clear differences between compared GO clustering and GO de-correlation methods. The implementation is coded in Matlab and is freely available at the dedicated website http://ekhidna.biocenter.helsinki.fi/users/petri/public/POSGODA/POSGODA.html.
Petri Törönen, Petri Pehkonen, Liisa Holm
BMC Bioinform.3
2008 Searching protein structure databases with DaliLite v.3
abstract
UNLABELLED: The Red Queen said, 'It takes all the running you can do, to keep in the same place.' Lewis Carrol MOTIVATION: Newly solved protein structures are routinely scanned against structures already in the Protein Data Bank (PDB) using Internet servers. In favourable cases, comparing 3D structures may reveal biologically interesting similarities that are not detectable by comparing sequences. The number of known structures continues to grow exponentially. Sensitive-thorough but slow-search algorithms are challenged to deliver results in a reasonable time, as there are now more structures in the PDB than seconds in a day. The brute-force solution would be to distribute the individual comparisons on a massively parallel computer. A frugal solution, as implemented in the Dali server, is to reduce the total computational cost by pruning search space using prior knowledge about the distribution of structures in fold space. This note reports paradigm revisions that enable maintaining such a knowledge base up-to-date on a PC. AVAILABILITY: The Dali server for protein structure database searching at http://ekhidna.biocenter.helsinki.fi/dali_server is running DaliLite v.3. The software can be downloaded for academic use from http://ekhidna.biocenter.helsinki.fi/dali_lite/downloads/v3.
Liisa Holm, S. Kääriäinen, Päivi Rosenström, A. Schenkel
Bioinform.1
2007 The global trace graph, a novel paradigm for searching protein sequence databases
abstract
MOTIVATION: Propagating functional annotations to sequence-similar, presumably homologous proteins lies at the heart of the bioinformatics industry. Correct propagation is crucially dependent on the accurate identification of subtle sequence motifs that are conserved in evolution. The evolutionary signal can be difficult to detect because functional sites may consist of non-contiguous residues while segments in-between may be mutated without affecting fold or function. RESULTS: Here, we report a novel graph clustering algorithm in which all known protein sequences simultaneously self-organize into hypothetical multiple sequence alignments. This eliminates noise so that non-contiguous sequence motifs can be tracked down between extremely distant homologues. The novel data structure enables fast sequence database searching methods which are superior to profile-profile comparison at recognizing distant homologues. This study will boost the leverage of structural and functional genomics and opens up new avenues for data mining a complete set of functional signature motifs. AVAILABILITY: http://www.bioinfo.biocenter.helsinki.fi/gtg. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Andreas Heger, Swapan Mallick, Christopher Andrew Wilton, Liisa Holm
Bioinform.4
2006 Bayesian search of functionally divergent protein subgroups and their function specific residues
abstract
MOTIVATION: The rapid increase in the amount of protein sequence data has created a need for an automated identification of evolutionarily related subgroups from large datasets. The existing methods typically require a priori specification of the number of putative groups, which defines the resolution of the classification solution. RESULTS: We introduce a Bayesian model-based approach to simultaneous identification of evolutionary groups and conserved parts of the protein sequences. The model-based approach provides an intuitive and efficient way of determining the number of groups from the sequence data, in contrast to the ad hoc methods often exploited for similar purposes. Our model recognizes the areas in the sequences that are relevant for the clustering and regards other areas as noise. We have implemented the method using a fast stochastic optimization algorithm which yields a clustering associated with the estimated maximum posterior probability. The method has been shown to have high specificity and sensitivity in simulated and real clustering tasks. With real datasets the method also highlights the residues close to the active site. AVAILABILITY: Software 'kPax' is available at http://www.rni.helsinki.fi/jic/softa.html
Pekka Marttinen, Jukka Corander, Petri Törönen, Liisa Holm
Bioinform.4
2005 PSIbase: a database of Protein Structural Interactome map (PSIMAP)
abstract
UNLABELLED: Protein Structural Interactome map (PSIMAP) is a global interaction map that describes domain-domain and protein-protein interaction information for known Protein Data Bank structures. It calculates the Euclidean distance to determine interactions between possible pairs of structural domains in proteins. PSIbase is a database and file server for protein structural interaction information calculated by the PSIMAP algorithm. PSIbase also provides an easy-to-use protein domain assignment module, interaction navigation and visual tools. Users can retrieve possible interaction partners of their proteins of interests if a significant homology assignment is made with their query sequences. AVAILABILITY: http://psimap.org and http://psibase.kaist.ac.kr/
Sungsam Gong, Giseok Yoon, Insoo Jang, Dan M. Bolser, Panos Dafas, Michael Schroeder 0001, Hansol Choi, Yoobok Cho, Kyungsook Han, Sunghoon Lee, Hwanho Choi, Michael Lappe, Liisa Holm, Sangsoo Kim, Donghoon Oh, Jonghwa Bhak
Bioinform.13
2003 Accurate detection of very sparse sequence motifs
abstract
Protein sequence alignments are more reliable the shorter the evolutionary distance. Here, we align distantly related proteins using many closely spaced intermediate sequences as stepping stones. Such transitive alignments can be generated between any two proteins in a connected set, whether they are direct or indirect sequence neighbours in the underlying library of pairwise alignments. We have implemented a greedy algorithm, MaxFlow, using a novel consistency score to estimate the relative likelihood of alternative paths of transitive alignment. In contrast to traditional profile models of amino acid preferences, MaxFlow models the probability that two positions are structurally equivalent and retains high information content across large distances in sequence space. Thus, MaxFlow is able to identify sparse and narrow active-site sequence signatures which are embedded in high-entropy sequence segments in the structure-based multiple alignment of large diverse enzyme superfamilies. In a challenging benchmark, MaxFlow yields better reliability and double coverage compared to available sequence alignment software. This promises to increase information returns from functional and structural genomics, where reliable sequence alignment is a bottleneck to transferring the functional or structural characterization of model proteins to entire protein families and superfamilies.
Andreas Heger, Michael Lappe, Liisa Holm
RECOMB3
2001 Picasso: generating a covering set of protein family profiles
abstract
MOTIVATION: Evolutionary classification leads to an economical description of protein sequence data because attributes of function and structure are inherited in protein families. This paper presents Picasso, a procedure for deriving a minimal set of protein family profiles that cover all known protein sequences. RESULTS: Picasso starts from highly overlapping sequence neighbourhoods revealed by all-on-all pairwise Blast alignment. Overlaps are reduced by merging sequences or parts of sequences into multiple alignments. For maximum unification, the multiple alignments must reach into the twilight zone of sequence similarity. Sensitive and selective profile-profile comparison allows unification down to about 15% pairwise sequence identity. Families unified through a short conserved sequence motif are associated with multiple full-length alignments describing different subfamilies. Domains that are mobile modules are identified based on their association with different sets of neighbours. The result is 10000 unified domain families (excluding singletons) representing functionally related proteins and recovering classical prolific domain types in high numbers. The classification is useful, for example, in developing strategies for efficient database searching and for selecting targets to complete the map of all 3-D structures.
Andreas Heger, Liisa Holm
Bioinform.2
2000 DaliLite workbench for protein structure comparison
abstract
Abstract Summary: DaliLite is a program for pairwise structure comparison and for structure database searching. It is a standalone version of the search engine of the popular Dali server. A web interface is provided to view the results, multiple alignments and 3D superimpositions of structures. Availability: DaliLite has been ported to the Linux and Irix operating systems and can be compiled in many other UNIX operating systems. It is found at http://www.embl-ebi.ac.uk/dali/DaliLite. Contact: [email protected]
Liisa Holm, Jong-Chan Park
Bioinform.1
2000 Estimating the significance of sequence order in protein secondary structure and prediction
abstract
MOTIVATION: How critical is the sequence order information in predicting protein secondary structure segments? We tried to get a rough insight on it from a theoretical approach using both a prediction algorithm and structural fragments from Protein Databank (PDB). RESULTS: Using reverse protein sequences and PDB structural fragments, we theoretically estimated the significance of the order for protein secondary structure and prediction. On average: (1) 79% of protein sequence segments resulted in the same prediction in both normal and reverse directions, which indicated a relatively high conservation of secondary structure propensity in the reverse direction; (2) the reversed sequence prediction alone performed less accurately than the normal forward sequence prediction, but comparably high (2% difference); (3) the commonly predicted regions showed a slightly higher prediction accuracy (4%) than the normal sequences prediction; and (4) structural fragments which have counterparts in reverse direction in the same protein showed a comparable degree of secondary structure conservation (73% identity with reversed structures on average for pentamers). CONTACT: [email protected]; [email protected]; [email protected]; [email protected]
Jong-Chan Park, Sabine Dietmann, Andreas Heger, Liisa Holm
Bioinform.4
2000 Sequence search algorithm assessment and testing toolkit (SAT)
abstract
MOTIVATION: The Sequence Search Algorithm Assessment and Testing Toolkit (SAT) aims to be a complete package for the comparison of different protein homology search algorithms. The structural classification of proteins can provide us with a clear criterion for judgment in homology detection. There have been several assessments based on structural sequences with classifications but a good deal of similar work is now being repeated with locally developed procedures and programs. The SAT will provide developers with a complete package which will save time and produce more comparable performance assessments for search algorithms. The package is complete in the sense that it provides a non-redundant large sequence resource database, a well-characterized query database of proteins domains, all the parsers and some previous results from PSI-BLAST and a hidden markov model algorithm. RESULTS: An analysis on two different data sets was carried out using the SAT package. It compared the performance of a full protein sequence database (RSDB100) with a non-redundant representative sequence database derived from it (RSDB50). The performance measurement indicated that the full database is sub-optimal for a homology search. This result justifies the use of much smaller and faster RSDB50 than RSDB100 for the SAT. AVAILABILITY: A web site is up. The whole packa ge is accessible via www and ftp. ftp://ftp.ebi.ac.uk/pub/contrib/jong/SAT http://cyrah.ebi.ac.uk:1111/Proj/Bio/SAT http://www.mrc-lmb.cam.ac.uk/genomes/SAT In the package, some previous assessment results produced by the package can also be found for reference. CONTACT: [email protected]
Jong-Chan Park, Liisa Holm, Cyrus Chothia
Bioinform.2
2000 RSDB: representative protein sequence databases have high information content
abstract
MOTIVATION: Biological sequence databases are highly redundant for two main reasons: 1. various databanks keep redundant sequences with many identical and nearly identical sequences 2. natural sequences often have high sequence identities due to gene duplication. We wanted to know how many sequences can be removed before the databases start losing homology information. Can a database of sequences with mutual sequence identity of 50% or less provide us with the same amount of biological information as the original full database? RESULTS: Comparisons of nine representative sequence databases (RSDB) derived from full protein databanks showed that the information content of sequence databases is not linearly proportional to its size. An RSDB reduced to mutual sequence identity of around 50% (RSDB50) was equivalent to the original full database in terms of the effectiveness of homology searching. It was a third of the full database size which resulted in a six times faster iterative profile searching. The RSDBs are produced at different granularity for efficient homology searching. AVAILABILITY: All the RSDB files generated and the full analysis results are available through internet: ftp://ftp.ebi.ac. uk/pub/contrib/jong/RSDB/http://cyrah.e bi.ac.uk:1111/Proj/Bio/RSDB
Jong-Chan Park, Liisa Holm, Andreas Heger, Cyrus Chothia
Bioinform.2
1998 Removing near-neighbour redundancy from large protein sequence collections
abstract
MOTIVATION: To maximize the chances of biological discovery, homology searching must use an up-to-date collection of sequences. However, the available sequence databases are growing rapidly and are partially redundant in content. This leads to increasing strain on CPU resources and decreasing density of first-hand annotation. RESULTS: These problems are addressed by clustering closely similar sequences to yield a covering of sequence space by a representative subset of sequences. No pair of sequences in the representative set has >90% mutual sequence identity. The representative set is derived by an exhaustive search for close similarities in the sequence database in which the need for explicit sequence alignment is significantly reduced by applying deca- and pentapeptide composition filters. The algorithm was applied to the union of the Swissprot, Swissnew, Trembl, Tremblnew, Genbank, PIR, Wormpep and PDB databases. The all-against-all comparison required to generate a representative set at 90% sequence identity was accomplished in 2 days CPU time, and the removal of fragments and close similarities yielded a size reduction of 46%, from 260 000 unique sequences to 140 000 representative sequences. The practical implications are (i) faster homology searches using, for example, Fasta or Blast, and (ii) unified annotation for all sequences clustered around a representative. As tens of thousands of sequence searches are performed daily world-wide, appropriate use of the non-redundant database can lead to major savings in computer resources, without loss of efficacy. AVAILABILITY: A regularly updated non-redundant protein sequence database (nrdb90), a server for homology searches against nrdb90, and a Perl script (nrdb90.pl) implementing the algorithm are available for academic use from http://www.embl-ebi.ac. uk/holm/nrdb90. CONTACT: [email protected]
Liisa Holm, Chris Sander
Bioinform.1
1998 COFFEE: an objective function for multiple sequence alignments
abstract
MOTIVATION: In order to increase the accuracy of multiple sequence alignments, we designed a new strategy for optimizing multiple sequence alignments by genetic algorithm. We named it COFFEE (Consistency based Objective Function For alignmEnt Evaluation). The COFFEE score reflects the level of consistency between a multiple sequence alignment and a library containing pairwise alignments of the same sequences. RESULTS: We show that multiple sequence alignments can be optimized for their COFFEE score with the genetic algorithm package SAGA. The COFFEE function is tested on 11 test cases made of structural alignments extracted from 3D_ali. These alignments are compared to those produced using five alternative methods. Results indicate that COFFEE outperforms the other methods when the level of identity between the sequences is low. Accuracy is evaluated by comparison with the structural alignments used as references. We also show that the COFFEE score can be used as a reliability index on multiple sequence alignments. Finally, we show that given a library of structure-based pairwise sequence alignments extracted from FSSP, SAGA can produce high-quality multiple sequence alignments. The main advantage of COFFEE is its flexibility. With COFFEE, any method suitable for making pairwise alignments can be extended to making multiple alignments. AVAILABILITY: The package is available along with the test cases through the WWW: http://www. ebi.ac.uk/cedric CONTACT: [email protected]
Cédric Notredame, Liisa Holm, Desmond G. Higgins
Bioinform.2
1997 Decision Support System for the Evolutionary Classification of Protein Structures
Liisa Holm, Chris Sander
ISMB1
1995 3-D Lookup: Fast Protein Structure Database Searches at 90% Reliability
Liisa Holm, Chris Sander
ISMB1