EDBT 2026 Demo / reviewers in the wild / expert
Geir Kjetil Sandve
dblp:62/4526 · also Geir Kjetil Ferkingstad Sandve
· DBLP profile ↗
32ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0002-4959-1409ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 29 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | inMOTIFin: a lightweight end-to-end simulation software for regulatory sequencesabstractSUMMARY: The accurate development, assessment, interpretation, and benchmarking of bioinformatics frameworks for analyzing transcriptional regulatory grammars rely on controlled simulations to validate the underlying methods. However, existing simulators often lack end-to-end flexibility or ease of integration, which limits their practical use. We present inMOTIFin, a lightweight, modular, and user-friendly Python-based software that addresses these gaps by providing versatile and efficient simulation and modification of DNA regulatory sequences. inMOTIFin enables users to simulate or modify regulatory sequences efficiently for the customizable generation of motifs and insertion of motif instances with precise control over their positions, co-occurrences, and spacing, as well as direct modification of real sequences, facilitating a comprehensive evaluation of motif-based methods and interpretation tools. We demonstrate inMOTIFin applications for the assessment of de novo motif discovery, the analysis of transcription factor cooperativity, and the support of explainability analyses for deep learning models. inMOTIFin ensures robust and reproducible analyses for studying transcriptional regulatory grammars. AVAILABILITY AND IMPLEMENTATION: inMOTIFin is available at PyPI https://pypi.org/project/inMOTIFin/ and Docker Hub https://hub.docker.com/r/cbgr/inmotifin. Detailed documentation is available at https://inmotifin.readthedocs.io/en/latest/. The code for use case analyses is available at https://bitbucket.org/CBGR/inmotifin_evaluation/src/main/. The version of the code used for this article has been uploaded to Zenodo with DOI: 10.5281/zenodo.17638579. Katalin Ferenc, Lorenzo Martini, Ieva Rauluseviciute, Geir Kjetil Sandve, Anthony Mathelier |
Bioinform. | 4 |
| 2026 | PEPE: scalable extraction of multi-modal protein language model representationsabstractSUMMARY: Protein language models (PLMs) capture intricate amino-acid dependencies, producing embeddings that encode rich structural, functional, and evolutionary information. Despite their potential, current extraction workflows rely on arbitrary choices, with respect to embedding layer, pooling, and padding, that frequently yield suboptimal representations for feature extraction and downstream analyses. Large-scale embedding generation is further limited by inefficiencies in computation and memory: (i) accumulating all model outputs in memory before writing to disk causes severe bottlenecks, and (ii) repeatedly embedding identical sequences to extract different modes introduces redundant computation and drastically reduces throughput and scalability. We introduce PEPE (Parallel Extraction for Protein Embeddings), a command-line tool and Python library that enables efficient, high-throughput, and multimodal extraction from protein language models. PEPE's parallelized and streaming-based architecture achieves runtimes several orders of magnitude faster than sequential approaches. Unlike conventional methods-whose peak memory usage scales linearly with output size and fails when memory capacity is exceeded-PEPE maintains stable, low memory consumption, enabling multimodal embedding extraction even beyond available RAM. PEPE supports a wide range of state-of-the-art and custom PLMs through a simple, flexible interface. By combining scalability, robustness, and ease of use, PEPE allows researchers to generate massive, information-rich embedding datasets efficiently, and facilitate the discovery of optimal representations for structural, functional, and evolutionary downstream tasks. By streamlining the generation of diverse embedding configurations, PEPE provides researchers with the necessary data to identify high-performing latent states for specific biological contexts without requiring additional computational resources. AVAILABILITY AND IMPLEMENTATION: PEPE is a command-line tool written in Python and published under MIT license. The source code and documentation are available at https://github.com/csi-greifflab/pepe-cli. PEPE is also available for installation from PyPI under https://pypi.org/project/pepe-cli and deposited on Zenodo at https://zenodo.org/records/20268104. Jahn Zhong, Niccolò Cardente, Geir Kjetil Sandve, Habib Bashour, Maria Francesca Abbate, Victor Greiff |
Bioinform. | 3 |
| 2024 | Incorporating probabilistic domain knowledge into deep multiple instance learningabstractDeep learning methods, including deep multiple instance learning methods, have been criticized for their limited ability to incorporate domain knowledge. A reason that knowledge incorporation is challenging in deep learning is that the models usually lack a mapping between their model components and the entities of the domain, making it a non-trivial task to incorporate probabilistic prior information. In this work, we show that such a mapping between domain entities and model components can be defined for a multiple instance learning setting and propose a framework DeeMILIP that encompasses multiple strategies to exploit this mapping for prior knowledge incorporation. We motivate and formalize these strategies from a probabilistic perspective. Experiments on an immune-based diagnostics case show that our proposed strategies allow to learn generalizable models even in settings with weak signals, limited dataset size, and limited compute. Ghadi S. Al Hajj, Aliaksandr Hubin, Chakravarthi Kanduri, Milena Pavlovic, Knut D. Rand, Michael Widrich, Anne H. Schistad Solberg, Victor Greiff, Johan Pensar, Günter Klambauer, Geir Kjetil Sandve |
ICML | 11 |
| 2024 | Predictability of antigen binding based on short motifs in the antibody CDRH3abstractAdaptive immune receptors, such as antibodies and T-cell receptors, recognize foreign threats with exquisite specificity. A major challenge in adaptive immunology is discovering the rules governing immune receptor-antigen binding in order to predict the antigen binding status of previously unseen immune receptors. Many studies assume that the antigen binding status of an immune receptor may be determined by the presence of a short motif in the complementarity determining region 3 (CDR3), disregarding other amino acids. To test this assumption, we present a method to discover short motifs which show high precision in predicting antigen binding and generalize well to unseen simulated and experimental data. Our analysis of a mutagenesis-based antibody dataset reveals 11 336 position-specific, mostly gapped motifs of 3-5 amino acids that retain high precision on independently generated experimental data. Using a subset of only 178 motifs, a simple classifier was made that on the independently generated dataset outperformed a deep learning model proposed specifically for such datasets. In conclusion, our findings support the notion that for some antibodies, antigen binding may be largely determined by a short CDR3 motif. As more experimental data emerge, our methodology could serve as a foundation for in-depth investigations into antigen binding signals. Lonneke Scheffer, Eric Emanuel Reber, Brij Bhushan Mehta, Milena Pavlovic, Maria Chernigovskaya, Eve Richardson, Rahmad Akbar, Fridtjof Lund-Johansen, Victor Greiff, Ingrid Hobæk Haff, Geir Kjetil Sandve |
Briefings Bioinform. | 11 |
| 2023 | Adjustment of spurious correlations in co-expression measurements from RNA-Sequencing dataabstractMOTIVATION: Gene co-expression measurements are widely used in computational biology to identify coordinated expression patterns across a group of samples. Coordinated expression of genes may indicate that they are controlled by the same transcriptional regulatory program, or involved in common biological processes. Gene co-expression is generally estimated from RNA-Sequencing data, which are commonly normalized to remove technical variability. Here, we demonstrate that certain normalization methods, in particular quantile-based methods, can introduce false-positive associations between genes. These false-positive associations can consequently hamper downstream co-expression network analysis. Quantile-based normalization can, however, be extremely powerful. In particular, when preprocessing large-scale heterogeneous data, quantile-based normalization methods such as smooth quantile normalization can be applied to remove technical variability while maintaining global differences in expression for samples with different biological attributes. RESULTS: We developed SNAIL (Smooth-quantile Normalization Adaptation for the Inference of co-expression Links), a normalization method based on smooth quantile normalization specifically designed for modeling of co-expression measurements. We show that SNAIL avoids formation of false-positive associations in co-expression as well as in downstream network analyses. Using SNAIL, one can avoid arbitrary gene filtering and retain associations to genes that only express in small subgroups of samples. This highlights the method's potential future impact on network modeling and other association-based approaches in large-scale heterogeneous data. AVAILABILITY AND IMPLEMENTATION: The implementation of the SNAIL algorithm and code to reproduce the analyses described in this work can be found in the GitHub repository https://github.com/kuijjerlab/PySNAIL. Ping-Han Hsieh, Camila Miranda Lopes-Ramos, Manuela Zucknick, Geir Kjetil Sandve, Kimberly Glass, Marieke L. Kuijjer |
Bioinform. | 4 |
| 2022 | TCRpower: quantifying the detection power of T-cell receptor sequencing with a novel computational pipeline calibrated by spike-in sequencesabstractT-cell receptor (TCR) sequencing has enabled the development of innovative diagnostic tests for cancers, autoimmune diseases and other applications. However, the rarity of many T-cell clonotypes presents a detection challenge, which may lead to misdiagnosis if diagnostically relevant TCRs remain undetected. To address this issue, we developed TCRpower, a novel computational pipeline for quantifying the statistical detection power of TCR sequencing methods. TCRpower calculates the probability of detecting a TCR sequence as a function of several key parameters: in-vivo TCR frequency, T-cell sample count, read sequencing depth and read cutoff. To calibrate TCRpower, we selected unique TCRs of 45 T-cell clones (TCCs) as spike-in TCRs. We sequenced the spike-in TCRs from TCCs, together with TCRs from peripheral blood, using a 5' RACE protocol. The 45 spike-in TCRs covered a wide range of sample frequencies, ranging from 5 per 100 to 1 per 1 million. The resulting spike-in TCR read counts and ground truth frequencies allowed us to calibrate TCRpower. In our TCR sequencing data, we observed a consistent linear relationship between sample and sequencing read frequencies. We were also able to reliably detect spike-in TCRs with frequencies as low as one per million. By implementing an optimized read cutoff, we eliminated most of the falsely detected sequences in our data (TCR α-chain 99.0% and TCR β-chain 92.4%), thereby improving diagnostic specificity. TCRpower is publicly available and can be used to optimize future TCR sequencing experiments, and thereby enable reliable detection of disease-relevant TCRs for diagnostic applications. Shiva Dahal-Koirala, Gabriel Balaban, Ralf Stefan Neumann, Lonneke Scheffer, Knut Erik Aslaksen Lundin, Victor Greiff, Ludvig Magne Sollid, Shuo-Wang Qiao, Geir Kjetil Sandve |
Briefings Bioinform. | 9 |
| 2022 | CompAIRR: ultra-fast comparison of adaptive immune receptor repertoires by exact and approximate sequence matchingabstractMOTIVATION: Adaptive immune receptor (AIR) repertoires (AIRRs) record past immune encounters with exquisite specificity. Therefore, identifying identical or similar AIR sequences across individuals is a key step in AIRR analysis for revealing convergent immune response patterns that may be exploited for diagnostics and therapy. Existing methods for quantifying AIRR overlap scale poorly with increasing dataset numbers and sizes. To address this limitation, we developed CompAIRR, which enables ultra-fast computation of AIRR overlap, based on either exact or approximate sequence matching. RESULTS: CompAIRR improves computational speed 1000-fold relative to the state of the art and uses only one-third of the memory: on the same machine, the exact pairwise AIRR overlap of 104 AIRRs with 105 sequences is found in ∼17 min, while the fastest alternative tool requires 10 days. CompAIRR has been integrated with the machine learning ecosystem immuneML to speed up commonly used AIRR-based machine learning applications. AVAILABILITY AND IMPLEMENTATION: CompAIRR code and documentation are available at https://github.com/uio-bmi/compairr. Docker images are available at https://hub.docker.com/r/torognes/compairr. The code to replicate the synthetic datasets, scripts for benchmarking and creating figures, and all raw data underlying the figures are available at https://github.com/uio-bmi/compairr-benchmarking. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Torbjørn Rognes, Lonneke Scheffer, Victor Greiff, Geir Kjetil Sandve |
Bioinform. | 4 |
| 2022 | Access to ground truth at unconstrained size makes simulated data as indispensable as experimental data for bioinformatics methods development and benchmarkingabstractThe author instructions of this journal (OUP Bioinformatics) include several detailed requirements for manuscripts to be considered for publication (https://academic.oup.com/bioinformatics/pages/instructions_for_authors). Such requirements clarify expectations for authors and establish important peer-review standards. Importantly, it also allows for an open discussion of these requirements, which for a leading journal such as Bioinformatics contributes to shaping the fields of computational biology and bioinformatics. One of the requirements is that manuscripts presenting new methodology must include ‘actual biological data’ as opposed to simulated data. We find this requirement potentially counterproductive for several reasons and argue for emphasizing the complementarity of simulated and experimental data. Before we outline our argument, we would like to remark that terminology has connotations that may influence how different types of scientific evidence are valued. Specifically, we find the term ‘actual biological data’ to be problematic because primary data in the biological research domain is generated either from wet-lab experiments or computational simulations, both of which have their idiosyncrasies relative to the underlying biology (Leek et al., 2010). The aim, and often the very purpose, of computational methodology in bioinformatics is to model biological phenomena at a resolution and scale that transcends the current experimental state-of-the-art, to prepare for the arrival of biological data generation at a scale that allows better coverage of biological phenomena. Thus, a term such as ‘experimental calibration’ may be more appropriate to describe the use and purpose of experimental data in the majority of the bioinformatics literature, as opposed to the currently predominant denomination of ‘experimental validation’ (Jafari et al., 2021). Here, we argue that in the majority of bioinformatics settings, available experimental data do not have the size, resolution and sufficient set of controls (hereafter referred to as ‘limited data’) that would allow for rigorous method assessment. With limited data, performance estimates may be uncertain and sensitive to external factors such as parameter choices. This makes it challenging to judge whether observed improvements over previous methods are substantial, that is, biologically relevant, or merely the result of deliberate tuning of a method to perform particularly well on the experimental dataset(s) at hand (Castaldi et al., 2011; Salzberg, 1997). As a reviewer or critical reader, it is usually unfeasible to generate corroborating (or falsifying) experimental data with similar properties and thus not possible to rule out chance or tuning. Therefore, we argue that increased emphasis on experimental data may lead to insufficient and potentially misleading method evaluation. In contrast, simulation enables the generation of datasets of virtually unconstrained size, with precise control over introduced signals (ground truth) (Morris et al., 2019). This confers a critical reader the competence to challenge a reported assessment (Meyer and Birney, 2018) and rule out chance results and inappropriate tuning by simply generating new data from the same simulation process, thereby ensuring that conclusions can be meaningfully reproduced and cover a biologically relevant parameter range. Additionally, the specification of a simulation algorithm makes data assumptions for a method explicit and thus contributes to the transparency of a method both in terms of advances over the state of the art as well as its limitations. For example, in immunoinformatics of adaptive immunity, natural immune receptor sequence diversity, which is of the order of >1013, is routinely modeled using simulation frameworks for testing biological assumptions and the benchmarking of novel methods (Davidsen et al., 2019; Marcou et al., 2018; Pavlovic et al., 2021; Safonova et al., 2015; Weber et al., 2020). We stress that simulated data are only meaningful for bioinformatics method development and assessment if it reflects method-relevant underlying biology. The same criterion should be applied to experimental data. We agree that, if available at a sufficient scale, resolution and quality, experimental data are unsurpassed for assessing the capacity of bioinformatics methods to handle the types of signal complexities and data distributions that distinguishes bioinformatics method development from general informatics. Thus, novel methods should indeed be required to be evaluated on experimental data in domains where the available data are sufficiently robust to admit a rigorous assessment. However, we have several concerns with the quality of typically available experimental data for assessment purposes. A first concern is one of size. The performance of a method on a given dataset is always an uncertain estimate of its true performance on the underlying distribution that the observed data reflects. For many biological problems, available experimental datasets are so small that estimate uncertainties can easily be larger than any performance differences observed between competing methods. We fear that a strict journal requirement of employing experimental data may push authors to draw unwarranted conclusions from too small datasets and that reviewers may allow this to pass through due to the lack of good alternatives for the authors. An author’s requirement to include at least rudimentary measures of uncertainty for any reported performance measurement could alleviate these concerns (Walsh et al., 2021). A second concern is that experimental data are often available only for one particular problem setup. This leaves no opportunities to test the sensitivity of a method to variation in problem configuration (to assess how broadly it generalizes) or to test how it performs on data outside the training distribution (whether it is robust to domain shift). A third concern is that there is no possibility to know whether the patterns that a method extracts from experimental data reflect underlying causal relations or not. Furthermore, suboptimal study designs may introduce spurious correlations in datasets, and there is a risk that the best-performing methods at least partly exploit such artificial data patterns. Also, we fear that a strict, general requirement to assess novel methodology on experimental data may impede progress in method development in the many domains where available data are scarce. In data-scarce domains, we hold that authors should instead be urged to provide a rigorous assessment on simulated data, where authors should explicitly argue for the biological relevance based on either underlying mechanistic knowledge or by calibrating their simulation procedure with experimental data. In particular, we consider such experimentally calibrated simulation to often provide a better assessment of the capabilities of a method than the (in our opinion) too common reliance on anecdotal findings on small experimental datasets, which may reveal more about the ingenuity of the authors than of the proposed method. The discovery of novel biological knowledge does not in itself establish the usefulness or novelty of a new bioinformatics method—it is merely a corollary of, for example, a new method’s greater sensitivity, applicability to a wider parameter range or scalability. While it may be tempting to boost impact by combining novel methodology and novel biological findings in the same paper, this interferes with the assessment of methods on their own merit and thus undermines the selection pressures for the evolutionary process of method improvement in the field. Our view on the complementarity of simulated and experimental data is in line with the approach to method assessment taken in the machine learning field. Here, the evaluation of simulated data has always had a prominent role. Nowadays, the availability of large and well-curated databases such as ImageNet (Deng et al., 2009) and MNIST (Deng, 2012) make it natural to expect that novel methodologies are also assessed in such real-world data collections. However, when for instance the long short-term memory model was introduced in 1997 (Hochreiter and Schmidhuber, 1997), the authors explicitly asked in their paper ‘which tasks are appropriate to demonstrate the quality of a novel long-time-lag algorithm’ and answered their question based on a collection of exclusively synthetic datasets. Years later, improved data availability revealed that the model is indeed able to learn relevant patterns in a wide variety of real-world domains. The top-cited paper of the present journal (according to ISI web of science) (Li and Durbin, 2009) contains two sections in the Results section entitled ‘Evaluation on simulated data’ and ‘Evaluation on real data’. While we suggest that well-argued exceptions to the inclusion of experimental data assessment should be allowed for bioinformatics methods research, we can hardly think of any circumstance with compelling reasons for not including any assessment on simulated data. Since method developers should always have a conscious relationship to the data assumptions that they build their models and algorithms on, it should usually be straightforward to implement a simulation of data according to these same assumptions. This allows method developers to confirm that their method behaves as expected (e.g. identification of any bugs), it reveals to developers and readers the range of data parameters within which the method provides sensible results (method transparency) and allows developers or readers to reproduce assessments under identical or modified data assumptions (method reproducibility). We thus encourage basic assessment on simulated data to be considered an integral part of good bioinformatics method craftsmanship. Once simulated data have shown that a method works as intended, experimental data may be used to show that the software works on the field-specific experimental data formats and, ideally, recovers orthogonally validated biological or technological signals. We have argued that simulated and experimental data should be considered complementary and of equal importance for assessing methods in typical bioinformatics settings. They should both be strongly encouraged as part of a rigorous review process of novel methodology, where reviewers should ensure that a given paper exploits the best available data sources for assessment (be it experimental or well-established simulated datasets) and when necessary combines data sources for a comprehensive assessment. When available in high quantity, fidelity and generality, experimental data may ensure assessment validity—that a method handles relevant signals and noise profiles from the biological domain. But for many bioinformatics application areas, experimental data are not available at sufficient scale or annotation quality to allow conclusive assessment. Through full control over ground truth and unconstrained data size, simulated data may ensure assessment reliability—that the reported performance of a given method is representative and can be reproduced under the same or modified assumptions of the underlying data generating process. Importantly, sophisticated simulation processes, where signals and noise are calibrated by experimental data or knowledge of underlying mechanisms (Cao et al., 2021; Prakash et al., 2021; Schuler et al., 2017), allows methodology to be developed, assessed and improved early in a field so as to reach a good level of maturity at the time large-scale experimental data starts to become available. Well-calibrated simulation schemes may even be used to explore targeted hypotheses relating to complex biological systems in a way that can guide future experimental data collection (Azencott et al., 2017). In addition, simulation makes explicit the assumptions and layers of biological complexity understood so far and helps identify methodological errors or software bugs. In summary, we suggest that new bioinformatics methods should be shown to perform comparatively well on ground truth data of a size that allows reliable assessment, be it experimental or simulated. Method developers should be encouraged to make use of both simulated and experimental data, in complementary ways, to cover the multiple purposes of method assessment. When certain roles of assessments are not fully covered, method developers should be expected to provide compelling, explicit reasons—be it reasons for not including assessments involving simulated or experimental data. We would like to thank Michael Widrich and Günter Klambauer for their helpful suggestions. This work was supported by the Research Council of Norway [IKTPLUSS project (#311341 to G.K.S. and V.G.)]. Conflict of Interest: V.G. declares advisory board positions in aiNET GmbH, Enpicom B.V, Specifica Inc, Adaptyv Biosystems and EVQLV. V.G. is a consultant for Roche/Genentech. No new data were generated or analyzed in support of this research. Geir Kjetil Sandve, Victor Greiff |
Bioinform. | 1 |
| 2021 | Ten simple rules for quick and dirty scientific programming
Gabriel Balaban, Ivar Grytten, Knut D. Rand, Lonneke Scheffer, Geir Kjetil Sandve |
PLoS Comput. Biol. | 5 |
| 2020 | Modern Hopfield Networks and Attention for Immune Repertoire ClassificationabstractA central mechanism in machine learning is to identify, store, and recognize patterns. How to learn, access, and retrieve such patterns is crucial in Hopfield networks and the more recent transformer architectures. We show that the attention mechanism of transformer architectures is actually the update rule of modern Hopfield networks that can store exponentially many patterns. We exploit this high storage capacity of modern Hopfield networks to solve a challenging multiple instance learning (MIL) problem in computational biology: immune repertoire classification. In immune repertoire classification, a vast number of immune receptors are used to predict the immune status of an individual. This constitutes a MIL problem with an unprecedentedly massive number of instances, two orders of magnitude larger than currently considered problems, and with an extremely low witness rate. Accurate and interpretable machine learning methods solving this problem could pave the way towards new vaccines and therapies, which is currently a very relevant research topic intensified by the COVID-19 crisis. In this work, we present our novel method DeepRC that integrates transformer-like attention, or equivalently modern Hopfield networks, into deep learning architectures for massive MIL such as immune repertoire classification. We demonstrate that DeepRC outperforms all other methods with respect to predictive performance on large-scale experiments including simulated and real-world virus infection data and enables the extraction of sequence motifs that are connected to a given disease class. Source code and datasets: https://github.com/ml-jku/DeepRC Michael Widrich, Bernhard Schäfl, Milena Pavlovic, Hubert Ramsauer, Lukas Gruber, Markus Holzleitner, Johannes Brandstetter, Geir Kjetil Sandve, Victor Greiff, Sepp Hochreiter, Günter Klambauer |
NeurIPS | 8 |
| 2020 | Beware the Jaccard: the choice of similarity measure is important and non-trivial in genomic colocalisation analysisabstractThe generation and systematic collection of genome-wide data is ever-increasing. This vast amount of data has enabled researchers to study relations between a variety of genomic and epigenomic features, including genetic variation, gene regulation and phenotypic traits. Such relations are typically investigated by comparatively assessing genomic co-occurrence. Technically, this corresponds to assessing the similarity of pairs of genome-wide binary vectors. A variety of similarity measures have been proposed for this problem in other fields like ecology. However, while several of these measures have been employed for assessing genomic co-occurrence, their appropriateness for the genomic setting has never been investigated. We show that the choice of similarity measure may strongly influence results and propose two alternative modelling assumptions that can be used to guide this choice. On both simulated and real genomic data, the Jaccard index is strongly altered by dataset size and should be used with caution. The Forbes coefficient (fold change) and tetrachoric correlation are less influenced by dataset size, but one should be aware of increased variance for small datasets. All results on simulated and real data can be inspected and reproduced at https://hyperbrowser.uio.no/sim-measure. Stefania Salvatore, Knut D. Rand, Ivar Grytten, Egil Ferkingstad, Diana Domanska, Lars Holden, Marius Gheorghe, Anthony Mathelier, Ingrid Kristine Glad, Geir Kjetil Sandve |
Briefings Bioinform. | 10 |
| 2020 | immuneSIM: tunable multi-feature simulation of B- and T-cell receptor repertoires for immunoinformatics benchmarkingabstractSUMMARY: B- and T-cell receptor repertoires of the adaptive immune system have become a key target for diagnostics and therapeutics research. Consequently, there is a rapidly growing number of bioinformatics tools for immune repertoire analysis. Benchmarking of such tools is crucial for ensuring reproducible and generalizable computational analyses. Currently, however, it remains challenging to create standardized ground truth immune receptor repertoires for immunoinformatics tool benchmarking. Therefore, we developed immuneSIM, an R package that allows the simulation of native-like and aberrant synthetic full-length variable region immune receptor sequences by tuning the following immune receptor features: (i) species and chain type (BCR, TCR, single and paired), (ii) germline gene usage, (iii) occurrence of insertions and deletions, (iv) clonal abundance, (v) somatic hypermutation and (vi) sequence motifs. Each simulated sequence is annotated by the complete set of simulation events that contributed to its in silico generation. immuneSIM permits the benchmarking of key computational tools for immune receptor analysis, such as germline gene annotation, diversity and overlap estimation, sequence similarity, network architecture, clustering analysis and machine learning methods for motif detection. AVAILABILITY AND IMPLEMENTATION: The package is available via https://github.com/GreiffLab/immuneSIM and on CRAN at https://cran.r-project.org/web/packages/immuneSIM. The documentation is hosted at https://immuneSIM.readthedocs.io. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Cédric R. Weber, Rahmad Akbar, Alexander Yermanos, Milena Pavlovic, Igor Snapkov, Geir Kjetil Sandve, Sai T. Reddy, Victor Greiff |
Bioinform. | 6 |
| 2020 | NucBreak: location of structural errors in a genome assembly by using paired-end Illumina readsabstractBACKGROUND: Advances in whole genome sequencing strategies have provided the opportunity for genomic and comparative genomic analysis of a vast variety of organisms. The analysis results are highly dependent on the quality of the genome assemblies used. Assessment of the assembly accuracy may significantly increase the reliability of the analysis results and is therefore of great importance. RESULTS: Here, we present a new tool called NucBreak aimed at localizing structural errors in assemblies, including insertions, deletions, duplications, inversions, and different inter- and intra-chromosomal rearrangements. The approach taken by existing alternative tools is based on analysing reads that do not map properly to the assembly, for instance discordantly mapped reads, soft-clipped reads and singletons. NucBreak uses an entirely different and unique method to localise the errors. It is based on analysing the alignments of reads that are properly mapped to an assembly and exploit information about the alternative read alignments. It does not annotate detected errors. We have compared NucBreak with other existing assembly accuracy assessment tools, namely Pilon, REAPR, and FRCbam as well as with several structural variant detection tools, including BreakDancer, Lumpy, and Wham, by using both simulated and real datasets. CONCLUSIONS: The benchmarking results have shown that NucBreak in general predicts assembly errors of different types and sizes with relatively high sensitivity and with lower false discovery rate than the other tools. Such a balance between sensitivity and false discovery rate makes NucBreak a good alternative to the existing assembly accuracy assessment tools and SV detection tools. NucBreak is freely available at https://github.com/uio-bmi/NucBreak under the MPL license. Ksenia Khelik, Geir Kjetil Sandve, Alexander Johan Nederbragt, Torbjørn Rognes |
BMC Bioinform. | 2 |
| 2019 | Colocalization analyses of genomic elements: approaches, recommendations and challengesabstractMOTIVATION: Many high-throughput methods produce sets of genomic regions as one of their main outputs. Scientists often use genomic colocalization analysis to interpret such region sets, for example to identify interesting enrichments and to understand the interplay between the underlying biological processes. Although widely used, there is little standardization in how these analyses are performed. Different practices can substantially affect the conclusions of colocalization analyses. RESULTS: Here, we describe the different approaches and provide recommendations for performing genomic colocalization analysis, while also discussing common methodological challenges that may influence the conclusions. As illustrated by concrete example cases, careful attention to analysis details is needed in order to meet these challenges and to obtain a robust and biologically meaningful interpretation of genomic region set data. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chakravarthi Kanduri, Christoph Bock, Sveinung Gundersen, Eivind Hovig, Geir Kjetil Sandve |
Bioinform. | 5 |
| 2019 | Graph Peak Caller: Calling ChIP-seq peaks on graph-based reference genomesabstractGraph-based representations are considered to be the future for reference genomes, as they allow integrated representation of the steadily increasing data on individual variation. Currently available tools allow de novo assembly of graph-based reference genomes, alignment of new read sets to the graph representation as well as certain analyses like variant calling and haplotyping. We here present a first method for calling ChIP-Seq peaks on read data aligned to a graph-based reference genome. The method is a graph generalization of the peak caller MACS2, and is implemented in an open source tool, Graph Peak Caller. By using the existing tool vg to build a pan-genome of Arabidopsis thaliana, we validate our approach by showing that Graph Peak Caller with a pan-genome reference graph can trace variants within peaks that are not part of the linear reference genome, and find peaks that in general are more motif-enriched than those found by MACS2. Ivar Grytten, Knut D. Rand, Alexander Johan Nederbragt, Geir Storvik, Ingrid Kristine Glad, Geir Kjetil Sandve |
PLoS Comput. Biol. | 6 |
| 2018 | Mind the gaps: overlooking inaccessible regions confounds statistical testing in genome analysisabstractBACKGROUND: The current versions of reference genome assemblies still contain gaps represented by stretches of Ns. Since high throughput sequencing reads cannot be mapped to those gap regions, the regions are depleted of experimental data. Moreover, several technology platforms assay a targeted portion of the genomic sequence, meaning that regions from the unassayed portion of the genomic sequence cannot be detected in those experiments. We here refer to all such regions as inaccessible regions, and hypothesize that ignoring these regions in the null model may increase false findings in statistical testing of colocalization of genomic features. RESULTS: Our explorative analyses confirm that the genomic regions in public genomic tracks intersect very little with assembly gaps of human reference genomes (hg19 and hg38). The little intersection was observed only at the beginning and end portions of the gap regions. Further, we simulated a set of synthetic tracks by matching the properties of real genomic tracks in a way that nullified any true association between them. This allowed us to test our hypothesis that not avoiding inaccessible regions (as represented by assembly gaps) in the null model would result in spurious inflation of statistical significance. We contrasted the distributions of test statistics and p-values of Monte Carlo-based permutation tests that either avoided or did not avoid assembly gaps in the null model when testing colocalization between a pair of tracks. We observed that the statistical tests that did not account for assembly gaps in the null model resulted in a distribution of the test statistic that is shifted to the right and a distribution of p-values that is shifted to the left (indicating inflated significance). We observed a similar level of inflated significance in hg19 and hg38, despite assembly gaps covering a smaller proportion of the latter reference genome. CONCLUSION: We provide empirical evidence demonstrating that inaccessible regions, even when covering only a few percentages of the genome, can lead to a substantial amount of false findings if not accounted for in statistical colocalization analysis. Diana Domanska, Chakravarthi Kanduri, Boris Simovski, Geir Kjetil Sandve |
BMC Bioinform. | 4 |
| 2017 | The rainfall plot: its motivation, characteristics and pitfallsabstractBACKGROUND: A visualization referred to as rainfall plot has recently gained popularity in genome data analysis. The plot is mostly used for illustrating the distribution of somatic cancer mutations along a reference genome, typically aiming to identify mutation hotspots. In general terms, the rainfall plot can be seen as a scatter plot showing the location of events on the x-axis versus the distance between consecutive events on the y-axis. Despite its frequent use, the motivation for applying this particular visualization and the appropriateness of its usage have never been critically addressed in detail. RESULTS: We show that the rainfall plot allows visual detection even for events occurring at high frequency over very short distances. In addition, event clustering at multiple scales may be detected as distinct horizontal bands in rainfall plots. At the same time, due to the limited size of standard figures, rainfall plots might suffer from inability to distinguish overlapping events, especially when multiple datasets are plotted in the same figure. We demonstrate the consequences of plot congestion, which results in obscured visual data interpretations. CONCLUSIONS: This work provides the first comprehensive survey of the characteristics and proper usage of rainfall plots. We find that the rainfall plot is able to convey a large amount of information without any need for parameterization or tuning. However, we also demonstrate how plot congestion and the use of a logarithmic y-axis may result in obscured visual data interpretations. To aid the productive utilization of rainfall plots, we demonstrate their characteristics and potential pitfalls using both simulated and real data, and provide a set of practical guidelines for their proper interpretation and usage. Diana Domanska, Daniel Vodák, Christin Lund-Andersen, Stefania Salvatore, Eivind Hovig, Geir Kjetil Sandve |
BMC Bioinform. | 6 |
| 2017 | NucDiff: in-depth characterization and annotation of differences between two sets of DNA sequencesabstractBACKGROUND: Comparing sets of sequences is a situation frequently encountered in bioinformatics, examples being comparing an assembly to a reference genome, or two genomes to each other. The purpose of the comparison is usually to find where the two sets differ, e.g. to find where a subsequence is repeated or deleted, or where insertions have been introduced. Such comparisons can be done using whole-genome alignments. Several tools for making such alignments exist, but none of them 1) provides detailed information about the types and locations of all differences between the two sets of sequences, 2) enables visualisation of alignment results at different levels of detail, and 3) carefully takes genomic repeats into consideration. RESULTS: We here present NucDiff, a tool aimed at locating and categorizing differences between two sets of closely related DNA sequences. NucDiff is able to deal with very fragmented genomes, repeated sequences, and various local differences and structural rearrangements. NucDiff determines differences by a rigorous analysis of alignment results obtained by the NUCmer, delta-filter and show-snps programs in the MUMmer sequence alignment package. All differences found are categorized according to a carefully defined classification scheme covering all possible differences between two sequences. Information about the differences is made available as GFF3 files, thus enabling visualisation using genome browsers as well as usage of the results as a component in an analysis pipeline. NucDiff was tested with varying parameters for the alignment step and compared with existing alternatives, called QUAST and dnadiff. CONCLUSIONS: We have developed a whole genome alignment difference classification scheme together with the program NucDiff for finding such differences. The proposed classification scheme is comprehensive and can be used by other tools. NucDiff performs comparably to QUAST and dnadiff but gives much more detailed results that can easily be visualized. NucDiff is freely available on https://github.com/uio-cels/NucDiff under the MPL license. Ksenia Khelik, Karin Lagesen, Geir Kjetil Sandve, Torbjørn Rognes, Alexander Johan Nederbragt |
BMC Bioinform. | 3 |
| 2017 | Coordinates and intervals in graph-based reference genomesabstractBACKGROUND: It has been proposed that future reference genomes should be graph structures in order to better represent the sequence diversity present in a species. However, there is currently no standard method to represent genomic intervals, such as the positions of genes or transcription factor binding sites, on graph-based reference genomes. RESULTS: We formalize offset-based coordinate systems on graph-based reference genomes and introduce methods for representing intervals on these reference structures. We show the advantage of our methods by representing genes on a graph-based representation of the newest assembly of the human genome (GRCh38) and its alternative loci for regions that are highly variable. CONCLUSION: More complex reference genomes, containing alternative loci, require methods to represent genomic data on these structures. Our proposed notation for genomic intervals makes it possible to fully utilize the alternative loci of the GRCh38 assembly and potential future graph-based reference genomes. We have made a Python package for representing such intervals on offset-based coordinate systems, available at https://github.com/uio-cels/offsetbasedgraph . An interactive web-tool using this Python package to visualize genes on a graph created from GRCh38 is available at https://github.com/uio-cels/genomicgraphcoords . Knut D. Rand, Ivar Grytten, Alexander Johan Nederbragt, Geir Storvik, Ingrid Kristine Glad, Geir Kjetil Sandve |
BMC Bioinform. | 6 |
| 2016 | In the loop: promoter-enhancer interactions and bioinformaticsabstractEnhancer-promoter regulation is a fundamental mechanism underlying differential transcriptional regulation. Spatial chromatin organization brings remote enhancers in contact with target promoters in cis to regulate gene expression. There is considerable evidence for promoter-enhancer interactions (PEIs). In the recent years, genome-wide analyses have identified signatures and mapped novel enhancers; however, being able to precisely identify their target gene(s) requires massive biological and bioinformatics efforts. In this review, we give a short overview of the chromatin landscape and transcriptional regulation. We discuss some key concepts and problems related to chromatin interaction detection technologies, and emerging knowledge from genome-wide chromatin interaction data sets. Then, we critically review different types of bioinformatics analysis methods and tools related to representation and visualization of PEI data, raw data processing and PEI prediction. Lastly, we provide specific examples of how PEIs have been used to elucidate a functional role of non-coding single-nucleotide polymorphisms. The topic is at the forefront of epigenetic research, and by highlighting some future bioinformatics challenges in the field, this review provides a comprehensive background for future PEI studies. Antonio Mora, Geir Kjetil Sandve, Odd Stokke Gabrielsen, Ragnhild Eskeland |
Briefings Bioinform. | 2 |
| 2016 | Galaxy Portal: interacting with the galaxy platform through mobile devicesabstractUNLABELLED: : We present Galaxy Portal app, an open source interface to the Galaxy system through smart phones and tablets. The Galaxy Portal provides convenient and efficient monitoring of job completion, as well as opportunities for inspection of results and execution history. In addition to being useful to the Galaxy community, we believe that the app also exemplifies a useful way of exploiting mobile interfaces for research/high-performance computing resources in general. AVAILABILITY AND IMPLEMENTATION: The source is freely available under a GPL license on GitHub, along with user documentation and pre-compiled binaries and instructions for several platforms: https://github.com/Tarostar/QMLGalaxyPortal It is available for iOS version 7 (and newer) through the Apple App Store, and for Android through Google Play for version 4.1 (API 16) or newer. CONTACT: [email protected]. Claus Børnich, Ivar Grytten, Eivind Hovig, Jonas Paulsen, Martin Cech, Geir Kjetil Sandve |
Bioinform. | 6 |
| 2014 | HiBrowse: multi-purpose statistical analysis of genome-wide chromatin 3D organizationabstractUNLABELLED: Recently developed methods that couple next-generation sequencing with chromosome conformation capture-based techniques, such as Hi-C and ChIA-PET, allow for characterization of genome-wide chromatin 3D structure. Understanding the organization of chromatin in three dimensions is a crucial next step in the unraveling of global gene regulation, and methods for analyzing such data are needed. We have developed HiBrowse, a user-friendly web-tool consisting of a range of hypothesis-based and descriptive statistics, using realistic assumptions in null-models. AVAILABILITY AND IMPLEMENTATION: HiBrowse is supported by all major browsers, and is freely available at http://hyperbrowser.uio.no/3d. Software is implemented in Python, and source code is available for download by following instructions on the main site. Jonas Paulsen, Geir Kjetil Sandve, Sveinung Gundersen, Tonje Lien, Kai Trengereid, Eivind Hovig |
Bioinform. | 2 |
| 2013 | Ten Simple Rules for Reproducible Computational ResearchabstractReplication is the cornerstone of a cumulative science [1]. However, new tools and technologies, massive amounts of data, interdisciplinary approaches, and the complexity of the questions being asked are complicating replication efforts, as are increased pressures on scientists to advance their research [2]. As full replication of studies on independently collected data is often not feasible, there has recently been a call for reproducible research as an attainable minimum standard for assessing the value of scientific claims [3]. This requires that papers in experimental science describe the results and provide a sufficiently clear protocol to allow successful repetition and extension of analyses based on original data [4].
The importance of replication and reproducibility has recently been exemplified through studies showing that scientific papers commonly leave out experimental details essential for reproduction [5], studies showing difficulties with replicating published experimental results [6], an increase in retracted papers [7], and through a high number of failing clinical trials [8], [9]. This has led to discussions on how individual researchers, institutions, funding bodies, and journals can establish routines that increase transparency and reproducibility. In order to foster such aspects, it has been suggested that the scientific community needs to develop a “culture of reproducibility” for computational science, and to require it for published claims [3].
We want to emphasize that reproducibility is not only a moral responsibility with respect to the scientific field, but that a lack of reproducibility can also be a burden for you as an individual researcher. As an example, a good practice of reproducibility is necessary in order to allow previously developed methodology to be effectively applied on new data, or to allow reuse of code and results for new projects. In other words, good habits of reproducibility may actually turn out to be a time-saver in the longer run.
We further note that reproducibility is just as much about the habits that ensure reproducible research as the technologies that can make these processes efficient and realistic. Each of the following ten rules captures a specific aspect of reproducibility, and discusses what is needed in terms of information handling and tracking of procedures. If you are taking a bare-bones approach to bioinformatics analysis, i.e., running various custom scripts from the command line, you will probably need to handle each rule explicitly. If you are instead performing your analyses through an integrated framework (such as GenePattern [10], Galaxy [11], LONI pipeline [12], or Taverna [13]), the system may already provide full or partial support for most of the rules. What is needed on your part is then merely the knowledge of how to exploit these existing possibilities.
In a pragmatic setting, with publication pressure and deadlines, one may face the need to make a trade-off between the ideals of reproducibility and the need to get the research out while it is still relevant. This trade-off becomes more important when considering that a large part of the analyses being tried out never end up yielding any results. However, frequently one will, with the wisdom of hindsight, contemplate the missed opportunity to ensure reproducibility, as it may already be too late to take the necessary notes from memory (or at least much more difficult than to do it while underway). We believe that the rewards of reproducibility will compensate for the risk of having spent valuable time developing an annotated catalog of analyses that turned out as blind alleys.
As a minimal requirement, you should at least be able to reproduce the results yourself. This would satisfy the most basic requirements of sound research, allowing any substantial future questioning of the research to be met with a precise explanation. Although it may sound like a very weak requirement, even this level of reproducibility will often require a certain level of care in order to be met. There will for a given analysis be an exponential number of possible combinations of software versions, parameter values, pre-processing steps, and so on, meaning that a failure to take notes may make exact reproduction essentially impossible.
With this basic level of reproducibility in place, there is much more that can be wished for. An obvious extension is to go from a level where you can reproduce results in case of a critical situation to a level where you can practically and routinely reuse your previous work and increase your productivity. A second extension is to ensure that peers have a practical possibility of reproducing your results, which can lead to increased trust in, interest for, and citations of your work [6], [14].
We here present ten simple rules for reproducibility of computational research. These rules can be at your disposal for whenever you want to make your research more accessible—be it for peers or for your future self. Geir Kjetil Sandve, Anton Nekrutenko, James Taylor 0001, Eivind Hovig |
PLoS Comput. Biol. | 1 |
| 2011 | Sequential Monte Carlo multiple testingabstractMOTIVATION: In molecular biology, as in many other scientific fields, the scale of analyses is ever increasing. Often, complex Monte Carlo simulation is required, sometimes within a large-scale multiple testing setting. The resulting computational costs may be prohibitively high. RESULTS: We here present MCFDR, a simple, novel algorithm for false discovery rate (FDR) modulated sequential Monte Carlo (MC) multiple hypothesis testing. The algorithm iterates between adding MC samples across tests and calculating intermediate FDR values for the collection of tests. MC sampling is stopped either by sequential MC or based on a threshold on FDR. An essential property of the algorithm is that it limits the total number of MC samples whatever the number of true null hypotheses. We show on both real and simulated data that the proposed algorithm provides large gains in computational efficiency. AVAILABILITY: MCFDR is implemented in the Genomic HyperBrowser (http://hyperbrowser.uio.no/mcfdr), a web-based system for genome analysis. All input data and results are available and can be reproduced through a Galaxy Pages document at: http://hyperbrowser.uio.no/mcfdr/u/sandve/p/mcfdr. CONTACT: [email protected]. Geir Kjetil Sandve, Egil Ferkingstad, Ståle Nygård |
Bioinform. | 1 |
| 2011 | Identifying elemental genomic track types and representing them uniformlyabstractBACKGROUND: With the recent advances and availability of various high-throughput sequencing technologies, data on many molecular aspects, such as gene regulation, chromatin dynamics, and the three-dimensional organization of DNA, are rapidly being generated in an increasing number of laboratories. The variation in biological context, and the increasingly dispersed mode of data generation, imply a need for precise, interoperable and flexible representations of genomic features through formats that are easy to parse. A host of alternative formats are currently available and in use, complicating analysis and tool development. The issue of whether and how the multitude of formats reflects varying underlying characteristics of data has to our knowledge not previously been systematically treated. RESULTS: We here identify intrinsic distinctions between genomic features, and argue that the distinctions imply that a certain variation in the representation of features as genomic tracks is warranted. Four core informational properties of tracks are discussed: gaps, lengths, values and interconnections. From this we delineate fifteen generic track types. Based on the track type distinctions, we characterize major existing representational formats and find that the track types are not adequately supported by any single format. We also find, in contrast to the XML formats, that none of the existing tabular formats are conveniently extendable to support all track types. We thus propose two unified formats for track data, an improved XML format, BioXSD 1.1, and a new tabular format, GTrack 1.0. CONCLUSIONS: The defined track types are shown to capture relevant distinctions between genomic annotation tracks, resulting in varying representational needs and analysis possibilities. The proposed formats, GTrack 1.0 and BioXSD 1.1, cater to the identified track distinctions and emphasize preciseness, flexibility and parsing convenience. Sveinung Gundersen, Matús Kalas, Osman Abul, Arnoldo Frigessi, Eivind Hovig, Geir Kjetil Sandve |
BMC Bioinform. | 6 |
| 2008 | BayCis: A Bayesian Hierarchical HMM for Cis-Regulatory Module Decoding in Metazoan Genomes
Tien-ho Lin, Pradipta Ray, Geir Kjetil Sandve, Selen Uguroglu, Eric P. Xing |
RECOMB | 3 |
| 2008 | Assessment of composite motif discovery methodsabstractBACKGROUND: Computational discovery of regulatory elements is an important area of bioinformatics research and more than a hundred motif discovery methods have been published. Traditionally, most of these methods have addressed the problem of single motif discovery - discovering binding motifs for individual transcription factors. In higher organisms, however, transcription factors usually act in combination with nearby bound factors to induce specific regulatory behaviours. Hence, recent focus has shifted from single motifs to the discovery of sets of motifs bound by multiple cooperating transcription factors, so called composite motifs or cis-regulatory modules. Given the large number and diversity of methods available, independent assessment of methods becomes important. Although there have been several benchmark studies of single motif discovery, no similar studies have previously been conducted concerning composite motif discovery. RESULTS: We have developed a benchmarking framework for composite motif discovery and used it to evaluate the performance of eight published module discovery tools. Benchmark datasets were constructed based on real genomic sequences containing experimentally verified regulatory modules, and the module discovery programs were asked to predict both the locations of these modules and to specify the single motifs involved. To aid the programs in their search, we provided position weight matrices corresponding to the binding motifs of the transcription factors involved. In addition, selections of decoy matrices were mixed with the genuine matrices on one dataset to test the response of programs to varying levels of noise. CONCLUSION: Although some of the methods tested tended to score somewhat better than others overall, there were still large variations between individual datasets and no single method performed consistently better than the rest in all situations. The variation in performance on individual datasets also shows that the new benchmark datasets represents a suitable variety of challenges to most methods for module discovery. Kjetil Klepper, Geir Kjetil Sandve, Osman Abul, Jostein Johansen, Finn Drabløs |
BMC Bioinform. | 2 |
| 2008 | Compo: composite motif discovery using discrete modelsabstractBACKGROUND: Computational discovery of motifs in biomolecular sequences is an established field, with applications both in the discovery of functional sites in proteins and regulatory sites in DNA. In recent years there has been increased attention towards the discovery of composite motifs, typically occurring in cis-regulatory regions of genes. RESULTS: This paper describes Compo: a discrete approach to composite motif discovery that supports richer modeling of composite motifs and a more realistic background model compared to previous methods. Furthermore, multiple parameter and threshold settings are tested automatically, and the most interesting motifs across settings are selected. This avoids reliance on single hard thresholds, which has been a weakness of previous discrete methods. Comparison of motifs across parameter settings is made possible by the use of p-values as a general significance measure. Compo can either return an ordered list of motifs, ranked according to the general significance measure, or a Pareto front corresponding to a multi-objective evaluation on sensitivity, specificity and spatial clustering. CONCLUSION: Compo performs very competitively compared to several existing methods on a collection of benchmark data sets. These benchmarks include a recently published, large benchmark suite where the use of support across sequences allows Compo to correctly identify binding sites even when the relevant PWMs are mixed with a large number of noise PWMs. Furthermore, the possibility of parameter-free running offers high usability, the support for multi-objective evaluation allows a rich view of potential regulators, and the discrete model allows flexibility in modeling and interpretation of motifs. Geir Kjetil Sandve, Osman Abul, Finn Drabløs |
BMC Bioinform. | 1 |
| 2007 | False Discovery Rates in Identifying Functional DNA MotifsabstractThere are several methods for scoring a set of upstream DNA sequences against a given motif. Typically, significance of raw scores are based on p-values, measured by statistical hypothesis testing. As an extension, multiple hypothesis testing is adopted in cases where there are multiple motifs to be evaluated in parallel. In this way significant motifs are identified for a given significance level. However, a set of significantly identified motifs can contain false positives. In this work, we introduce a false discovery rate estimation problem for significantly predicted motifs. An explorative method for this problem is presented. We test the method using TRANSFAC and JASPAR motif libraries on several upstream DNA subsets of S.cerevisiae. The results show the effectiveness of the method. Osman Abul, Geir Kjetil Sandve, Finn Drabløs |
BIBE | 2 |
| 2007 | Improved benchmarks for computational motif discoveryabstractBACKGROUND: An important step in annotation of sequenced genomes is the identification of transcription factor binding sites. More than a hundred different computational methods have been proposed, and it is difficult to make an informed choice. Therefore, robust assessment of motif discovery methods becomes important, both for validation of existing tools and for identification of promising directions for future research. RESULTS: We use a machine learning perspective to analyze collections of transcription factors with known binding sites. Algorithms are presented for finding position weight matrices (PWMs), IUPAC-type motifs and mismatch motifs with optimal discrimination of binding sites from remaining sequence. We show that for many data sets in a recently proposed benchmark suite for motif discovery, none of the common motif models can accurately discriminate the binding sites from remaining sequence. This may obscure the distinction between the potential performance of the motif discovery tool itself versus the intrinsic complexity of the problem we are trying to solve. Synthetic data sets may avoid this problem, but we show on some previously proposed benchmarks that there may be a strong bias towards a presupposed motif model. We also propose a new approach to benchmark data set construction. This approach is based on collections of binding site fragments that are ranked according to the optimal level of discrimination achieved with our algorithms. This allows us to select subsets with specific properties. We present one benchmark suite with data sets that allow good discrimination between positive and negative instances with the common motif models. These data sets are suitable for evaluating algorithms for motif discovery that rely on these models. We present another benchmark suite where PWM, IUPAC and mismatch motif models are not able to discriminate reliably between positive and negative instances. This suite could be used for evaluating more powerful motif models. CONCLUSION: Our improved benchmark suites have been designed to differentiate between the performance of motif discovery algorithms and the power of motif models. We provide a web server where users can download our benchmark suites, submit predictions and visualize scores on the benchmarks. Geir Kjetil Sandve, Osman Abul, Vegard Walseng, Finn Drabløs |
BMC Bioinform. | 1 |
| 2006 | Accelerating Motif Discovery: Motif Matching on Parallel Hardware
Geir Kjetil Sandve, Magnar Nedland, Øyvind Bø Syrstad, Lars Andreas Eidsheim, Osman Abul, Finn Drabløs |
WABI | 1 |
| 2005 | Generalized Composite Motif Discovery
Geir Kjetil Sandve, Finn Drabløs |
KES (3) | 1 |