EDBT 2026 Demo / reviewers in the wild / expert
Thorfinn Korneliussen
dblp:29/9718 · also Thorfinn Sand Korneliussen
· DBLP profile ↗
9ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0001-7576-5380ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
6 papers |
Bioinformatics and computational biology · 100% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › molecular evolution › evolutionary bioinformatics
ancient DNA analysis |
1.3 | 2 | 2025 | <tt>AdDeam</tt> : a fast and scalable tool for estimating and clustering reference-level damage profiles · Bioinform. 2025 A likelihood method for estimating present-day human contamination in ancient male samples using low-depth X-chromosome data · Bioinform. 2020 |
Bioinformatics and computational biology › statistical genetics › genotype calling
genotype likelihood estimation |
0.9 | 1 | 2025 | vcfgl: a flexible genotype likelihood simulator for VCF/BCF files · Bioinform. 2025 |
Bioinformatics and computational biology
population genetics |
0.8 | 2 | 2022 | LocalNgsRelate: a software tool for inferring IBD sharing along the genome between pairs of individuals from low-depth NGS data · Bioinform. 2022 NgsRelate: a software tool for estimating pairwise relatedness from next-generation sequencing data · Bioinform. 2015 |
Bioinformatics and computational biology › genomics
sequencing |
0.7 | 1 | 2023 | NGSNGS: next-generation simulator for next-generation sequencing data · Bioinform. 2023 |
Bioinformatics and computational biology
sequencing simulation |
0.7 | 1 | 2023 | NGSNGS: next-generation simulator for next-generation sequencing data · Bioinform. 2023 |
Bioinformatics and computational biology › population genetics › identity by descent
identity-by-descent detection |
0.6 | 1 | 2022 | LocalNgsRelate: a software tool for inferring IBD sharing along the genome between pairs of individuals from low-depth NGS data · Bioinform. 2022 |
Bioinformatics and computational biology
metagenomics |
0.3 | 1 | 2025 | <tt>AdDeam</tt> : a fast and scalable tool for estimating and clustering reference-level damage profiles · Bioinform. 2025 |
Bioinformatics and computational biology › genomics
next-generation sequencing data analysis |
0.1 | 1 | 2015 | NgsRelate: a software tool for estimating pairwise relatedness from next-generation sequencing data · Bioinform. 2015 |
Methods — techniques the papers use, named apart from their topics
principal component analysis · 0.9genotype likelihood models · 0.9clustering · 0.9beta distribution simulation · 0.9genotype likelihood · 0.8multithreading · 0.7probabilistic inference · 0.6x-chromosome analysis · 0.4maximum likelihood · 0.4maximum likelihood estimation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | vcfgl: a flexible genotype likelihood simulator for VCF/BCF filesabstractMOTIVATION: Accurate quantification of genotype uncertainty is pivotal in ensuring the reliability of genetic inferences drawn from NGS data. Genotype uncertainty is typically modeled using Genotype Likelihoods (GLs), which can help propagate measures of statistical uncertainty in base calls to downstream analyses. However, the effects of errors and biases in the estimation of GLs, introduced by biases in the original base call quality scores or the discretization of quality scores, as well as the choice of the GL model, remain under-explored. RESULTS: We present vcfgl, a versatile tool for simulating genotype likelihoods associated with simulated read data. It offers a framework for researchers to simulate and investigate the uncertainties and biases associated with the quantification of uncertainty, thereby facilitating a deeper understanding of their impacts on downstream analytical methods. Through simulations, we demonstrate the utility of vcfgl in benchmarking GL-based methods. The program can calculate GLs using various widely used genotype likelihood models and can simulate the errors in quality scores using a Beta distribution. It is compatible with modern simulators such as msprime and SLiM, and can output data in pileup, Variant Call Format (VCF)/BCF, and genomic VCF file formats, supporting a wide range of applications. The vcfgl program is freely available as an efficient and user-friendly software written in C/C++. AVAILABILITY AND IMPLEMENTATION: vcfgl is freely available at https://github.com/isinaltinkaya/vcfgl. Isin Altinkaya, Rasmus Nielsen, Thorfinn Korneliussen |
Bioinform. | 3 |
| 2025 | <tt>AdDeam</tt> : a fast and scalable tool for estimating and clustering reference-level damage profilesabstractMOTIVATION: DNA damage patterns, such as increased frequencies of C→T and G→A substitutions at fragment ends, are widely used in ancient DNA studies to assess authenticity and detect contamination. In metagenomic studies, fragments can be mapped against multiple references or de novo assembled contigs to identify those likely to be ancient. Generating and comparing damage profiles, however, can be both tedious and time-consuming. Although tools exist for estimating damage in single reference genomes and metagenomic datasets, none efficiently cluster damage patterns. RESULTS: To address this methodological gap, we developed AdDeam, a tool that combines rapid damage estimation with clustering for streamlined analyses and easy identification of potential contaminants or outliers. Our tool takes aligned ancient DNA (aDNA) fragments from various samples or contigs as input, computes damage patterns, clusters them, and outputs representative damage profiles per cluster, a probability of each sample pertaining to a cluster, as well as a Principal Component Analysis of the damage patterns for each sample for fast visualisation. We evaluated AdDeam on both simulated and empirical datasets. AdDeam effectively distinguishes different damage levels, such as uracil-DNA glycosylase-treated samples, sample-specific damages from specimens of different time periods, and can also distinguish between contigs containing modern or ancient fragments, providing a clear framework for aDNA authentication and facilitating large-scale analyses. AVAILABILITY AND IMPLEMENTATION: AdDeam is publicly available at https://github.com/LouisPwr/AdDeam and can also be installed via Bioconda. It is implemented in Python and C++. All analysis scripts and datasets are available at https://github.com/LouisPwr/AdDeamAnalysis and on Zenodo under: 10.5281/zenodo.15052427. Louis Kraft, Thorfinn Korneliussen, Peter Wad Sackett, Gabriel Renaud |
Bioinform. | 2 |
| 2023 | NGSNGS: next-generation simulator for next-generation sequencing dataabstractSUMMARY: With the rapid expansion of the capabilities of the DNA sequencers throughout the different sequencing generations, the quantity of generated data has likewise increased. This evolution has also led to new bioinformatical methods, for which in silico data have become crucial when verifying the accuracy of a model or the robustness of a genomic analysis pipeline. Here, we present a multithreaded next-generation simulator for next-generation sequencing data (NGSNGS), which simulates reads faster than currently available methods and programs. NGSNGS can simulate reads with platform-specific characteristics based on nucleotide quality score profiles as well as including a post-mortem damage model which is relevant for simulating ancient DNA. The simulated sequences are sampled (with replacement) from a reference DNA genome, which can represent a haploid genome, polyploid assemblies or even population haplotypes and allows the user to simulate known variable sites directly. The program is implemented in a multithreading framework and is factors faster than currently available tools while extending their feature set and possible output formats. AVAILABILITY AND IMPLEMENTATION: The method and associated programs are released as open-source software, code and user manual are available at https://github.com/RAHenriksen/NGSNGS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rasmus Amund Henriksen, Thorfinn Korneliussen |
Bioinform. | 3 |
| 2022 | LocalNgsRelate: a software tool for inferring IBD sharing along the genome between pairs of individuals from low-depth NGS dataabstractMOTIVATION: Inference of identity-by-descent (IBD) sharing along the genome between pairs of individuals has important uses. But all existing inference methods are based on genotypes, which is not ideal for low-depth Next Generation Sequencing (NGS) data from which genotypes can only be called with high uncertainty. RESULTS: We present a new probabilistic software tool, LocalNgsRelate, for inferring IBD sharing along the genome between pairs of individuals from low-depth NGS data. Its inference is based on genotype likelihoods instead of genotypes, and thereby it takes the uncertainty of the genotype calling into account. Using real data from the 1000 Genomes project, we show that LocalNgsRelate provides more accurate IBD inference for low-depth NGS data than two state-of-the-art genotype-based methods, Albrechtsen et al. (2009) and hap-IBD. We also show that the method works well for NGS data down to a depth of 2×. AVAILABILITY AND IMPLEMENTATION: LocalNgsRelate is freely available at https://github.com/idamoltke/LocalNgsRelate. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alissa L. Severson, Thorfinn Korneliussen, Ida Moltke |
Bioinform. | 2 |
| 2020 | A likelihood method for estimating present-day human contamination in ancient male samples using low-depth X-chromosome dataabstractMOTIVATION: The presence of present-day human contaminating DNA fragments is one of the challenges defining ancient DNA (aDNA) research. This is especially relevant to the ancient human DNA field where it is difficult to distinguish endogenous molecules from human contaminants due to their genetic similarity. Recently, with the advent of high-throughput sequencing and new aDNA protocols, hundreds of ancient human genomes have become available. Contamination in those genomes has been measured with computational methods often developed specifically for these empirical studies. Consequently, some of these methods have not been implemented and tested for general use while few are aimed at low-depth nuclear data, a common feature in aDNA datasets. RESULTS: We develop a new X-chromosome-based maximum likelihood method for estimating present-day human contamination in low-depth sequencing data from male individuals. We implement our method for general use, assess its performance under conditions typical of ancient human DNA research, and compare it to previous nuclear data-based methods through extensive simulations. For low-depth data, we show that existing methods can produce unusable estimates or substantially underestimate contamination. In contrast, our method provides accurate estimates for a depth of coverage as low as 0.5× on the X-chromosome when contamination is below 25%. Moreover, our method still yields meaningful estimates in very challenging situations, i.e. when the contaminant and the target come from closely related populations or with increased error rates. With a running time below 5 min, our method is applicable to large scale aDNA genomic studies. AVAILABILITY AND IMPLEMENTATION: The method is implemented in C++ and R and is available in github.com/sapfo/contaminationX and popgen.dk/angsd. José Víctor Moreno-Mayar, Thorfinn Korneliussen, Jyoti Dalal, Gabriel Renaud, Anders Albrechtsen, Rasmus Nielsen, Anna-Sapfo Malaspinas |
Bioinform. | 2 |
| 2015 | NgsRelate: a software tool for estimating pairwise relatedness from next-generation sequencing dataabstractMOTIVATION: Pairwise relatedness estimation is important in many contexts such as disease mapping and population genetics. However, all existing estimation methods are based on called genotypes, which is not ideal for next-generation sequencing (NGS) data of low depth from which genotypes cannot be called with high certainty. RESULTS: We present a software tool, NgsRelate, for estimating pairwise relatedness from NGS data. It provides maximum likelihood estimates that are based on genotype likelihoods instead of genotypes and thereby takes the inherent uncertainty of the genotypes into account. Using both simulated and real data, we show that NgsRelate provides markedly better estimates for low-depth NGS data than two state-of-the-art genotype-based methods. AVAILABILITY: NgsRelate is implemented in C++ and is available under the GNU license at www.popgen.dk/software. Thorfinn Korneliussen, Ida Moltke |
Bioinform. | 1 |
| 2014 | ANGSD: Analysis of Next Generation Sequencing DataabstractBACKGROUND: High-throughput DNA sequencing technologies are generating vast amounts of data. Fast, flexible and memory efficient implementations are needed in order to facilitate analyses of thousands of samples simultaneously. RESULTS: We present a multithreaded program suite called ANGSD. This program can calculate various summary statistics, and perform association mapping and population genetic analyses utilizing the full information in next generation sequencing data by working directly on the raw sequencing data or by using genotype likelihoods. CONCLUSIONS: The open source c/c++ program ANGSD is available at http://www.popgen.dk/angsd . The program is tested and validated on GNU/Linux systems. The program facilitates multiple input formats including BAM and imputed beagle genotype probability files. The program allow the user to choose between combinations of existing methods and can perform analysis that is not implemented elsewhere. Thorfinn Korneliussen, Anders Albrechtsen, Rasmus Nielsen |
BMC Bioinform. | 1 |
| 2013 | Calculation of Tajima's D and other neutrality test statistics from low depth next-generation sequencing dataabstractBACKGROUND: A number of different statistics are used for detecting natural selection using DNA sequencing data, including statistics that are summaries of the frequency spectrum, such as Tajima's D. These statistics are now often being applied in the analysis of Next Generation Sequencing (NGS) data. However, estimates of frequency spectra from NGS data are strongly affected by low sequencing coverage; the inherent technology dependent variation in sequencing depth causes systematic differences in the value of the statistic among genomic regions. RESULTS: We have developed an approach that accommodates the uncertainty of the data when calculating site frequency based neutrality test statistics. A salient feature of this approach is that it implicitly solves the problems of varying sequencing depth, missing data and avoids the need to infer variable sites for the analysis and thereby avoids ascertainment problems introduced by a SNP discovery process. CONCLUSION: Using an empirical Bayes approach for fast computations, we show that this method produces results for low-coverage NGS data comparable to those achieved when the genotypes are known without uncertainty. We also validate the method in an analysis of data from the 1000 genomes project. The method is implemented in a fast framework which enables researchers to perform these neutrality tests on a genome-wide scale. Thorfinn Korneliussen, Ida Moltke, Anders Albrechtsen, Rasmus Nielsen |
BMC Bioinform. | 1 |
| 2011 | Estimation of allele frequency and association mapping using next-generation sequencing dataabstractBACKGROUND: Estimation of allele frequency is of fundamental importance in population genetic analyses and in association mapping. In most studies using next-generation sequencing, a cost effective approach is to use medium or low-coverage data (e.g., < 15X). However, SNP calling and allele frequency estimation in such studies is associated with substantial statistical uncertainty because of varying coverage and high error rates. RESULTS: We evaluate a new maximum likelihood method for estimating allele frequencies in low and medium coverage next-generation sequencing data. The method is based on integrating over uncertainty in the data for each individual rather than first calling genotypes. This method can be applied to directly test for associations in case/control studies. We use simulations to compare the likelihood method to methods based on genotype calling, and show that the likelihood method outperforms the genotype calling methods in terms of: (1) accuracy of allele frequency estimation, (2) accuracy of the estimation of the distribution of allele frequencies across neutrally evolving sites, and (3) statistical power in association mapping studies. Using real re-sequencing data from 200 individuals obtained from an exon-capture experiment, we show that the patterns observed in the simulations are also found in real data. CONCLUSIONS: Overall, our results suggest that association mapping and estimation of allele frequencies should not be based on genotype calling in low to medium coverage data. Furthermore, if genotype calling methods are used, it is usually better not to filter genotypes based on the call confidence score. Su Yeon Kim, Kirk E. Lohmueller, Anders Albrechtsen, Yingrui Li, Thorfinn Korneliussen, Geng Tian, Niels Grarup, Gitte Andersen, Daniel Witte, Torben Jørgensen, Torben Hansen, Oluf Pedersen, Jun Wang 0004, Rasmus Nielsen |
BMC Bioinform. | 5 |