VLDB 2026 Research / reviewers in the wild / expert
Rasmus Nielsen
dblp:51/725
· DBLP profile ↗
20ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0003-0513-6591ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 4 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | vcfgl: a flexible genotype likelihood simulator for VCF/BCF filesabstractMOTIVATION: Accurate quantification of genotype uncertainty is pivotal in ensuring the reliability of genetic inferences drawn from NGS data. Genotype uncertainty is typically modeled using Genotype Likelihoods (GLs), which can help propagate measures of statistical uncertainty in base calls to downstream analyses. However, the effects of errors and biases in the estimation of GLs, introduced by biases in the original base call quality scores or the discretization of quality scores, as well as the choice of the GL model, remain under-explored. RESULTS: We present vcfgl, a versatile tool for simulating genotype likelihoods associated with simulated read data. It offers a framework for researchers to simulate and investigate the uncertainties and biases associated with the quantification of uncertainty, thereby facilitating a deeper understanding of their impacts on downstream analytical methods. Through simulations, we demonstrate the utility of vcfgl in benchmarking GL-based methods. The program can calculate GLs using various widely used genotype likelihood models and can simulate the errors in quality scores using a Beta distribution. It is compatible with modern simulators such as msprime and SLiM, and can output data in pileup, Variant Call Format (VCF)/BCF, and genomic VCF file formats, supporting a wide range of applications. The vcfgl program is freely available as an efficient and user-friendly software written in C/C++. AVAILABILITY AND IMPLEMENTATION: vcfgl is freely available at https://github.com/isinaltinkaya/vcfgl. Isin Altinkaya, Rasmus Nielsen, Thorfinn Korneliussen |
Bioinform. | 2 |
| 2022 | SCONCE: a method for profiling copy number alterations in cancer evolution using single-cell whole genome sequencingabstractMOTIVATION: Copy number alterations (CNAs) are a significant driver in cancer growth and development, but remain poorly characterized on the single cell level. Although genome evolution in cancer cells is Markovian through evolutionary time, CNAs are not Markovian along the genome. However, existing methods call copy number profiles with Hidden Markov Models or change point detection algorithms based on changes in observed read depth, corrected by genome content and do not account for the stochastic evolutionary process. RESULTS: We present a theoretical framework to use tumor evolutionary history to accurately call CNAs in a principled manner. To model the tumor evolutionary process and account for technical noise from low coverage single-cell whole genome sequencing data, we developed SCONCE, a method based on a Hidden Markov Model to analyze read depth data from tumor cells using matched normal cells as negative controls. Using a combination of public data sets and simulations, we show SCONCE accurately decodes copy number profiles, and provides a useful tool for understanding tumor evolution. AVAILABILITYAND IMPLEMENTATION: SCONCE is implemented in C++11 and is freely available from https://github.com/NielsenBerkeleyLab/sconce. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sandra Hui, Rasmus Nielsen |
Bioinform. | 2 |
| 2022 | AncestralClust: clustering of divergent nucleotide sequences by ancestral sequence reconstruction using phylogenetic treesabstractMOTIVATION: Clustering is a fundamental task in the analysis of nucleotide sequences. Despite the exponential increase in the size of sequence databases of homologous genes, few methods exist to cluster divergent sequences. Traditional clustering methods have mostly focused on optimizing high speed clustering of highly similar sequences. We develop a phylogenetic clustering method which infers ancestral sequences for a set of initial clusters and then uses a greedy algorithm to cluster sequences. RESULTS: We describe a clustering program AncestralClust, which is developed for clustering divergent sequences. We compare this method with other state-of-the-art clustering methods using datasets of homologous sequences from different species. We show that, in divergent datasets, AncestralClust has higher accuracy and more even cluster sizes than current popular methods. AVAILABILITY AND IMPLEMENTATION: AncestralClust is an Open Source program available at https://github.com/lpipes/ancestralclust. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lenore Pipes, Rasmus Nielsen |
Bioinform. | 2 |
| 2022 | SCONCE2: jointly inferring single cell copy number profiles and tumor evolutionary distancesabstractBACKGROUND: Single cell whole genome tumor sequencing can yield novel insights into the evolutionary history of somatic copy number alterations. Existing single cell copy number calling methods do not explicitly model the shared evolutionary process of multiple cells, and generally analyze cells independently. Additionally, existing methods for estimating tumor cell phylogenies using copy number profiles are sensitive to profile estimation errors. RESULTS: We present SCONCE2, a method for jointly calling copy number alterations and estimating pairwise distances for single cell sequencing data. Using simulations, we show that SCONCE2 has higher accuracy in copy number calling and phylogeny estimation than competing methods. We apply SCONCE2 to previously published single cell sequencing data to illustrate the utility of the method. CONCLUSIONS: SCONCE2 jointly estimates copy number profiles and a distance metric for inferring tumor phylogenies in single cell whole genome tumor sequencing across multiple cells, enabling deeper understandings of tumor evolution. Sandra Hui, Rasmus Nielsen |
BMC Bioinform. | 2 |
| 2020 | A likelihood method for estimating present-day human contamination in ancient male samples using low-depth X-chromosome dataabstractMOTIVATION: The presence of present-day human contaminating DNA fragments is one of the challenges defining ancient DNA (aDNA) research. This is especially relevant to the ancient human DNA field where it is difficult to distinguish endogenous molecules from human contaminants due to their genetic similarity. Recently, with the advent of high-throughput sequencing and new aDNA protocols, hundreds of ancient human genomes have become available. Contamination in those genomes has been measured with computational methods often developed specifically for these empirical studies. Consequently, some of these methods have not been implemented and tested for general use while few are aimed at low-depth nuclear data, a common feature in aDNA datasets. RESULTS: We develop a new X-chromosome-based maximum likelihood method for estimating present-day human contamination in low-depth sequencing data from male individuals. We implement our method for general use, assess its performance under conditions typical of ancient human DNA research, and compare it to previous nuclear data-based methods through extensive simulations. For low-depth data, we show that existing methods can produce unusable estimates or substantially underestimate contamination. In contrast, our method provides accurate estimates for a depth of coverage as low as 0.5× on the X-chromosome when contamination is below 25%. Moreover, our method still yields meaningful estimates in very challenging situations, i.e. when the contaminant and the target come from closely related populations or with increased error rates. With a running time below 5 min, our method is applicable to large scale aDNA genomic studies. AVAILABILITY AND IMPLEMENTATION: The method is implemented in C++ and R and is available in github.com/sapfo/contaminationX and popgen.dk/angsd. José Víctor Moreno-Mayar, Thorfinn Korneliussen, Jyoti Dalal, Gabriel Renaud, Anders Albrechtsen, Rasmus Nielsen, Anna-Sapfo Malaspinas |
Bioinform. | 6 |
| 2020 | Inferring the ancestry of parents and grandparents from genetic dataabstractInference of admixture proportions is a classical statistical problem in population genetics. Standard methods implicitly assume that both parents of an individual have the same admixture fraction. However, this is rarely the case in real data. In this paper we show that the distribution of admixture tract lengths in a genome contains information about the admixture proportions of the ancestors of an individual. We develop a Hidden Markov Model (HMM) framework for estimating the admixture proportions of the immediate ancestors of an individual, i.e. a type of decomposition of an individual's admixture proportions into further subsets of ancestral proportions in the ancestors. Based on a genealogical model for admixture tracts, we develop an efficient algorithm for computing the sampling probability of the genome from a single individual, as a function of the admixture proportions of the ancestors of this individual. This allows us to perform probabilistic inference of admixture proportions of ancestors only using the genome of an extant individual. We perform extensive simulations to quantify the error in the estimation of ancestral admixture proportions under various conditions. To illustrate the utility of the method, we apply it to real genetic data. Jingwen Pei, Yiming Zhang 0007, Rasmus Nielsen, Yufeng Wu 0001 |
PLoS Comput. Biol. | 3 |
| 2017 | Fast admixture analysis and population tree estimation for SNP and NGS dataabstractMOTIVATION: Structure methods are highly used population genetic methods for classifying individuals in a sample fractionally into discrete ancestry components. CONTRIBUTION: We introduce a new optimization algorithm for the classical STRUCTURE model in a maximum likelihood framework. Using analyses of real data we show that the new method finds solutions with higher likelihoods than the state-of-the-art method in the same computational time. The optimization algorithm is also applicable to models based on genotype likelihoods, that can account for the uncertainty in genotype-calling associated with Next Generation Sequencing (NGS) data. We also present a new method for estimating population trees from ancestry components using a Gaussian approximation. Using coalescence simulations of diverging populations, we explore the adequacy of the STRUCTURE-style models and the Gaussian assumption for identifying ancestry components correctly and for inferring the correct tree. In most cases, ancestry components are inferred correctly, although sample sizes and times since admixture can influence the results. We show that the popular Gaussian approximation tends to perform poorly under extreme divergence scenarios e.g. with very long branch lengths, but the topologies of the population trees are accurately inferred in all scenarios explored. The new methods are implemented together with appropriate visualization tools in the software package Ohana. AVAILABILITY AND IMPLEMENTATION: Ohana is publicly available at https://github.com/jade-cheng/ohana . In addition to source code and installation instructions, we also provide example work-flows in the project wiki site. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jade Yu Cheng, Thomas Mailund, Rasmus Nielsen |
Bioinform. | 3 |
| 2016 | Workshop message: IoT-SoS 2016abstractIt is our great pleasure to welcome you to the 5th edition of the successful workshop on the Internet of Things: Smart Objects and Services (IoT-SoS), which is organized this year in conjunction with WoWMoM 2016. The main focus of this wokshop is to bring together experts from the research community, the industry and standardisation bodies and discuss in the context of heterogeneous networking technologies for the Internet of Things (IoT), aiming to identify solutions for addressing the key challenges that IoT brings to the networking domain. Elias Z. Tragos, Rasmus Nielsen, Adam Kapovits, Claudio Cicconetti, Enzo Mingozzi, Jaudelice Cavalcante de Oliveira, Xiaohua Jia, Stefano Iellamo, Vangelis Angelakis |
WoWMoM | 2 |
| 2016 | SweepFinder2: increased sensitivity, robustness and flexibilityabstractUNLABELLED: SweepFinder is a widely used program that implements a powerful likelihood-based method for detecting recent positive selection, or selective sweeps. Here, we present SweepFinder2, an extension of SweepFinder with increased sensitivity and robustness to the confounding effects of mutation rate variation and background selection. Moreover, SweepFinder2 has increased flexibility that enables the user to specify test sites, set the distance between test sites and utilize a recombination map. AVAILABILITY AND IMPLEMENTATION: SweepFinder2 is a freely-available (www.personal.psu.edu/mxd60/sf2.html) software package that is written in C and can be run from a Unix command line. CONTACT: [email protected]. Michael Degiorgio, Christian D. Huber, Melissa J. Hubisz, Ines Hellmann, Rasmus Nielsen |
Bioinform. | 5 |
| 2016 | Estimating IBD tracts from low coverage NGS dataabstractMOTIVATION: The amount of IBD in an individual depends on the relatedness of the individual's parents. However, it can also provide information regarding mating system, past history and effective size of the population from which the individual has been sampled. RESULTS: Here, we present a new method for estimating inbreeding IBD tracts from low coverage NGS data. Contrary to other methods that use genotype data, the one presented here uses genotype likelihoods to take the uncertainty of the data into account. We benchmark it under a wide range of biologically relevant conditions and show that the new method provides a marked increase in accuracy even at low coverage. AVAILABILITY AND IMPLEMENTATION: The methods presented in this work were implemented in C/C ++ and are freely available for non-commercial use from https://github.com/fgvieira/ngsF-HMM CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Filipe G. Vieira, Anders Albrechtsen, Rasmus Nielsen |
Bioinform. | 3 |
| 2015 | Association Mapping for Compound Heterozygous Traits Using Phenotypic Distance and Integer Programming
Dan Gusfield, Rasmus Nielsen |
WABI | 2 |
| 2014 | ngsTools: methods for population genetics analyses from next-generation sequencing dataabstractSUMMARY: Next-generation sequencing technologies produce short reads that are either de novo assembled or mapped to a reference genome. Genotypes and/or single-nucleotide polymorphisms are then determined from the read composition at each site, which become the basis for many downstream analyses. However, for low sequencing depths, e.g. , there is considerable statistical uncertainty in the assignment of genotypes because of random sampling of homologous base pairs in heterozygotes and sequencing or alignment errors. Recently, several probabilistic methods have been proposed to account for this uncertainty and make accurate inferences from low quality and/or coverage sequencing data. We present ngsTools, a collection of programs to perform population genetics analyses from next-generation sequencing data. The methods implemented in these programs do not rely on single-nucleotide polymorphism or genotype calling and are particularly suitable for low sequencing depth data. AVAILABILITY: Programs included in ngsTools are implemented in C/C++ and are freely available for noncommercial use at https://github.com/mfumagalli/ngsTools. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary materials are available at Bioinformatics online. Matteo Fumagalli 0002, Filipe G. Vieira, Tyler Linderoth, Rasmus Nielsen |
Bioinform. | 4 |
| 2014 | bammds: a tool for assessing the ancestry of low-depth whole-genome data using multidimensional scaling (MDS)abstractSUMMARY: We present bammds, a practical tool that allows visualization of samples sequenced by second-generation sequencing when compared with a reference panel of individuals (usually genotypes) using a multidimensional scaling algorithm. Our tool is aimed at determining the ancestry of unknown samples-typical of ancient DNA data-particularly when only low amounts of data are available for those samples. AVAILABILITY AND IMPLEMENTATION: The software package is available under GNU General Public License v3 and is freely available together with test datasets https://savannah.nongnu.org/projects/bammds/. It is using R (http://www.r-project.org/), parallel (http://www.gnu.org/software/parallel/), samtools (https://github.com/samtools/samtools). CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Anna-Sapfo Malaspinas, Ole Tange, José Víctor Moreno-Mayar, Morten Rasmussen, Michael Degiorgio, Yong Wang 0075, Cristina E. Valdiosera, Gustavo Politis, Eske Willerslev, Rasmus Nielsen |
Bioinform. | 10 |
| 2014 | ANGSD: Analysis of Next Generation Sequencing DataabstractBACKGROUND: High-throughput DNA sequencing technologies are generating vast amounts of data. Fast, flexible and memory efficient implementations are needed in order to facilitate analyses of thousands of samples simultaneously. RESULTS: We present a multithreaded program suite called ANGSD. This program can calculate various summary statistics, and perform association mapping and population genetic analyses utilizing the full information in next generation sequencing data by working directly on the raw sequencing data or by using genotype likelihoods. CONCLUSIONS: The open source c/c++ program ANGSD is available at http://www.popgen.dk/angsd . The program is tested and validated on GNU/Linux systems. The program facilitates multiple input formats including BAM and imputed beagle genotype probability files. The program allow the user to choose between combinations of existing methods and can perform analysis that is not implemented elsewhere. Thorfinn Korneliussen, Anders Albrechtsen, Rasmus Nielsen |
BMC Bioinform. | 3 |
| 2013 | Calculation of Tajima's D and other neutrality test statistics from low depth next-generation sequencing dataabstractBACKGROUND: A number of different statistics are used for detecting natural selection using DNA sequencing data, including statistics that are summaries of the frequency spectrum, such as Tajima's D. These statistics are now often being applied in the analysis of Next Generation Sequencing (NGS) data. However, estimates of frequency spectra from NGS data are strongly affected by low sequencing coverage; the inherent technology dependent variation in sequencing depth causes systematic differences in the value of the statistic among genomic regions. RESULTS: We have developed an approach that accommodates the uncertainty of the data when calculating site frequency based neutrality test statistics. A salient feature of this approach is that it implicitly solves the problems of varying sequencing depth, missing data and avoids the need to infer variable sites for the analysis and thereby avoids ascertainment problems introduced by a SNP discovery process. CONCLUSION: Using an empirical Bayes approach for fast computations, we show that this method produces results for low-coverage NGS data comparable to those achieved when the genotypes are known without uncertainty. We also validate the method in an analysis of data from the 1000 genomes project. The method is implemented in a fast framework which enables researchers to perform these neutrality tests on a genome-wide scale. Thorfinn Korneliussen, Ida Moltke, Anders Albrechtsen, Rasmus Nielsen |
BMC Bioinform. | 4 |
| 2011 | Estimation of allele frequency and association mapping using next-generation sequencing dataabstractBACKGROUND: Estimation of allele frequency is of fundamental importance in population genetic analyses and in association mapping. In most studies using next-generation sequencing, a cost effective approach is to use medium or low-coverage data (e.g., < 15X). However, SNP calling and allele frequency estimation in such studies is associated with substantial statistical uncertainty because of varying coverage and high error rates. RESULTS: We evaluate a new maximum likelihood method for estimating allele frequencies in low and medium coverage next-generation sequencing data. The method is based on integrating over uncertainty in the data for each individual rather than first calling genotypes. This method can be applied to directly test for associations in case/control studies. We use simulations to compare the likelihood method to methods based on genotype calling, and show that the likelihood method outperforms the genotype calling methods in terms of: (1) accuracy of allele frequency estimation, (2) accuracy of the estimation of the distribution of allele frequencies across neutrally evolving sites, and (3) statistical power in association mapping studies. Using real re-sequencing data from 200 individuals obtained from an exon-capture experiment, we show that the patterns observed in the simulations are also found in real data. CONCLUSIONS: Overall, our results suggest that association mapping and estimation of allele frequencies should not be based on genotype calling in low to medium coverage data. Furthermore, if genotype calling methods are used, it is usually better not to filter genotypes based on the call confidence score. Su Yeon Kim, Kirk E. Lohmueller, Anders Albrechtsen, Yingrui Li, Thorfinn Korneliussen, Geng Tian, Niels Grarup, Gitte Andersen, Daniel Witte, Torben Jørgensen, Torben Hansen, Oluf Pedersen, Jun Wang 0004, Rasmus Nielsen |
BMC Bioinform. | 15 |
| 2007 | Finding cis-regulatory modules in Drosophila using phylogenetic hidden Markov modelsabstractMOTIVATION: Finding the regulatory modules for transcription factors binding is an important step in elucidating the complex molecular mechanisms underlying regulation of gene expression. There are numerous methods available for solving this problem, however, very few of them take advantage of the increasing availability of comparative genomic data. RESULTS: We develop a method for finding regulatory modules in Eukaryotic species using phylogenetic data. Using computer simulations and analysis of real data, we show that the use of phylogenetic hidden Markov model can lead to an increase in accuracy of prediction over methods that do not take advantage of the data from multiple species. AVAILABILITY: The new method is made accessible under GPL in a new publicly available JAVA program: EvoPromoter. It can be downloaded at http://sourceforge.net/projects/evopromoter/. Wendy S. W. Wong, Rasmus Nielsen |
Bioinform. | 2 |
| 2007 | Dependence of paracentric inversion rate on tract lengthabstractBACKGROUND: We develop a Bayesian method based on MCMC for estimating the relative rates of pericentric and paracentric inversions from marker data from two species. The method also allows estimation of the distribution of inversion tract lengths. RESULTS: We apply the method to data from Drosophila melanogaster and D. yakuba. We find that pericentric inversions occur at a much lower rate compared to paracentric inversions. The average paracentric inversion tract length is approx. 4.8 Mb with small inversions being more frequent than large inversions. If the two breakpoints defining a paracentric inversion tract are uniformly and independently distributed over chromosome arms there will be more short tract-length inversions than long; we find an even greater preponderance of short tract lengths than this would predict. Thus there appears to be a correlation between the positions of breakpoints which favors shorter tract lengths. CONCLUSION: The method developed in this paper provides the first statistical estimator for estimating the distribution of inversion tract lengths from marker data. Application of this method for a number of data sets may help elucidate the relationship between the length of an inversion and the chance that it will get accepted. Thomas L. York, Richard Durrett, Rasmus Nielsen |
BMC Bioinform. | 3 |
| 2006 | Identification of physicochemical selective pressure on protein encoding nucleotide sequencesabstractBACKGROUND: Statistical methods for identifying positively selected sites in protein coding regions are one of the most commonly used tools in evolutionary bioinformatics. However, they have been limited by not taking the physiochemical properties of amino acids into account. RESULTS: We develop a new codon-based likelihood model for detecting site-specific selection pressures acting on specific physicochemical properties. Nonsynonymous substitutions are divided into substitutions that differ with respect to the physicochemical properties of interest, and those that do not. The substitution rates of these two types of changes, relative to the synonymous substitution rate, are then described by two parameters, gamma and omega respectively. The new model allows us to perform likelihood ratio tests for positive selection acting on specific physicochemical properties of interest. The new method is first used to analyze simulated data and is shown to have good power and accuracy in detecting physicochemical selective pressure. We then re-analyze data from the class-I alleles of the human Major Histocompatibility Complex (MHC) and from the abalone sperm lysine. CONCLUSION: Our new method allows a more flexible framework to identify selection pressure on particular physicochemical properties. Wendy S. W. Wong, Raazesh Sainudiin, Rasmus Nielsen |
BMC Bioinform. | 3 |
| 2002 | PATRI-paternity inference using genetic dataabstractAbstract Summary: PATRI is a new application for paternity analysis using genetic data that accounts for the sampling fraction of potential fathers. Availability: Executables are available at www.biom.cornell.edu/Homepages/Rasmus_Nielsen/ Contact: [email protected] * To whom correspondence should be addressed. J. Signorovitch, Rasmus Nielsen |
Bioinform. | 2 |