EDBT 2026 Demo / reviewers in the wild / expert
Mikko J. Sillanpää
dblp:80/2571
· DBLP profile ↗
19ranked-venue papers
0as first author
14since 2021 · last 2026
0000-0003-2808-2768ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Proximal regularization of deep residual neural networks applied to high-dimensional genomic dataabstractHigh-dimensional genomic datasets contain complex patterns shaped by substantial biological noise, which pose major challenges for predictive modeling in genetics and breeding. Residual neural networks (ResNets) provide a powerful framework for capturing nonlinear genomic effects, but often overfit in settings where marker numbers greatly exceed sample sizes. As a solution, a range of regularization methods have been proposed. One promising approach relies on the proximal mapping technique, which is computationally efficient since it can be directly incorporated into the optimization algorithm. However, the performance of ResNets with various convex or non-convex proximal regularizers remains under-explored on high-dimensional data. In this study, we propose an extended stochastic adaptive proximal gradient ResNet method that can handle both convex and non-convex regularizers that range from $L_{0}$ to $L_{\infty }$ and give more analysis of the convergence guarantee for the convex and non-convex regularizers. Moreover, we evaluate the prediction performance in a supervised regression setting on four real high-dimensional genomic datasets from mice, pig, wheat, and loblolly pine. For comparison, we also implement and evaluate traditional sparse linear proximal methods with the same regularizers, as well as LightGBM. Experimental results demonstrate that an 18-layer ResNet with $L_{\frac{1}{2}}$ regularization outperforms other configurations on both mice and pig datasets. For the wheat and loblolly pine data, the 15-layer ResNet $L_{\frac{1}{2}}$ configuration achieves the lowest test mean squared errors and the highest distance correlation (dCor). These findings highlight the effectiveness of the regularized adaptive proximal gradient ResNet method and its potential for prediction tasks on high-dimensional genomic data. Yuhua Fan, Ilkka Launonen, Mikko J. Sillanpää, Patrik Waldmann |
Briefings Bioinform. | 3 |
| 2026 | GHSCM: Efficient maximum a posteriori inference for biological networks with the graphical horseshoe priorabstractGaussian graphical models are a popular choice for estimating biological networks, such as gene co-expression and protein-protein interaction networks. The data are assumed to follow a multivariate normal distribution, and the network is constructed from the precision matrix (inverse of the covariance matrix). The frequentist network estimation literature is vast and well-studied, with well-known and widely used methods, such as graphical LASSO. However, there is still room for new studies and methods in Bayesian literature, particularly in terms of accelerating computation. In this work, we propose a computationally efficient conditional maximization algorithm for Bayesian network estimation that utilizes the graphical horseshoe prior. We derive an approximation for the posterior edge probabilities, which are used to construct a sparse network estimate. We also introduce a simple, data-dimension-dependent formula for choosing the global scale parameter in the prior, which helps reduce computation time. With simulated and real-world examples, we show that the network estimates of our method are highly similar to those of the fully Bayesian Gibbs sampler with the graphical horseshoe prior. Still, our method is 50–100 times faster and arguably easier to use. We implemented the method in the R package GHSCM, available on GitHub at https://github.com/THautamaki/GHSCM . Tuomas Hautamäki, Aapo E. Korhonen, Olli Sarala, Markku Kuismin, Mikko J. Sillanpää |
Inf. Sci. | 5 |
| 2025 | HMFGraph: Novel Bayesian approach for recovering biological networksabstractGaussian graphical models (GGM) are powerful tools to examine partial correlation structures in high-dimensional omics datasets. Partial correlation networks can explain complex relationships between genes or other biological variables. Bayesian implementations of GGMs have recently received more attention. Usually, the most demanding parts of GGM implementations are: (i) hyperparameter tuning, (ii) edge selection, (iii) scalability for large datasets, and (iv) the prior choice for Bayesian GGM. To address these limitations, we introduce a novel Bayesian GGM using a hierarchical matrix-F prior with a fast implementation. We show, with extensive simulations and biological example analyses, that this prior has competitive network recovery capabilities compared to state-of-the-art approaches and good properties for recovering meaningful networks. We present a new way of tuning the shrinkage hyperparameter by constraining the condition number of the estimated precision matrix. For edge selection, we propose using approximated credible intervals (CI) whose width is controlled by the false discovery rate. An optimal CI is selected by maximizing an estimated F1-score via permutations. In addition, a specific choice of hyperparameter can make the proposed prior better suited for clustering and community detection. Our method, with a generalized expectation-maximization algorithm, computationally outperforms existing Bayesian GGM approaches that use Markov chain Monte Carlo algorithms. The method is implemented in the R package HMFGraph, found on GitHub at https://github.com/AapoKorhonen/HMFGraph. All codes to reproduce the results are found on GitHub at https://github.com/AapoKorhonen/HMFGraph-Supplementary. Aapo E. Korhonen, Olli Sarala, Tuomas Hautamäki, Markku Kuismin, Mikko J. Sillanpää |
PLoS Comput. Biol. | 5 |
| 2025 | tvsfglasso: Time-varying scale-free graphical lasso for network estimation from time-series dataabstractIn high-dimensional gene co-expression network analysis, capturing the temporal changes of gene associations is crucial for unveiling dynamic regulatory mechanisms inherent in biological systems. Examining how these interactions change over time offers valuable insights into the developmental and adaptive processes that drive an organism's lifecycle. Moreover, incorporating structural prior information can substantially enhance the accuracy and interpretability of the estimated sparse dynamic gene network. Methods previously proposed in the literature cannot simultaneously model sparse time-varying co-expression network structure and have the power-law degree distribution. Additionally, there is a demand of time-efficient, memory-light software implementations and possibility to utilize repeated measures at each time-point (if available). In this paper, we introduce the time-varying scale-free graphical lasso (tvsfglasso), a novel scalable framework for estimating high-dimensional time-varying gene co-expression networks under the assumption that these networks simultaneously exhibit sparse and a scale-free structure. We utilize fast algorithms developed for the graphical lasso (glasso), which makes tvsfglasso a scalable tool for high-dimensional problems. We evaluate the performance of tvsfglasso using both simulated and real-world dynamic gene expression time series datasets, demonstrating its capability to detect temporal changes in gene associations. Our results highlight the potential of tvsfglasso to advance the understanding of dynamic gene networks, making this estimator useful for more accurate modeling of complex biological processes. Markku Kuismin, Mikko J. Sillanpää |
PLoS Comput. Biol. | 2 |
| 2025 | Adaptive and Self-Tuning SBL With Total Variation Priors for Block-Sparse Signal RecoveryabstractThis letter addresses the problem of estimating block sparse signal with unknown group partitions in a multiple measurement vector (MMV) setup. We propose a Bayesian framework by applying an adaptive total variation (TV) penalty on the hyper-parameter space of the sparse signal. The main contributions are two-fold. 1) We extend the TV penalty beyond the immediate neighbor, thus enabling better capture of the signal structure. 2) A dynamic framework is provided to learn the regularization weights for the TV penalty based on the statistical dependencies between the entries of tentative blocks, thus eliminating the need for fine-tuning. The superior performance of the proposed method is empirically demonstrated by extensive computer simulations with the state-of-art benchmarks. The proposed solution exhibits both excellent performance and robustness against sparsity model mismatch. Hamza Djelouat, Reijo Leinonen, Mikko J. Sillanpää, Bhaskar D. Rao, Markku Juntti |
IEEE Signal Process. Lett. | 3 |
| 2024 | Capacitated spatial clustering with multiple constraints and attributesabstractCapacitated spatial clustering, a type of unsupervised machine learning method, is often used to tackle problems in compressing data, classification, logistic optimization and infrastructure optimization. Depending on the application at hand, a multitude of extensions to the clustering problem may be necessary. In this article, we propose a number of novel extensions to PACK, a recent capacitated partitional spatial clustering method which uses an optimization algorithm that is based on linear programming tasks. These extensions relate to the relocation and location preference of cluster centers, outliers, and non-spatial attributes, and they can be considered jointly. In the context of edge server placement, these improve the spatial location of servers while considering, for example, application placement on the servers in response to spatial application usage patterns. We demonstrate the usefulness of an extended version of PACK with an example with simulated data, as well as a real world example in edge server placement for a city region with various different setups. These setups are evaluated with summary statistics about spatial proximity and attribute similarity. As a result, the similarity of the clusters was improved by 53% at best while simultaneously the proximity degraded only by 18%. The extensions provide valuable means for including non-spatial information in the cluster analysis, and to attain better overall proximity and similarity. Tero Lähderanta, Lauri Lovén, Leena Ruha, Teemu Leppänen, Ilkka Launonen, Jukka Riekki, Mikko J. Sillanpää |
Eng. Appl. Artif. Intell. | 7 |
| 2024 | Hierarchical MTC User Activity Detection and Channel Estimation With Unknown Spatial CovarianceabstractThis paper addresses the joint user identification and channel estimation (JUICE) problem in machine-type communications under the practical spatially correlated channels model with unknown covariance matrices. Furthermore, we consider an MTC network with hierarchical user activity patterns following an event-triggered traffic mode. Therein the users are distributed over clusters with a structured sporadic activity behavior that exhibits both cluster-level and intra-cluster sparsity patterns. To solve the JUICE problem, we first leverage the concept of strong priors and propose a hierarchical-sparsity-inducing spike-and-slab prior to model the structured sparse activity pattern. Subsequently, we derive a Bayesian inference scheme by coupling the expectation propagation (EP) algorithm with the expectation maximization (EM) framework. Second, we reformulate the JUICE as a maximum a posteriori (MAP) estimation problem and propose a computationally-efficient solution based on the alternating direction method of multipliers (ADMM). More precisely, we relax the strong spike-and-slab prior with a cluster-sparsity-promoting prior based on the long-sum penalty. We then derive an ADMM algorithm that solves the MAP problem through a sequence of closed-form updates. Numerical results highlight the significant performance gains obtained by the proposed algorithms, as well as their robustness against various assumptions on the sparse activity behavior of the users. Hamza Djelouat, Mikko J. Sillanpää, Markus Leinonen, Markku Juntti |
IEEE Trans. Wirel. Commun. | 2 |
| 2023 | Machine Learning-Aided Piece-Wise Modeling Technique of Power Amplifier for Digital PredistortionabstractWe propose a new power amplifier (PA) behavioral modeling approach, to characterize and compensate for the signal quality degrading effects induced by a PA with a machine learning (ML) aided piece-wise (PW) modeling approach. Instead of using a single pruned Volterra model, we use multiple small-size pruned Volterra models by classifying the input data into different classes. For that purpose, an ML classifier model is trained by extracting some crucial features from both the input signal statistics and the PA operating point. The simulation results indicate that our approach contributes to an improved performance/complexity trade-off than a single generalized memory polynomial (GMP) model in terms of PA behavior modeling and linearization. S. S. Krishna Chaitanya Bulusu, Nuutti Tervo, Praneeth Susarla, Mikko J. Sillanpää, Olli Silvén, Markku Juntti, Aarno Pärssinen |
ICASSP | 4 |
| 2023 | Genetic fine-mapping from summary data using a nonlocal prior improves the detection of multiple causal variantsabstractMOTIVATION: Genome-wide association studies (GWAS) have been successful in identifying genomic loci associated with complex traits. Genetic fine-mapping aims to detect independent causal variants from the GWAS-identified loci, adjusting for linkage disequilibrium patterns. RESULTS: We present "FiniMOM" (fine-mapping using a product inverse-moment prior), a novel Bayesian fine-mapping method for summarized genetic associations. For causal effects, the method uses a nonlocal inverse-moment prior, which is a natural prior distribution to model non-null effects in finite samples. A beta-binomial prior is set for the number of causal variants, with a parameterization that can be used to control for potential misspecifications in the linkage disequilibrium reference. The results of simulations studies aimed to mimic a typical GWAS on circulating protein levels show improved credible set coverage and power of the proposed method over current state-of-the-art fine-mapping method SuSiE, especially in the case of multiple causal variants within a locus. AVAILABILITY AND IMPLEMENTATION: https://vkarhune.github.io/finimom/. Ville Karhunen, Ilkka Launonen, Marjo-Riitta Järvelin, Sylvain Sebert, Mikko J. Sillanpää |
Bioinform. | 5 |
| 2023 | BELMM: Bayesian model selection and random walk smoothing in time-series clusteringabstractMOTIVATION: Due to advances in measuring technology, many new phenotype, gene expression, and other omics time-course datasets are now commonly available. Cluster analysis may provide useful information about the structure of such data. RESULTS: In this work, we propose BELMM (Bayesian Estimation of Latent Mixture Models): a flexible framework for analysing, clustering, and modelling time-series data in a Bayesian setting. The framework is built on mixture modelling: first, the mean curves of the mixture components are assumed to follow random walk smoothing priors. Second, we choose the most plausible model and the number of mixture components using the Reversible-jump Markov chain Monte Carlo. Last, we assign the individual time series into clusters based on the similarity to the cluster-specific trend curves determined by the latent random walk processes. We demonstrate the use of fast and slow implementations of our approach on both simulated and real time-series data using widely available software R, Stan, and CU-MSDSp. AVAILABILITY AND IMPLEMENTATION: The French mortality dataset is available at http://www.mortality.org, the Drosophila melanogaster embryogenesis gene expression data at https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE121160. Details on our simulated datasets are available in the Supplementary Material, and R scripts and a detailed tutorial on GitHub at https://github.com/ollisa/BELMM. The software CU-MSDSp is available on GitHub at https://github.com/jtchavisIII/CU-MSDSp. Olli Sarala, Tanja Pyhäjärvi, Mikko J. Sillanpää |
Bioinform. | 3 |
| 2023 | Sequential Model Correction for Nonlinear Inverse ProblemsabstractAbstract. Inverse problems are in many cases solved with optimization techniques. When the underlying model is linear, first-order gradient methods are usually sufficient. With nonlinear models, due to nonconvexity, one must often resort to second-order methods that are computationally more expensive. In this work we aim to approximate a nonlinear model with a linear one and correct the resulting approximation error. We develop a sequential method that iteratively solves a linear inverse problem and updates the approximation error by evaluating it at the new solution. This treatment convexifies the problem and allows us to benefit from established convex optimization methods. We separately consider cases where the approximation is fixed over iterations and where the approximation is adaptive. In the fixed case we show theoretically under what assumptions the sequence converges. In the adaptive case, particularly considering the special case of approximation by first-order Taylor expansion, we show that with certain assumptions the sequence converges to a critical point of the original nonconvex functional. Furthermore, we show that with quadratic objective functions the sequence corresponds to the Gauss–Newton method. Finally, we showcase numerical results superior to the conventional model correction method. We also show that a fixed approximation can provide competitive results with considerable computational speed-up. Arttu Arjas, Mikko J. Sillanpää, Andreas Hauptmann |
SIAM J. Imaging Sci. | 2 |
| 2021 | MCPeSe: Monte Carlo penalty selection for graphical lassoabstractMOTIVATION: Graphical lasso (Glasso) is a widely used tool for identifying gene regulatory networks in systems biology. However, its computational efficiency depends on the choice of regularization parameter (tuning parameter), and selecting this parameter can be highly time consuming. Although fully Bayesian implementations of Glasso alleviate this problem somewhat by specifying a priori distribution for the parameter, these approaches lack the scalability of their frequentist counterparts. RESULTS: Here, we present a new Monte Carlo Penalty Selection method (MCPeSe), a computationally efficient approach to regularization parameter selection for Glasso. MCPeSe combines the scalability and low computational cost of the frequentist Glasso with the ability to automatically choose the regularization by Bayesian Glasso modeling. MCPeSe provides a state-of-the-art 'tuning-free' model selection criterion for Glasso and allows exploration of the posterior probability distribution of the tuning parameter. AVAILABILITY AND IMPLEMENTATION: R source code of MCPeSe, a step by step example showing how to apply MCPeSe and a collection of scripts used to prepare the material in this article are publicly available at GitHub under GPL (https://github.com/markkukuismin/MCPeSe/). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Markku Kuismin, Mikko J. Sillanpää |
Bioinform. | 2 |
| 2021 | Edge computing server placement with capacitated location allocationabstractThe deployment of edge computing infrastructure requires a careful placement of the edge servers, with an aim to improve application latencies and reduce data transfer load in opportunistic Internet of Things systems. In the edge server placement, it is important to consider computing capacity, available deployment budget, and hardware requirements for the edge servers and the underlying backbone network topology. In this paper, we thoroughly survey the existing literature in edge server placement, identify gaps and present an extensive set of parameters to be considered. We then develop a novel algorithm, called PACK, for server placement as a capacitated location–allocation problem. PACK minimizes the distances between servers and their associated access points, while taking into account capacity constraints for load balancing and enabling workload sharing between servers. Moreover, PACK considers practical issues such as prioritized locations and reliability. We evaluate the algorithm in two distinct scenarios: one with high capacity servers for edge computing in general, and one with low capacity servers for Fog computing. Evaluations are performed with a data set collected in a real-world network, consisting of both dense and sparse deployments of access points across a city area. The resulting algorithm and related tools are publicly available as open source software. Tero Lähderanta, Teemu Leppänen, Leena Ruha, Lauri Lovén, Erkki Harjula, Mika Ylianttila, Jukka Riekki, Mikko J. Sillanpää |
J. Parallel Distributed Comput. | 8 |
| 2021 | Model guided trait-specific co-expression network estimation as a new perspective for identifying molecular interactions and pathwaysabstractA wide variety of 1) parametric regression models and 2) co-expression networks have been developed for finding gene-by-gene interactions underlying complex traits from expression data. While both methodological schemes have their own well-known benefits, little is known about their synergistic potential. Our study introduces their methodological fusion that cross-exploits the strengths of individual approaches via a built-in information-sharing mechanism. This fusion is theoretically based on certain trait-conditioned dependency patterns between two genes depending on their role in the underlying parametric model. Resulting trait-specific co-expression network estimation method 1) serves to enhance the interpretation of biological networks in a parametric sense, and 2) exploits the underlying parametric model itself in the estimation process. To also account for the substantial amount of intrinsic noise and collinearities, often entailed by expression data, a tailored co-expression measure is introduced along with this framework to alleviate related computational problems. A remarkable advance over the reference methods in simulated scenarios substantiate the method's high-efficiency. As proof-of-concept, this synergistic approach is successfully applied in survival analysis, with acute myeloid leukemia data, further highlighting the framework's versatility and broad practical relevance. Juho A. J. Kontio, Tanja Pyhäjärvi, Mikko J. Sillanpää |
PLoS Comput. Biol. | 3 |
| 2020 | Estimation of dynamic SNP-heritability with Bayesian Gaussian process modelsabstractMOTIVATION: Improved DNA technology has made it practical to estimate single-nucleotide polymorphism (SNP)-heritability among distantly related individuals with unknown relationships. For growth- and development-related traits, it is meaningful to base SNP-heritability estimation on longitudinal data due to the time-dependency of the process. However, only few statistical methods have been developed so far for estimating dynamic SNP-heritability and quantifying its full uncertainty. RESULTS: We introduce a completely tuning-free Bayesian Gaussian process (GP)-based approach for estimating dynamic variance components and heritability as their function. For parameter estimation, we use a modern Markov Chain Monte Carlo method which allows full uncertainty quantification. Several datasets are analysed and our results clearly illustrate that the 95% credible intervals of the proposed joint estimation method (which 'borrows strength' from adjacent time points) are significantly narrower than of a two-stage baseline method that first estimates the variance components at each time point independently and then performs smoothing. We compare the method with a random regression model using MTG2 and BLUPF90 software and quantitative measures indicate superior performance of our method. Results are presented for simulated and real data with up to 1000 time points. Finally, we demonstrate scalability of the proposed method for simulated data with tens of thousands of individuals. AVAILABILITY AND IMPLEMENTATION: The C++ implementation dynBGP and simulated data are available in GitHub: https://github.com/aarjas/dynBGP. The programmes can be run in R. Real datasets are available in QTL archive: https://phenome.jax.org/centers/QTLA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Arttu Arjas, Andreas Hauptmann, Mikko J. Sillanpää |
Bioinform. | 3 |
| 2019 | A Gaussian process model and Bayesian variable selection for mapping function-valued quantitative traits with incomplete phenotypic dataabstractMOTIVATION: Recent advances in high dimensional phenotyping bring time as an extra dimension into the phenotypes. This promotes the quantitative trait locus (QTL) studies of function-valued traits such as those related to growth and development. Existing approaches for analyzing functional traits utilize either parametric methods or semi-parametric approaches based on splines and wavelets. However, very limited choices of software tools are currently available for practical implementation of functional QTL mapping and variable selection. RESULTS: We propose a Bayesian Gaussian process (GP) approach for functional QTL mapping. We use GPs to model the continuously varying coefficients which describe how the effects of molecular markers on the quantitative trait are changing over time. We use an efficient gradient based algorithm to estimate the tuning parameters of GPs. Notably, the GP approach is directly applicable to the incomplete datasets having even larger than 50% missing data rate (among phenotypes). We further develop a stepwise algorithm to search through the model space in terms of genetic variants, and use a minimal increase of Bayesian posterior probability as a stopping rule to focus on only a small set of putative QTL. We also discuss the connection between GP and penalized B-splines and wavelets. On two simulated and three real datasets, our GP approach demonstrates great flexibility for modeling different types of phenotypic trajectories with low computational cost. The proposed model selection approach finds the most likely QTL reliably in tested datasets. AVAILABILITY AND IMPLEMENTATION: Software and simulated data are available as a MATLAB package 'GPQTLmapping', and they can be downloaded from GitHub (https://github.com/jpvanhat/GPQTLmapping). Real datasets used in case studies are publicly available at QTL Archive. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jarno Vanhatalo, Mikko J. Sillanpää |
Bioinform. | 3 |
| 2011 | Estimating Haplotype Frequencies by Combining Data from Large DNA Pools with Database InformationabstractWe assume that allele frequency data have been extracted from several large DNA pools, each containing genetic material of up to hundreds of sampled individuals. Our goal is to estimate the haplotype frequencies among the sampled individuals by combining the pooled allele frequency data with prior knowledge about the set of possible haplotypes. Such prior information can be obtained, for example, from a database such as HapMap. We present a Bayesian haplotyping method for pooled DNA based on a continuous approximation of the multinomial distribution. The proposed method is applicable when the sizes of the DNA pools and/or the number of considered loci exceed the limits of several earlier methods. In the example analyses, the proposed model clearly outperforms a deterministic greedy algorithm on real data from the HapMap database. With a small number of loci, the performance of the proposed method is similar to that of an EM-algorithm, which uses a multinormal approximation for the pooled allele frequencies, but which does not utilize prior information about the haplotypes. The method has been implemented using Matlab and the code is available upon request from the authors. Dario Gasbarra, Sangita Kulathinal, Matti Pirinen, Mikko J. Sillanpää |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2007 | Estimating genealogies from linked marker data: a Bayesian approachabstractBACKGROUND: Answers to several fundamental questions in statistical genetics would ideally require knowledge of the ancestral pedigree and of the gene flow therein. A few examples of such questions are haplotype estimation, relatedness and relationship estimation, gene mapping by combining pedigree and linkage disequilibrium information, and estimation of population structure. RESULTS: We present a probabilistic method for genealogy reconstruction. Starting with a group of genotyped individuals from some population isolate, we explore the state space of their possible ancestral histories under our Bayesian model by using Markov chain Monte Carlo (MCMC) sampling techniques. The main contribution of our work is the development of sampling algorithms in the resulting vast state space with highly dependent variables. The main drawback is the computational complexity that limits the time horizon within which explicit reconstructions can be carried out in practice. CONCLUSION: The estimates for IBD (identity-by-descent) and haplotype distributions are tested in several settings using simulated data. The results appear to be promising for a further development of the method. Dario Gasbarra, Matti Pirinen, Mikko J. Sillanpää, Elja Arjas |
BMC Bioinform. | 3 |
| 2004 | BAPS 2: enhanced possibilities for the analysis of genetic population structureabstractUNLABELLED: Bayesian statistical methods based on simulation techniques have recently been shown to provide powerful tools for the analysis of genetic population structure. We have previously developed a Markov chain Monte Carlo (MCMC) algorithm for characterizing genetically divergent groups based on molecular markers and geographical sampling design of the dataset. However, for large-scale datasets such algorithms may get stuck to local maxima in the parameter space. Therefore, we have modified our earlier algorithm to support multiple parallel MCMC chains, with enhanced features that enable considerably faster and more reliable estimation compared to the earlier version of the algorithm. We consider also a hierarchical tree representation, from which a Bayesian model-averaged structure estimate can be extracted. The algorithm is implemented in a computer program that features a user-friendly interface and built-in graphics. The enhanced features are illustrated by analyses of simulated data and an extensive human molecular dataset. AVAILABILITY: Freely available at http://www.rni.helsinki.fi/~jic/bapspage.html. Jukka Corander, Patrik Waldmann, Pekka Marttinen, Mikko J. Sillanpää |
Bioinform. | 4 |