Jukka Corander

dblp:94/2872 · DBLP profile ↗
← Back
47ranked-venue papers
6as first author
3since 2021 · last 2024
0000-0002-7752-1942ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 22 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 15 · 1 first-author · 1 since 2021Theory of computation · 5 · 2 first-authorDatabases, data management, data science and information retrieval · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3Computer networks · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Probabilistic and Bayesian machine learning · 94% Optimization for machine learning · 6%
Interdisciplinary, comprehensive, and emerging computing
8 papers
Bioinformatics and computational biology · 98% Computational science and engineering · 2%
Theoretical computer science
4 papers
Coding theory · 76% Information theory · 20% Computational complexity · 2%
Computer networks
2 papers
Physical-layer communications · 82% Datacenter networks · 18%
Human-computer interaction and pervasive computing
1 paper
Usability and user experience research · 50% Interaction techniques and input · 50%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 30 heaviest of 44, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › approximate bayesian inference
simulation-based inference
1.332024
Approximate Bayesian inference from noisy likelihoods with Gaussian process emulated MCMC · J. Mach. Learn. Res. 2024
ELFI: Engine for Likelihood-Free Inference · J. Mach. Learn. Res. 2018
Bayesian Optimization for Likelihood-Free Inference of Simulator-Based Statistical Models · J. Mach. Learn. Res. 2016
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
approximate bayesian inference
0.812024
Approximate Bayesian inference from noisy likelihoods with Gaussian process emulated MCMC · J. Mach. Learn. Res. 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.812024
Approximate Bayesian inference from noisy likelihoods with Gaussian process emulated MCMC · J. Mach. Learn. Res. 2024
Bioinformatics and computational biology
population genetics
0.732018
Bacmeta: simulator for genomic evolution in bacterial metapopulations · Bioinform. 2018
Kpax3: Bayesian bi-clustering of large sequence datasets · Bioinform. 2018
BAPS 2: enhanced possibilities for the analysis of genetic population structure · Bioinform. 2004
Physical-layer communications
MIMO
0.522017
Asymptotic Analysis of Rayleigh Product Channels: A Free Probability Approach · IEEE Trans. Inf. Theory 2017
On the Outage Capacity of Orthogonal Space-Time Block Codes Over Multi-Cluster Scattering MIMO Channels · IEEE Trans. Commun. 2015
Bioinformatics and computational biology › sequence analysis › sequence assembly
plasmid reconstruction
0.412020
gplas: a comprehensive tool for plasmid analysis using short-read graphs · Bioinform. 2020
Bioinformatics and computational biology › population genetics
ancestry inference
0.422018
Kpax3: Bayesian bi-clustering of large sequence datasets · Bioinform. 2018
BAPS 2: enhanced possibilities for the analysis of genetic population structure · Bioinform. 2004
Bioinformatics and computational biology › gene expression analysis
biclustering
0.312018
Kpax3: Bayesian bi-clustering of large sequence datasets · Bioinform. 2018
Bioinformatics and computational biology › genomics
genome-wide association study
0.312018
pyseer: a comprehensive tool for microbial pangenome-wide association studies · Bioinform. 2018
Bioinformatics and computational biology › population genetics
wright-fisher model
0.312018
Bacmeta: simulator for genomic evolution in bacterial metapopulations · Bioinform. 2018
Machine learning › Probabilistic and Bayesian machine learning › clustering
bayesian clustering
0.322015
A Bayesian Predictive Model for Clustering Data of Mixed Discrete and Continuous Type · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Bayesian Clustering of Fuzzy Feature Vectors Using a Quasi-Likelihood Approach · IEEE Trans. Pattern Anal. Mach. Intell. 2009
Usability and user experience research
cognitive modeling
0.312017
Inferring Cognitive Models from Data using Approximate Bayesian Computation · CHI 2017
Interaction techniques and input › selection techniques › command selection
menu interaction
0.312017
Inferring Cognitive Models from Data using Approximate Bayesian Computation · CHI 2017
Physical-layer communications › MIMO › MIMO capacity
mutual information distribution
0.312017
Asymptotic Analysis of Rayleigh Product Channels: A Free Probability Approach · IEEE Trans. Inf. Theory 2017
Datacenter networks
remote procedure calls
0.312017
Asymptotic Analysis of Rayleigh Product Channels: A Free Probability Approach · IEEE Trans. Inf. Theory 2017
Coding theory › error-correcting codes
coding bounds
0.312017
From Random Matrix Theory to Coding Theory: Volume of a Metric Ball in Unitary Group · IEEE Trans. Inf. Theory 2017
Coding theory › error-correcting codes › coding bounds › minimum distance bounds
gilbert-varshamov bound
0.312017
From Random Matrix Theory to Coding Theory: Volume of a Metric Ball in Unitary Group · IEEE Trans. Inf. Theory 2017
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization
0.212016
Bayesian Optimization for Likelihood-Free Inference of Simulator-Based Statistical Models · J. Mach. Learn. Res. 2016
Data mining
clustering
0.212016
Low-Rank Doubly Stochastic Matrix Decomposition for Cluster Analysis · J. Mach. Learn. Res. 2016
Data mining › clustering › clustering evaluation
cluster number estimation
0.212016
Low-Rank Doubly Stochastic Matrix Decomposition for Cluster Analysis · J. Mach. Learn. Res. 2016
Coding theory › error-correcting codes › code construction
codebook design
0.212016
Volume of Metric Balls in High-Dimensional Complex Grassmann Manifolds · IEEE Trans. Inf. Theory 2016
Coding theory › network coding › subspace codes
grassmannian codes
0.212016
Volume of Metric Balls in High-Dimensional Complex Grassmann Manifolds · IEEE Trans. Inf. Theory 2016
Coding theory › source coding
rate-distortion theory
0.212016
Volume of Metric Balls in High-Dimensional Complex Grassmann Manifolds · IEEE Trans. Inf. Theory 2016
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.212024
Approximate Bayesian inference from noisy likelihoods with Gaussian process emulated MCMC · J. Mach. Learn. Res. 2024
Physical-layer communications
channel modeling
0.212015
On the Outage Capacity of Orthogonal Space-Time Block Codes Over Multi-Cluster Scattering MIMO Channels · IEEE Trans. Commun. 2015
Physical-layer communications › MIMO › space-time coding
space-time block codes
0.212015
On the Outage Capacity of Orthogonal Space-Time Block Codes Over Multi-Cluster Scattering MIMO Channels · IEEE Trans. Commun. 2015
Information theory
channel capacity
0.212015
On the Outage Capacity of Orthogonal Space-Time Block Codes Over Multi-Cluster Scattering MIMO Channels · IEEE Trans. Commun. 2015
Information theory › channel capacity › fading channel
outage capacity
0.212015
On the Outage Capacity of Orthogonal Space-Time Block Codes Over Multi-Cluster Scattering MIMO Channels · IEEE Trans. Commun. 2015
Bioinformatics and computational biology
metagenomics
0.212014
SEK: sparsity exploiting k-mer-based estimation of bacterial community composition · Bioinform. 2014
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.212013
Learning Chordal Markov Networks by Constraint Satisfaction · NIPS 2013

Methods — techniques the papers use, named apart from their topics

gaussian process · 1.0sequential experimental design · 0.8metropolis-hastings · 0.8random matrix theory · 0.7bayesian optimization · 0.6asymptotic analysis · 0.5closed-form approximation · 0.4network partitioning · 0.4assembly graph analysis · 0.4wright-fisher model · 0.3surrogate modeling · 0.3stochastic simulation · 0.3split-merge sampler · 0.3population structure correction · 0.3linear mixed model · 0.3gibbs sampler · 0.3bayesian inference · 0.3answer set programming · 0.3
YearPublicationVenuePosition
2024 Approximate Bayesian inference from noisy likelihoods with Gaussian process emulated MCMC
abstract
We present a framework for approximate Bayesian inference intended for a situation where only a limited number of noisy log-likelihood evaluations can be obtained due to constraints on the available computational budget, which is becoming increasingly common for expensive simulator-based models. We model the log-likelihood function using a Gaussian process (GP) and our main methodological innovation is to apply this model to emulate the progression that an exact Metropolis-Hastings (MH) sampler would take if it was applicable. Informative log-likelihood evaluation locations are selected using a sequential experimental design strategy until the MH accept/reject decisions are performed with sufficient level of accuracy based on a prespecified error tolerance criterion. The resulting approximate sampler is conceptually simple and shown to be sample-efficient. It is also more robust compared with earlier “Bayesian optimisation-like” methods tailored for approximate Bayesian inference, which generally assume a global surrogate model across the parameter space that can be challenging to fit well. We discuss some theoretical aspects and various interpretations of the resulting approximate MH sampler, and demonstrate its benefits in the context of Bayesian and generalised Bayesian likelihood-free inference for simulator-based statistical models.
Marko Järvenpää, Jukka Corander
J. Mach. Learn. Res.2
2022 Identification of multiplicatively acting modulatory mutational signatures in cancer
abstract
BACKGROUND: A deep understanding of carcinogenesis at the DNA level underpins many advances in cancer prevention and treatment. Mutational signatures provide a breakthrough conceptualisation, as well as an analysis framework, that can be used to build such understanding. They capture somatic mutation patterns and at best identify their causes. Most studies in this context have focused on an inherently additive analysis, e.g. by non-negative matrix factorization, where the mutations within a cancer sample are explained by a linear combination of independent mutational signatures. However, other recent studies show that the mutational signatures exhibit non-additive interactions. RESULTS: We carefully analysed such additive model fits from the PCAWG study cataloguing mutational signatures as well as their activities across thousands of cancers. Our analysis identified systematic and non-random structure of residuals that is left unexplained by the additive model. We used hierarchical clustering to identify cancer subsets with similar residual profiles to show that both systematic mutation count overestimation and underestimation take place. We propose an extension to the additive mutational signature model-multiplicatively acting modulatory processes-and develop a maximum-likelihood framework to identify such modulatory mutational signatures. The augmented model is expressive enough to almost fully remove the observed systematic residual patterns. CONCLUSION: We suggest the modulatory processes biologically relate to sample specific DNA repair propensities with cancer or tissue type specific profiles. Overall, our results identify an interesting direction where to expand signature analysis.
Dovydas Kiciatovas, Qingli Guo, Miika Kailas, Henri Pesonen, Jukka Corander, Samuel Kaski, Esa Pitkänen, Ville Mustonen
BMC Bioinform.5
2021 Boosting heritability: estimating the genetic component of phenotypic variation with multiple sample splitting
abstract
BACKGROUND: Heritability is a central measure in genetics quantifying how much of the variability observed in a trait is attributable to genetic differences. Existing methods for estimating heritability are most often based on random-effect models, typically for computational reasons. The alternative of using a fixed-effect model has received much more limited attention in the literature. RESULTS: In this paper, we propose a generic strategy for heritability inference, termed as "boosting heritability", by combining the advantageous features of different recent methods to produce an estimate of the heritability with a high-dimensional linear model. Boosting heritability uses in particular a multiple sample splitting strategy which leads in general to a stable and accurate estimate. We use both simulated data and real antibiotic resistance data from a major human pathogen, Sptreptococcus pneumoniae, to demonstrate the attractive features of our inference strategy. CONCLUSIONS: Boosting is shown to offer a reliable and practically useful tool for inference about heritability.
The Tien Mai, Jukka Corander
BMC Bioinform.3
2020 gplas: a comprehensive tool for plasmid analysis using short-read graphs
abstract
SUMMARY: Plasmids can horizontally transmit genetic traits, enabling rapid bacterial adaptation to new environments and hosts. Short-read whole-genome sequencing data are often applied to large-scale bacterial comparative genomics projects but the reconstruction of plasmids from these data is facing severe limitations, such as the inability to distinguish plasmids from each other in a bacterial genome. We developed gplas, a new approach to reliably separate plasmid contigs into discrete components using sequence composition, coverage, assembly graph information and network partitioning based on a pruned network of plasmid unitigs. Gplas facilitates the analysis of large numbers of bacterial isolates and allows a detailed analysis of plasmid epidemiology based solely on short-read sequence data. AVAILABILITY AND IMPLEMENTATION: Gplas is written in R, Bash and uses a Snakemake pipeline as a workflow management system. Gplas is available under the GNU General Public License v3.0 at https://gitlab.com/sirarredondo/gplas.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sergio Arredondo-Alonso, Martin Bootsma, Yaïr Hein, Malbert R. C. Rogers, Jukka Corander, Rob J. L. Willems, Anita C. Schürch
Bioinform.5
2019 A logical approach to context-specific independence
Jukka Corander, Antti Hyttinen, Juha Kontinen, Johan Pensar, Jouko A. Väänänen
Ann. Pure Appl. Log.1
2019 Doubly Stochastic Neighbor Embedding on Spheres
abstract
Stochastic Neighbor Embedding (SNE) methods minimize the divergence between the similarity matrix of a high-dimensional data set and its counterpart from a low-dimensional embedding, leading to widely applied tools for data visualization. Despite their popularity, the current SNE methods experience a crowding problem when the data include highly imbalanced similarities. This implies that the data points with higher total similarity tend to get crowded around the display center. To solve this problem, we introduce a fast normalization method and normalize the similarity matrix to be doubly stochastic such that all the data points have equal total similarities. Furthermore, we show empirically and theoretically that the doubly stochasticity constraint often leads to embeddings which are approximately spherical. This suggests replacing a flat space with spheres as the embedding space. The spherical embedding eliminates the discrepancy between the center and the periphery in visualization, which efficiently resolves the crowding problem. We compared the proposed method (DOSNES) with the state-of-the-art SNE method on three real-world datasets and the results clearly indicate that our method is more favorable in terms of visualization quality. DOSNES is freely available at http://yaolubrain.github.io/dosnes/.
Yao Lu 0027, Jukka Corander, Zhirong Yang
Pattern Recognit. Lett.2
2018 pyseer: a comprehensive tool for microbial pangenome-wide association studies
abstract
Summary: Genome-wide association studies (GWAS) in microbes have different challenges to GWAS in eukaryotes. These have been addressed by a number of different methods. pyseer brings these techniques together in one package tailored to microbial GWAS, allows greater flexibility of the input data used, and adds new methods to interpret the association results. Availability and implementation: pyseer is written in python and is freely available at https://github.com/mgalardini/pyseer, or can be installed through pip. Documentation and a tutorial are available at http://pyseer.readthedocs.io. Supplementary information: Supplementary data are available at Bioinformatics online.
John A. Lees, Marco Galardini, Stephen D. Bentley, Jeffrey N. Weiser, Jukka Corander
Bioinform.5
2018 Kpax3: Bayesian bi-clustering of large sequence datasets
abstract
Motivation: Estimation of the hidden population structure is an important step in many genetic studies. Often the aim is also to identify which sequence locations are the most discriminative between groups of samples for a given data partition. Automated discovery of interesting patterns that are present in the data can help to generate new biological hypotheses. Results: We introduce Kpax3, a Bayesian method for bi-clustering multiple sequence alignments. Influence of individual sites will be determined in a supervised manner by using informative prior distributions for the model parameters. Our inference method uses an implementation of both split-merge and Gibbs sampler type MCMC algorithms to traverse the joint posterior of partitions of samples and variables. We use a large Rotavirus sequence dataset to demonstrate the ability of Kpax3 to generate biologically important hypotheses about differential selective pressures across a virus protein. Availability and implementation: Kpax3 is implemented as a Julia package and released under the MIT license. Source code and documentation are available at: https://github.com/albertopessia/Kpax3.jl. Supplementary information: Supplementary data are available at Bioinformatics online.
Alberto Pessia, Jukka Corander
Bioinform.2
2018 Bacmeta: simulator for genomic evolution in bacterial metapopulations
abstract
Summary: The advent of genomic data from densely sampled bacterial populations has created a need for flexible simulators by which models and hypotheses can be efficiently investigated in the light of empirical observations. Bacmeta provides fast stochastic simulation of neutral evolution within a large collection of interconnected bacterial populations with completely adjustable connectivity network. Stochastic events of mutations, recombinations, insertions/deletions, migrations and micro-epidemics can be simulated in discrete non-overlapping generations with a Wright-Fisher model that operates on explicit sequence data of any desired genome length. Each model component, including locus, bacterial strain, population and ultimately the whole metapopulation, is efficiently simulated using C++ objects and detailed metadata from each level can be acquired. The software can be executed in a cluster environment using simple textual input files, enabling, e.g. large-scale simulations and likelihood-free inference. Availability and implementation: Bacmeta is implemented with C++ for Linux, Mac and Windows. It is available at https://bitbucket.org/aleksisipola/bacmeta under the BSD 3-clause license. Supplementary information: Supplementary data are available at Bioinformatics online.
Aleksi Sipola, Pekka Marttinen, Jukka Corander
Bioinform.3
2018 ELFI: Engine for Likelihood-Free Inference
abstract
Engine for Likelihood-Free Inference (ELFI) is a Python software library for performing likelihood-free inference (LFI). ELFI provides a convenient syntax for arranging components in LFI, such as priors, simulators, summaries or distances, to a network called ELFI graph. The components can be implemented in a wide variety of languages. The stand-alone ELFI graph can be used with any of the available inference methods without modifications. A central method implemented in ELFI is Bayesian Optimization for Likelihood-Free Inference (BOLFI), which has recently been shown to accelerate likelihood-free inference up to several orders of magnitude by surrogate-modelling the distance. ELFI also has an inbuilt support for output data storing for reuse and analysis, and supports parallelization of computation from multiple cores up to a cluster environment. ELFI is designed to be extensible and provides interfaces for widening its functionality. This makes the adding of new inference methods to ELFI straightforward and automatically compatible with the inbuilt features.
Jarno Lintusaari, Henri Vuollekoski, Antti Kangasrääsiö, Kusti Skytén, Marko Järvenpää, Pekka Marttinen, Michael U. Gutmann, Aki Vehtari, Jukka Corander, Samuel Kaski
J. Mach. Learn. Res.9
2017 Inferring Cognitive Models from Data using Approximate Bayesian Computation
abstract
An important problem for HCI researchers is to estimate the parameter values of a cognitive model from behavioral data. This is a difficult problem, because of the substantial complexity and variety in human behavioral strategies. We report an investigation into a new approach using approximate Bayesian computation (ABC) to condition model parameters to data and prior knowledge. As the case study we examine menu interaction, where we have click time data only to infer a cognitive model that implements a search behaviour with parameters such as fixation duration and recall probability. Our results demonstrate that ABC (i) improves estimates of model parameter values, (ii) enables meaningful comparisons between model variants, and (iii) supports fitting models to individual users. ABC provides ample opportunities for theoretical HCI research by allowing principled inference of model parameter values and their uncertainty.
Antti Kangasrääsiö, Kumaripaba Athukorala, Andrew Howes 0001, Jukka Corander, Samuel Kaski, Antti Oulasvirta
CHI4
2017 Learning Gaussian graphical models with fractional marginal pseudo-likelihood
Janne Leppä-aho, Johan Pensar, Teemu Roos, Jukka Corander
Int. J. Approx. Reason.4
2017 From Random Matrix Theory to Coding Theory: Volume of a Metric Ball in Unitary Group
abstract
Volume estimates of metric balls in manifolds find diverse applications in information and coding theory. In this paper, new results for the volume of a metric ball in unitary group are derived via tools from random matrix theory. The first result is an integral representation of the exact volume, which involves a Toeplitz determinant of Bessel functions. A simple but accurate limiting volume formula is then obtained by invoking Szegö's strong limit theorem for large Toeplitz matrices. The derived asymptotic volume formula enables analytical evaluation of some coding-theoretic bounds of unitary codes. In particular, the Gilbert-Varshamov lower bound and the Hamming upper bound on the cardinality as well as the resulting bounds on code rate and minimum distance are derived. Moreover, bounds on the scaling law of code rate are found. Finally, a closed-form bound on the diversity sum relevant to unitary space-time codes is obtained, which was only computed numerically in the literature.
Lu Wei 0001, Renaud-Alexandre Pitaval, Jukka Corander, Olav Tirkkonen
IEEE Trans. Inf. Theory3
2017 Asymptotic Analysis of Rayleigh Product Channels: A Free Probability Approach
abstract
The Rayleigh product channel model is useful in capturing the performance degradation due to rank deficiency of MIMO channels. In this paper, such a performance degradation is investigated via the distribution of mutual information assuming the block fading channels and the uniform power transmission scheme. Using techniques of free probability theory, the asymptotic variance of mutual information is derived when the dimensions of the channel matrices approach infinity. In this asymptotic regime, the mutual information is rigorously proven to be Gaussian distributed. Using the obtained results, a fundamental tradeoff between multiplexing gain and diversity gain of Rayleigh product channels under the uniform power transmission can be characterized by the closed-form expression at any finite signal-to-noise ratio. Numerical results are provided to compare the outage performance between the Rayleigh product channels and the conventional Rayleigh MIMO channels.
Zhong Zheng 0001, Lu Wei 0001, Roland Speicher, Ralf R. Müller, Jyri Hämäläinen, Jukka Corander
IEEE Trans. Inf. Theory6
2016 Fast nearest neighbor search through sparse random projections and voting
abstract
Efficient index structures for fast approximate nearest neighbor queries are required in many applications such as recommendation systems. In high-dimensional spaces, many conventional methods suffer from excessive usage of memory and slow response times. We propose a method where multiple random projection trees are combined by a novel voting scheme. The key idea is to exploit the redundancy in a large number of candidate sets obtained by independently generated random projections in order to reduce the number of expensive exact distance evaluations. The method is straightforward to implement using sparse projections which leads to a reduced memory footprint and fast index construction. Furthermore, it enables grouping of the required computations into big matrix multiplications, which leads to additional savings due to cache effects and low-level parallelization. We demonstrate by extensive experiments on a wide variety of data sets that the method is faster than existing partitioning tree or hashing based approaches, making it the fastest available technique on high accuracy levels.
Ville Hyvönen, Teemu Pitkänen, Sotiris K. Tasoulis, Elias Jääsaari, Risto Tuomainen, Liang Wang 0009, Jukka Corander, Teemu Roos
IEEE BigData7
2016 A Logical Approach to Context-Specific Independence
Jukka Corander, Antti Hyttinen, Juha Kontinen, Johan Pensar, Jouko A. Väänänen
WoLLIC1
2016 On the inconsistency of ℓ 1-penalised sparse precision matrix estimation
abstract
Various ℓ 1-penalised estimation methods such as graphical lasso and CLIME are widely used for sparse precision matrix estimation and learning of undirected network structure from data. Many of these methods have been shown to be consistent under various quantitative assumptions about the underlying true covariance matrix. Intuitively, these conditions are related to situations where the penalty term will dominate the optimisation. We explore the consistency of ℓ 1-based methods for a class of bipartite graphs motivated by the structure of models commonly used for gene regulatory networks. We show that all ℓ 1-based methods fail dramatically for models with nearly linear dependencies between the variables. We also study the consistency on models derived from real gene expression data and note that the assumptions needed for consistency never hold even for modest sized gene networks and ℓ 1-based methods also become unreliable in practice for larger networks. Our results demonstrate that ℓ 1-penalised undirected network structure learning methods are unable to reliably learn many sparse bipartite graph structures, which arise often in gene expression data. Users of such methods should be aware of the consistency criteria of the methods and check if they are likely to be met in their application of interest.
Otte Heinävaara, Janne Leppä-aho, Jukka Corander, Antti Honkela
BMC Bioinform.3
2016 The role of local partial independence in learning of Bayesian networks
Johan Pensar, Henrik J. Nyman, Jarno Lintusaari, Jukka Corander
Int. J. Approx. Reason.4
2016 Bayesian Optimization for Likelihood-Free Inference of Simulator-Based Statistical Models
abstract
Our paper deals with inferring simulator-based statistical models given some observed data. A simulator-based model is a parametrized mechanism which specifies how data are generated. It is thus also referred to as generative model. We assume that only a finite number of parameters are of interest and allow the generative process to be very general; it may be a noisy nonlinear dynamical system with an unrestricted number of hidden variables. This weak assumption is useful for devising realistic models but it renders statistical inference very difficult. The main challenge is the intractability of the likelihood function. Several likelihood-free inference methods have been proposed which share the basic idea of identifying the parameters by finding values for which the discrepancy between simulated and observed data is small. A major obstacle to using these methods is their computational cost. The cost is largely due to the need to repeatedly simulate data sets and the lack of knowledge about how the parameters affect the discrepancy. We propose a strategy which combines probabilistic modeling of the discrepancy with optimization to facilitate likelihood-free inference. The strategy is implemented using Bayesian optimization and is shown to accelerate the inference through a reduction in the number of required simulations by several orders of magnitude.
Michael U. Gutmann, Jukka Corander
J. Mach. Learn. Res.2
2016 Low-Rank Doubly Stochastic Matrix Decomposition for Cluster Analysis
abstract
Cluster analysis by nonnegative low-rank approximations has experienced a remarkable progress in the past decade. However, the majority of such approximation approaches are still restricted to nonnegative matrix factorization (NMF) and suffer from the following two drawbacks: 1) they are unable to produce balanced partitions for large-scale manifold data which are common in real-world clustering tasks; 2) most existing NMF-type clustering methods cannot automatically determine the number of clusters. We propose a new low-rank learning method to address these two problems, which is beyond matrix factorization. Our method approximately decomposes a sparse input similarity in a normalized way and its objective can be used to learn both cluster assignments and the number of clusters. For efficient optimization, we use a relaxed formulation based on Data- Cluster-Data random walk, which is also shown to be equivalent to low-rank factorization of the doubly-stochastically normalized cluster incidence matrix. The probabilistic cluster assignments can thus be learned with a multiplicative majorization-minimization algorithm. Experimental results show that the new method is more accurate both in terms of clustering large-scale manifold data sets and of selecting the number of clusters.
Zhirong Yang, Jukka Corander, Erkki Oja
J. Mach. Learn. Res.2
2016 Volume of Metric Balls in High-Dimensional Complex Grassmann Manifolds
abstract
Volume of metric balls relates to rate-distortion theory and packing bounds on codes. In this paper, the volume of balls in complex Grassmann manifolds is evaluated for an arbitrary radius. The ball is defined as a set of hyperplanes of a fixed dimension with reference to a center of possibly different dimensions, and a generalized chordal distance for unequal dimensional subspaces is used. First, the volume is reduced to a 1-D integral representation. The overall problem boils down to evaluating a determinant of a matrix of the same size as the subspace dimensionality. Interpreting this determinant as a characteristic function of the Jacobi ensemble, an asymptotic analysis is carried out. The obtained asymptotic volume is moreover refined using moment-matching techniques to provide a tighter approximation in finite-size regimes. Finally, the pertinence of the derived results is shown by rate-distortion analysis of source coding on Grassmann manifolds.
Renaud-Alexandre Pitaval, Lu Wei 0001, Olav Tirkkonen, Jukka Corander
IEEE Trans. Inf. Theory4
2015 A gradient adaptive population importance sampler
abstract
Monte Carlo (MC) methods are widely used in signal processing and machine learning. A well-known class of MC methods is composed of importance sampling and its adaptive extensions (e.g., population Monte Carlo). In this paper, we introduce an adaptive importance sampler using a population of proposal densities. The novel algorithm dynamically optimizes the cloud of proposals, adapting them using information about the gradient and Hessian matrix of the target distribution. Moreover, a new kind of interaction in the adaptation of the proposal densities is introduced, establishing a trade-off between attaining a good performance in terms of mean square error and robustness to initialization.
Victor Elvira, Luca Martino, David Luengo, Jukka Corander
ICASSP4
2015 Smelly parallel MCMC chains
abstract
Monte Carlo (MC) methods are useful tools for Bayesian inference and stochastic optimization that have been widely applied in signal processing and machine learning. A well-known class of MC methods are Markov Chain Monte Carlo (MCMC) algorithms. In this work, we introduce a novel parallel interacting MCMC scheme, where the parallel chains share information, thus yielding a faster exploration of the state space. The interaction is carried out generating a dynamic repulsion among the “smelly” parallel chains that takes into account the entire population of current states. The ergodicity of the scheme and its relationship with other sampling methods are discussed. Numerical results show the advantages of the proposed approach in terms of mean square error, robustness w.r.t. to initial values and parameter choice.
Luca Martino, Victor Elvira, David Luengo, Antonio Artés-Rodríguez, Jukka Corander
ICASSP5
2015 Denoising Cluster Analysis
Ruqi Zhang, Zhirong Yang, Jukka Corander
ICONIP (3)3
2015 On the volume of a metric ball in unitary group
abstract
Volume estimates of metric balls in manifolds find diverse applications in communications and information theory. In this paper, we derive some new results for the volume of a metric ball in unitary group under Frobenius norm topological metric. Our first result is an integral representation of the exact volume, which involves a Toeplitz determinant of Bessel functions. The connection to matrix-variate hypergeometric functions leads from the exact finite size formula to an asymptotic one. The convergence of the obtained limiting formula is exceptionally fast due to the underlying mock-Gaussian behavior.
Lu Wei 0001, Renaud-Alexandre Pitaval, Jukka Corander, Olav Tirkkonen
ISIT3
2015 On the finite-SNR Diversity-Multiplexing Tradeoff in large Rayleigh product channels
abstract
The Diversity-Multiplexing Tradeoff (DMT) is studied for the large Rayleigh product channel at non-asymptotic SNRs. The first result is that, as matrix dimensions growing to infinity, the channel capacity converges to a Gaussian random variable. Based on this, we derive a compact expression for the finite-SNR DMT. From the analytical and numerical results, we gain useful insight into the fundamental tradeoff of the considered channel model in the realistic SNR regime.
Zhong Zheng 0001, Lu Wei 0001, Roland Speicher, Ralf R. Müller, Jyri Hämäläinen, Jukka Corander
ISIT6
2015 Labeled directed acyclic graphs: a generalization of context-specific independence in directed graphical models
Johan Pensar, Henrik J. Nyman, Timo Koski, Jukka Corander
Data Min. Knowl. Discov.4
2015 A Bayesian Predictive Model for Clustering Data of Mixed Discrete and Continuous Type
abstract
Advantages of model-based clustering methods over heuristic alternatives have been widely demonstrated in the literature. Most model-based clustering algorithms assume that the data are either discrete or continuous, possibly allowing both types to be present in separate features. In this paper, we introduce a model-based approach for clustering feature vectors of mixed type, allowing each feature to simultaneously take on both categorical and real values. Such data may be encountered, for instance, in chemical and biological analyses, in the analysis of survey data, as well as in image analysis. Our model is formulated within a Bayesian predictive framework, where clustering solutions correspond to random partitions of the data. Using conjugate analysis, the posterior probability for each possible partition can be determined analytically, enabling the utilization of efficient computational search strategies for finding the posterior optimal partition. The derived model is illustrated using several synthetic and real datasets.
Paul Blomstedt, Jing Tang 0002, Christian Granlund, Jukka Corander
IEEE Trans. Pattern Anal. Mach. Intell.5
2015 On the Outage Capacity of Orthogonal Space-Time Block Codes Over Multi-Cluster Scattering MIMO Channels
abstract
The multiple cluster scattering MIMO channel is a useful model for pico-cellular MIMO networks. In this paper, space-time coded transmission over such a channel is considered, where the effective channel corresponds to a product of complex Gaussian matrices. An accurate closed-form approximation to the channel outage capacity using orthogonal space-time block codes has been derived. The result is valid for an arbitrary number of clusters of scatterers and an arbitrary antenna configuration. From the analytical and numerical results, we study the relative outage performance between the multi-cluster MIMO channel and the special case of Rayleigh-fading MIMO channel.
Lu Wei 0001, Zhong Zheng 0001, Jukka Corander, Giorgio Taricco
IEEE Trans. Commun.3
2014 Preface
Samuel Kaski, Jukka Corander
AISTATS2
2014 Random projection based clustering for population genomics
abstract
Recent data revolution in population genomics for bacteria has increased the size of aligned sequence data sets by two-to-three orders of magnitude. This trend is expected to continue in the near future, putting an emphasis on applicability of big data techniques to leverage biologically important insights. Moreover, with the increasing density of sampling, it may also be necessary to consider alignment-free sequence analysis techniques combined with clustering to yield a sufficient insight to data. This leads to ultra high-dimensional data with tens of millions of variables, which can no longer be handled by the existing population genomic methods. Using the largest bacterial sequence data sets published to date, we demonstrate that random projection based clustering provides a highly accurate and several orders of magnitude faster approach to the analysis of both alignment-based and alignment-free genome data sets, compared with the Bayesian model-based analysis that is currently considered as the state-of-the-art. Hence, clustering methods for big data harbor considerable potential for important applications in genomics and could pave way for novel analysis pipelines even in the online setting when executed in a massively parallel computing environment.
Sotiris K. Tasoulis, Lu Cheng 0004, Niko Välimäki, Nicholas J. Croucher, Simon R. Harris, William P. Hanage, Teemu Roos, Jukka Corander
IEEE BigData8
2014 An adaptive population importance sampler
abstract
Monte Carlo (MC) methods are widely used in signal processing, machine learning and communications for statistical inference and stochastic optimization. A well-known class of MC methods is composed of importance sampling and its adaptive extensions (e.g., population Monte Carlo). In this work, we introduce an adaptive importance sampler using a population of proposal densities. The novel algorithm provides a global estimation of the variables of interest iteratively, using all the samples generated. The cloud of proposals is adapted by learning from a subset of previously generated samples, in such a way that local features of the target density can be better taken into account compared to single global adaptation procedures. Numerical results show the advantages of the proposed sampling scheme in terms of mean absolute error and robustness to initialization.
Luca Martino, Victor Elvira, David Luengo, Jukka Corander
ICASSP4
2014 Outage capacity of OSTBCs over pico-cellular MIMO channels
abstract
We consider orthogonal space-time block coded transmission over the multiple cluster scattering MIMO channels, where the effective channel equals the product of n complex Gaussian matrices. The considered channel model is typical in modeling pico-cellular MIMO propagations. In this setting, we derived a closed-form approximation to the channel outage capacity. The result is valid for an arbitrary number of clusters n-1 of scatterers and an arbitrary antenna configuration. Numerical results show the usefulness of the proposed approximation in diverse scenarios.
Lu Wei 0001, Zhong Zheng 0001, Jukka Corander, Giorgio Taricco
ISIT3
2014 SEK: sparsity exploiting k-mer-based estimation of bacterial community composition
abstract
MOTIVATION: Estimation of bacterial community composition from a high-throughput sequenced sample is an important task in metagenomics applications. As the sample sequence data typically harbors reads of variable lengths and different levels of biological and technical noise, accurate statistical analysis of such data is challenging. Currently popular estimation methods are typically time-consuming in a desktop computing environment. RESULTS: Using sparsity enforcing methods from the general sparse signal processing field (such as compressed sensing), we derive a solution to the community composition estimation problem by a simultaneous assignment of all sample reads to a pre-processed reference database. A general statistical model based on kernel density estimation techniques is introduced for the assignment task, and the model solution is obtained using convex optimization tools. Further, we design a greedy algorithm solution for a fast solution. Our approach offers a reasonably fast community composition estimation method, which is shown to be more robust to input data variation than a recently introduced related method. AVAILABILITY AND IMPLEMENTATION: A platform-independent Matlab implementation of the method is freely available at http://www.ee.kth.se/ctsoftware; source code that does not require access to Matlab is currently being tested and will be made available later through the above Web site.
Saikat Chatterjee, David Koslicki, Siyuan Dong, Nicolas Innocenti, Lu Cheng 0004, Yueheng Lan, Mikko Vehkaperä, Mikael Skoglund, Lars K. Rasmussen, Erik Aurell, Jukka Corander
Bioinform.11
2013 Learning Chordal Markov Networks by Constraint Satisfaction
abstract
We investigate the problem of learning the structure of a Markov network from data. It is shown that the structure of such networks can be described in terms of constraints which enables the use of existing solver technology with optimization capabilities to compute optimal networks starting from initial scores computed from the data. To achieve efficient encodings, we develop a novel characterization of Markov network structure using a balancing condition on the separators between cliques forming the network. The resulting translations into propositional satisfiability and its extensions such as maximum satisfiability, satisfiability modulo theories, and answer set programming, enable us to prove the optimality of networks which have been previously found by stochastic search.
Jukka Corander, Tomi Janhunen, Jussi Rintanen, Henrik J. Nyman, Johan Pensar
NIPS1
2013 Approximate Bayesian Computation
abstract
Approximate Bayesian computation (ABC) constitutes a class of computational methods rooted in Bayesian statistics. In all model-based statistical inference, the likelihood function is of central importance, since it expresses the probability of the observed data under a particular statistical model, and thus quantifies the support data lend to particular values of parameters and to choices among different models. For simple models, an analytical formula for the likelihood function can typically be derived. However, for more complex models, an analytical formula might be elusive or the likelihood function might be computationally very costly to evaluate. ABC methods bypass the evaluation of the likelihood function. In this way, ABC methods widen the realm of models for which statistical inference can be considered. ABC methods are mathematically well-founded, but they inevitably make assumptions and approximations whose impact needs to be carefully assessed. Furthermore, the wider application domain of ABC exacerbates the challenges of parameter estimation and model selection. ABC has rapidly gained popularity over the last years and in particular for the analysis of complex problems arising in biological sciences (e.g., in population genetics, ecology, epidemiology, and systems biology).
Mikael Sunnåker, Alberto Giovanni Busetto, Elina Numminen, Jukka Corander, Matthieu Foll, Christophe Dessimoz
PLoS Comput. Biol.4
2011 Bayesian semi-supervised classification of bacterial samples using MLST databases
abstract
BACKGROUND: Worldwide effort on sampling and characterization of molecular variation within a large number of human and animal pathogens has lead to the emergence of multi-locus sequence typing (MLST) databases as an important tool for studying the epidemiology and evolution of pathogens. Many of these databases are currently harboring several thousands of multi-locus DNA sequence types (STs) enriched with metadata over traits such as serotype, antibiotic resistance, host organism etc of the isolates. Curators of the databases have thus the possibility of dividing the pathogen populations into subsets representing different evolutionary lineages, geographically associated groups, or other subpopulations, which are defined in terms of molecular similarities and dissimilarities residing within a database. When combined with the existing metadata, such subsets may provide invaluable information for assessing the position of a new set of isolates in relation to the whole pathogen population. RESULTS: To enable users of MLST schemes to query the databases with sets of new bacterial isolates and to automatically analyze their relation to existing curated sequences, we introduce here a Bayesian model-based method for semi-supervised classification of MLST data. Our method can use an MLST database as a training set and assign simultaneously any set of query sequences into the earlier discovered lineages/populations, while also allowing some or all of these sequences to form previously undiscovered genetically distinct groups. This tool provides probabilistic quantification of the classification uncertainty and is highly efficient computationally, thus enabling rapid analyses of large databases and sets of query sequences. The latter feature is a necessary prerequisite for an automated access through the MLST web interface. We demonstrate the versatility of our approach by anayzing both real and synthesized data from MLST databases. The introduced method for semi-supervised classification of sets of query STs is freely available for Windows, Mac OS X and Linux operative systems in BAPS 5.4 software which is downloadable at http://web.abo.fi/fak/mnf/mate/jc/software/baps.html. The query functionality is also directly available for the Staphylococcus aureus database at http://www.mlst.net and shortly will be available for other species databases hosted at this web portal. CONCLUSIONS: We have introduced a model-based tool for automated semi-supervised classification of new pathogen samples that can be integrated into the web interface of the MLST databases. In particular, when combined with the existing metadata, the semi-supervised labeling may provide invaluable information for assessing the position of a new set of query strains in relation to the particular pathogen population represented by the curated database.Such information will be useful both for clinical and basic research purposes.
Lu Cheng 0004, Thomas R. Connor, David M. Aanensen, Brian G. Spratt, Jukka Corander
BMC Bioinform.5
2010 Efficient Bayesian approach for multilocus association mapping including gene-gene interactions
abstract
BACKGROUND: since the introduction of large-scale genotyping methods that can be utilized in genome-wide association (GWA) studies for deciphering complex diseases, statistical genetics has been posed with a tremendous challenge of how to most appropriately analyze such data. A plethora of advanced model-based methods for genetic mapping of traits has been available for more than 10 years in animal and plant breeding. However, most such methods are computationally intractable in the context of genome-wide studies. Therefore, it is hardly surprising that GWA analyses have in practice been dominated by simple statistical tests concerned with a single marker locus at a time, while the more advanced approaches have appeared only relatively recently in the biomedical and statistical literature. RESULTS: we introduce a novel Bayesian modeling framework for association mapping which enables the detection of multiple loci and their interactions that influence a dichotomous phenotype of interest. The method is shown to perform well in a simulation study when compared to widely used standard alternatives and its computational complexity is typically considerably smaller than that of a maximum likelihood based approach. We also discuss in detail the sensitivity of the Bayesian inferences with respect to the choice of prior distributions in the GWA context. CONCLUSIONS: our results show that the Bayesian model averaging approach which explicitly considers gene-gene interactions may improve the detection of disease associated genetic markers in two respects: first, by providing better estimates of the locations of the causal loci; second, by reducing the number of false positives. The benefits are most apparent when the interacting genes exhibit no main effects. However, our findings also illustrate that such an approach is somewhat sensitive to the prior distribution assigned on the model structure.
Pekka Marttinen, Jukka Corander
BMC Bioinform.2
2009 Bayesian clustering and feature selection for cancer tissue samples
abstract
BACKGROUND: The versatility of DNA copy number amplifications for profiling and categorization of various tissue samples has been widely acknowledged in the biomedical literature. For instance, this type of measurement techniques provides possibilities for exploring sets of cancerous tissues to identify novel subtypes. The previously utilized statistical approaches to various kinds of analyses include traditional algorithmic techniques for clustering and dimension reduction, such as independent and principal component analyses, hierarchical clustering, as well as model-based clustering using maximum likelihood estimation for latent class models. RESULTS: While purely algorithmic methods are usually easily applicable, their suboptimal performance and limitations in making formal inference have been thoroughly discussed in the statistical literature. Here we introduce a Bayesian model-based approach to simultaneous identification of underlying tissue groups and the informative amplifications. The model-based approach provides the possibility of using formal inference to determine the number of groups from the data, in contrast to the ad hoc methods often exploited for similar purposes. The model also automatically recognizes the chromosomal areas that are relevant for the clustering. CONCLUSION: Validatory analyses of simulated data and a large database of DNA copy number amplifications in human neoplasms are used to illustrate the potential of our approach. Our software implementation BASTA for performing Bayesian statistical tissue profiling is freely available for academic purposes at (http://web.abo.fi/fak/mnf/mate/jc/software/basta.html).
Pekka Marttinen, Samuel Myllykangas, Jukka Corander
BMC Bioinform.3
2009 Bayesian learning of graphical vector autoregressions with unequal lag-lengths
Pekka Marttinen, Jukka Corander
Mach. Learn.2
2009 Bayesian Clustering of Fuzzy Feature Vectors Using a Quasi-Likelihood Approach
abstract
Bayesian model-based classifiers, both unsupervised and supervised, have been studied extensively and their value and versatility have been demonstrated on a wide spectrum of applications within science and engineering. A majority of the classifiers are built on the assumption of intrinsic discreteness of the considered data features or on the discretization of them prior to the modeling. On the other hand, Gaussian mixture classifiers have also been utilized to a large extent for continuous features in the Bayesian framework. Often the primary reason for discretization in the classification context is the simplification of the analytical and numerical properties of the models. However, the discretization can be problematic due to its \textit{ad hoc} nature and the decreased statistical power to detect the correct classes in the resulting procedure. We introduce an unsupervised classification approach for fuzzy feature vectors that utilizes a discrete model structure while preserving the continuous characteristics of data. This is achieved by replacing the ordinary likelihood by a binomial quasi-likelihood to yield an analytical expression for the posterior probability of a given clustering solution. The resulting model can be justified from an information-theoretic perspective. Our method is shown to yield highly accurate clusterings for challenging synthetic and empirical data sets.
Pekka Marttinen, Jing Tang 0002, Bernard De Baets, Peter Dawyndt, Jukka Corander
IEEE Trans. Pattern Anal. Mach. Intell.5
2009 Identifying Currents in the Gene Pool for Bacterial Populations Using an Integrative Approach
abstract
The evolution of bacterial populations has recently become considerably better understood due to large-scale sequencing of population samples. It has become clear that DNA sequences from a multitude of genes, as well as a broad sample coverage of a target population, are needed to obtain a relatively unbiased view of its genetic structure and the patterns of ancestry connected to the strains. However, the traditional statistical methods for evolutionary inference, such as phylogenetic analysis, are associated with several difficulties under such an extensive sampling scenario, in particular when a considerable amount of recombination is anticipated to have taken place. To meet the needs of large-scale analyses of population structure for bacteria, we introduce here several statistical tools for the detection and representation of recombination between populations. Also, we introduce a model-based description of the shape of a population in sequence space, in terms of its molecular variability and affinity towards other populations. Extensive real data from the genus Neisseria are utilized to demonstrate the potential of an approach where these population genetic tools are combined with an phylogenetic analysis. The statistical tools introduced here are freely available in BAPS 5.2 software, which can be downloaded from http://web.abo.fi/fak/mnf/mate/jc/software/baps.html.
Jing Tang 0002, William P. Hanage, Christophe Fraser, Jukka Corander
PLoS Comput. Biol.4
2008 Enhanced Bayesian modelling in BAPS software for learning genetic structures of populations
abstract
BACKGROUND: During the most recent decade many Bayesian statistical models and software for answering questions related to the genetic structure underlying population samples have appeared in the scientific literature. Most of these methods utilize molecular markers for the inferences, while some are also capable of handling DNA sequence data. In a number of earlier works, we have introduced an array of statistical methods for population genetic inference that are implemented in the software BAPS. However, the complexity of biological problems related to genetic structure analysis keeps increasing such that in many cases the current methods may provide either inappropriate or insufficient solutions. RESULTS: We discuss the necessity of enhancing the statistical approaches to face the challenges posed by the ever-increasing amounts of molecular data generated by scientists over a wide range of research areas and introduce an array of new statistical tools implemented in the most recent version of BAPS. With these methods it is possible, e.g., to fit genetic mixture models using user-specified numbers of clusters and to estimate levels of admixture under a genetic linkage model. Also, alleles representing a different ancestry compared to the average observed genomic positions can be tracked for the sampled individuals, and a priori specified hypotheses about genetic population structure can be directly compared using Bayes' theorem. In general, we have improved further the computational characteristics of the algorithms behind the methods implemented in BAPS facilitating the analyses of large and complex datasets. In particular, analysis of a single dataset can now be spread over multiple computers using a script interface to the software. CONCLUSION: The Bayesian modelling methods introduced in this article represent an array of enhanced tools for learning the genetic structure of populations. Their implementations in the BAPS software are designed to meet the increasing need for analyzing large-scale population genetics data. The software is freely downloadable for Windows, Linux and Mac OS X systems at http://web.abo.fi/fak/mnf//mate/jc/software/baps.html.
Jukka Corander, Pekka Marttinen, Jukka Sirén, Jing Tang 0002
BMC Bioinform.1
2008 Bayesian modeling of recombination events in bacterial populations
abstract
BACKGROUND: We consider the discovery of recombinant segments jointly with their origins within multilocus DNA sequences from bacteria representing heterogeneous populations of fairly closely related species. The currently available methods for recombination detection capable of probabilistic characterization of uncertainty have a limited applicability in practice as the number of strains in a data set increases. RESULTS: We introduce a Bayesian spatial structural model representing the continuum of origins over sites within the observed sequences, including a probabilistic characterization of uncertainty related to the origin of any particular site. To enable a statistically accurate and practically feasible approach to the analysis of large-scale data sets representing a single genus, we have developed a novel software tool (BRAT, Bayesian Recombination Tracker) implementing the model and the corresponding learning algorithm, which is capable of identifying the posterior optimal structure and to estimate the marginal posterior probabilities of putative origins over the sites. CONCLUSION: A multitude of challenging simulation scenarios and an analysis of real data from seven housekeeping genes of 120 strains of genus Burkholderia are used to illustrate the possibilities offered by our approach. The software is freely available for download at URL http://web.abo.fi/fak/mnf//mate/jc/software/brat.html.
Pekka Marttinen, Adam Baldwin, William P. Hanage, Chris Dowson, Eshwar Mahenthiralingam, Jukka Corander
BMC Bioinform.6
2008 Parallell interacting MCMC for learning of topologies of graphical models
Jukka Corander, Magnus Ekdahl, Timo Koski
Data Min. Knowl. Discov.1
2006 Bayesian search of functionally divergent protein subgroups and their function specific residues
abstract
MOTIVATION: The rapid increase in the amount of protein sequence data has created a need for an automated identification of evolutionarily related subgroups from large datasets. The existing methods typically require a priori specification of the number of putative groups, which defines the resolution of the classification solution. RESULTS: We introduce a Bayesian model-based approach to simultaneous identification of evolutionary groups and conserved parts of the protein sequences. The model-based approach provides an intuitive and efficient way of determining the number of groups from the sequence data, in contrast to the ad hoc methods often exploited for similar purposes. Our model recognizes the areas in the sequences that are relevant for the clustering and regards other areas as noise. We have implemented the method using a fast stochastic optimization algorithm which yields a clustering associated with the estimated maximum posterior probability. The method has been shown to have high specificity and sensitivity in simulated and real clustering tasks. With real datasets the method also highlights the residues close to the active site. AVAILABILITY: Software 'kPax' is available at http://www.rni.helsinki.fi/jic/softa.html
Pekka Marttinen, Jukka Corander, Petri Törönen, Liisa Holm
Bioinform.2
2004 BAPS 2: enhanced possibilities for the analysis of genetic population structure
abstract
UNLABELLED: Bayesian statistical methods based on simulation techniques have recently been shown to provide powerful tools for the analysis of genetic population structure. We have previously developed a Markov chain Monte Carlo (MCMC) algorithm for characterizing genetically divergent groups based on molecular markers and geographical sampling design of the dataset. However, for large-scale datasets such algorithms may get stuck to local maxima in the parameter space. Therefore, we have modified our earlier algorithm to support multiple parallel MCMC chains, with enhanced features that enable considerably faster and more reliable estimation compared to the earlier version of the algorithm. We consider also a hierarchical tree representation, from which a Bayesian model-averaged structure estimate can be extracted. The algorithm is implemented in a computer program that features a user-friendly interface and built-in graphics. The enhanced features are illustrated by analyses of simulated data and an extensive human molecular dataset. AVAILABILITY: Freely available at http://www.rni.helsinki.fi/~jic/bapspage.html.
Jukka Corander, Patrik Waldmann, Pekka Marttinen, Mikko J. Sillanpää
Bioinform.1