EDBT 2026 Demo / reviewers in the wild / expert
Pekka Marttinen
dblp:32/894
· DBLP profile ↗
49ranked-venue papers
8as first author
27since 2021 · last 2026
0000-0001-7078-7927ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 2 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 6 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond the LUMIR challenge: The pathway to foundational registration models
Junyu Chen 0002, Shuwen Wei, Joel Honkamaa, Pekka Marttinen, Hang Zhang 0010, Min Liu 0008, Yichao Zhou 0002, Zuopeng Tan, Yi Wang 0028, Hongchao Zhou, Shunbo Hu, Yi Zhang 0120, Lukas Förner, Thomas Wendler 0001, Bailiang Jian, Benedikt Wiestler, Tim Hable, Dan Ruan, Frederic Madesta, Thilo Sentker, Wiebke Heyer, Lianrui Zuo, Yuwei Dai, Jerry L. Prince, Harrison X. Bai, Yong Du 0002, Yihao Liu 0003, Alessa Hering, Reuben Dorent, Lasse Hansen, Mattias P. Heinrich, Aaron Carass |
Medical Image Anal. | 4 |
| 2025 | Identifying latent state transitions in non-linear dynamical systemsabstractThis work aims to recover the underlying states and their time evolution in a latent dynamical system from high-dimensional sensory measurements. Previous works on identifiable representation learning in dynamical systems focused on identifying the latent states, often with linear transition approximations. As such, they cannot identify nonlinear transition dynamics, and hence fail to reliably predict complex future behavior. Inspired by the advances in nonlinear ICA, we propose a state-space modeling framework in which we can identify not just the latent states but also the unknown transition function that maps the past states to the present. Our identifiability theory relies on two key assumptions: (i) sufficient variability in the latent noise, and (ii) the bijectivity of the augmented transition function. Drawing from this theory, we introduce a practical algorithm based on variational auto-encoders. We empirically demonstrate that it improves generalization and interpretability of target dynamical systems by (i) recovering latent state dynamics with high accuracy, (ii) correspondingly achieving high future prediction accuracy, and (iii) adapting fast to new environments. Additionally, for complex real-world dynamics, (iv) it produces state-of the-art future prediction results for long horizons, highlighting its usefulness for practical scenarios. Caglar Hizli, Çagatay Yildiz, Matthias Bethge, S. T. John, Pekka Marttinen |
ICLR | 5 |
| 2025 | New Multimodal Similarity Measure for Image Registration via Modeling Local Functional Dependence with Linear Combination of Learned Basis Functions
Joel Honkamaa, Pekka Marttinen |
MICCAI (2) | 2 |
| 2024 | Knowledge-augmented Graph Neural Networks with Concept-aware Attention for Adverse Drug Event DetectionabstractAdverse drug events (ADEs) are an important aspect of drug safety. Various texts such as biomedical literature, drug reviews, and user posts on social media and medical forums contain a wealth of information about ADEs. Recent studies have applied word embedding and deep learning-based natural language processing to automate ADE detection from text. However, they did not explore incorporating explicit medical knowledge about drugs and adverse reactions or the corresponding feature learning. This paper adopts the heterogeneous text graph, which describes relationships between documents, words, and concepts, augments it with medical knowledge from the Unified Medical Language System, and proposes a concept-aware attention mechanism that learns features differently for the different types of nodes in the graph. We further utilize contextualized embeddings from pretrained language models and convolutional graph neural networks for effective feature representation and relational learning. Experiments on four public datasets show that our model performs competitively to the recent advances, and the concept-aware attention consistently outperforms other attention mechanisms. Ya Gao 0005, Shaoxiong Ji, Pekka Marttinen |
LREC/COLING | 3 |
| 2024 | Improving Medical Multi-modal Contrastive Learning with Expert Annotations
Pekka Marttinen |
ECCV (20) | 2 |
| 2024 | Generating Demonstrations for In-Context Compositional Generalization in Grounded Language LearningabstractIn-Context-learning and few-shot prompting are viable methods compositional output generation.However, these methods can be very sensitive to the choice of support examples used.Retrieving good supports from the training data for a given test query is already a difficult problem, but in some cases solving this may not even be enough.We consider the setting of grounded language learning problems where finding relevant supports in the same or similar states as the query may be difficult.We design an agent which instead generates possible supports inputs and targets current state of the world, then uses them in-context-learning to solve the test query.We show substantially improved performance on a previously unsolved compositional generalization test without a loss of performance in other areas.The approach is general and can even scale to instructions expressed in natural language.0.61 h 0.57 ± .50 (60) 0.61 h 0.86 ± .34(266) 0.64 h 0.59 ± .49(743) 0.65 h 0.87 ± .33 (2655) 0.68 h 0.59 ± .49(2907) 0.69 h 0.88 ± .32 (8144) 0.71 h 0.63 ± .48(4412) 0.73 h 0.86 ± .35(7480) 0.74 h 0.80 ± .40 (6025) 0.77 h 0.78 ± .41(7283) 0.78 h 0.81 ± .39 (10824) 0.81 h 0.76 ± .43 (5358) Sam Spilsbury, Pekka Marttinen, Alexander Ilin |
EMNLP | 2 |
| 2024 | VMS: Interactive Visualization to Support the Sensemaking and Selection of Predictive ModelsabstractTo compare and select machine learning models, relying on performance measures alone may not always be sufficient. This is particularly the case where different subsets, features, and predicted results may vary in importance relative to the task at hand. Explanation and visualization techniques are required to support model sensemaking and informed decision-making. However, a review shows that existing systems are mostly designed for model developers and not evaluated with target users in their effectiveness. To address this issue, this research proposes an interactive visualization, VMS (Visualization for Model Sensemaking and Selection), for users of the model to compare and select predictive models. VMS integrates performance-, instance-, and feature-level analysis to evaluate models from multiple angles. Particularly, a feature view integrating the value and contribution of hundreds of features supports model comparison on local and global scales. We exemplified VMS for comparing models predicting patients’ hospital length of stay through time-series health records and evaluated the prototype with 16 participants from the medical field. Results reveal evidence that VMS supports users to rationalize models in multiple ways and enables users to select the optimal models with a small sample size. User feedback suggests future directions on incorporating domain knowledge in model training, such as for different patient groups considering different sets of features as important. Chen He 0003, Vishnu Raj, Hans Moen, Tommi Gröhn, Chen Wang 0145, Laura-Maria Peltonen, Saila Koivusalo, Pekka Marttinen, Giulio Jacucci |
IUI | 8 |
| 2024 | Generating Code World Models with Large Language Models Guided by Monte Carlo Tree SearchabstractIn this work we consider Code World Models, world models generated by a Large Language Model (LLM) in the form of Python code for model-based Reinforcement Learning (RL). Calling code instead of LLMs for planning has potential to be more precise, reliable, interpretable, and extremely efficient.
However, writing appropriate Code World Models requires the ability to understand complex instructions, to generate exact code with non-trivial logic and to self-debug a long program with feedback from unit tests and environment trajectories. To address these challenges, we propose Generate, Improve and Fix with Monte Carlo Tree Search (GIF-MCTS), a new code generation strategy for LLMs. To test our approach in an offline RL setting, we introduce the Code World Models Benchmark (CWMB), a suite of program synthesis and planning tasks comprised of 18 diverse RL environments paired with corresponding textual descriptions and curated trajectories. GIF-MCTS surpasses all baselines on the CWMB and two other benchmarks, and we show that the Code World Models synthesized with it can be successfully used for planning, resulting in model-based RL agents with greatly improved sample efficiency and inference speed. Nicola Dainese, Matteo Merler, Minttu Alakuijala, Pekka Marttinen |
NeurIPS | 4 |
| 2024 | Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic SimilaritiesabstractUncertainty quantification in Large Language Models (LLMs) is crucial for applications where safety and reliability are important. In particular, uncertainty can be used to improve the trustworthiness of LLMs by detecting factually incorrect model responses, commonly called hallucinations. Critically, one should seek to capture the model's semantic uncertainty, i.e., the uncertainty over the meanings of LLM outputs, rather than uncertainty over lexical or syntactic variations that do not affect answer correctness.
To address this problem, we propose Kernel Language Entropy (KLE), a novel method for uncertainty estimation in white- and black-box LLMs. KLE defines positive semidefinite unit trace kernels to encode the semantic similarities of LLM outputs and quantifies uncertainty using the von Neumann entropy. It considers pairwise semantic dependencies between answers (or semantic clusters), providing more fine-grained uncertainty estimates than previous methods based on hard clustering of answers. We theoretically prove that KLE generalizes the previous state-of-the-art method called semantic entropy and empirically demonstrate that it improves uncertainty quantification performance across multiple natural language generation datasets and LLM architectures. Alexander Nikitin 0002, Jannik Kossen, Yarin Gal, Pekka Marttinen |
NeurIPS | 4 |
| 2023 | Incorporating functional summary information in Bayesian neural networks using a Dirichlet process likelihood approachabstractBayesian neural networks (BNNs) can account for both aleatoric and epistemic uncertainty. However, in BNNs the priors are often specified over the weights which rarely reflects true prior knowledge in large and complex neural network architectures. We present a simple approach to incorporate prior knowledge in BNNs based on external summary information about the predicted classification probabilities for a given dataset. The available summary information is incorporated as augmented data and modeled with a Dirichlet process, and we derive the corresponding Summary Evidence Lower BOund. The approach is founded on Bayesian principles, and all hyperparameters have a proper probabilistic interpretation. We show how the method can inform the model about task difficulty and class imbalance. Extensive experiments show that, with negligible computational overhead, our method parallels and in many cases outperforms popular alternatives in accuracy, uncertainty calibration, and robustness against corruptions with both balanced and imbalanced data. Vishnu Raj, Tianyu Cui, Markus Heinonen, Pekka Marttinen |
AISTATS | 4 |
| 2023 | Patient Outcome and Zero-shot Diagnosis Prediction with Hypernetwork-guided Multitask LearningabstractMultitask deep learning has been applied to patient outcome prediction from text, taking clinical notes as input and training deep neural networks with a joint loss function of multiple tasks.However, the joint training scheme of multitask learning suffers from inter-task interference, and diagnosis prediction among the multiple tasks has the generalizability issue due to rare diseases or unseen diagnoses.To solve these challenges, we propose a hypernetwork-based approach that generates task-conditioned parameters and coefficients of multitask prediction heads to learn task-specific prediction and balance the multitask learning.We also incorporate semantic task information to improve the generalizability of our task-conditioned multitask model.Experiments on early and discharge notes extracted from the real-world MIMIC database show our method can achieve better performance on multitask patient outcome prediction than strong baselines in most cases.Besides, our method can effectively handle the scenario with limited information and improve zero-shot prediction on unseen diagnosis categories. Shaoxiong Ji, Pekka Marttinen |
EACL | 2 |
| 2023 | Reader: Model-based language-instructed reinforcement learningabstractWe explore how we can build accurate world models, which are partially specified by language, and how we can plan with them in the face of novelty and uncertainty. We propose the first model-based reinforcement learning approach to tackle the environment Read To Fight Monsters (Zhong et al., 2019), a grounded policy learning problem. In RTFM an agent has to reason over a set of rules and a goal, both described in a language manual, and the observations, while taking into account the uncertainty arising from the stochasticity of the environment, in order to generalize successfully its policy to test episodes. We demonstrate the superior performance and sample efficiency of our model-based approach to the existing model-free SOTA agents in eight variants of RTFM. Furthermore, we show how the agent’s plans can be inspected, which represents progress towards more interpretable agents. Nicola Dainese, Pekka Marttinen, Alexander Ilin |
EMNLP | 2 |
| 2023 | Causal Modeling of Policy Interventions From Treatment-Outcome SequencesabstractA treatment policy defines when and what treatments are applied to affect some outcome of interest. Data-driven decision-making requires the ability to predict what happens if a policy is changed. Existing methods that predict how the outcome evolves under different scenarios assume that the tentative sequences of future treatments are fixed in advance, while in practice the treatments are determined stochastically by a policy and may depend, for example, on the efficiency of previous treatments. Therefore, the current methods are not applicable if the treatment policy is unknown or a counterfactual analysis is needed. To handle these limitations, we model the treatments and outcomes jointly in continuous time, by combining Gaussian processes and point processes. Our model enables the estimation of a treatment policy from observational sequences of treatments and outcomes, and it can predict the interventional and counterfactual progression of the outcome after an intervention on the treatment policy (in contrast with the causal effect of a single treatment). We show with real-world and semi-synthetic data on blood glucose progression that our method can answer causal queries more accurately than existing alternatives. Caglar Hizli, S. T. John, Anne Juuti, Tuure Saarinen, Kirsi Pietiläinen, Pekka Marttinen |
ICML | 6 |
| 2023 | Temporal Causal Mediation through a Point Process: Direct and Indirect Effects of Healthcare InterventionsabstractDeciding on an appropriate intervention requires a causal model of a treatment, the outcome, and potential mediators. Causal mediation analysis lets us distinguish between direct and indirect effects of the intervention, but has mostly been studied in a static setting. In healthcare, data come in the form of complex, irregularly sampled time-series, with dynamic interdependencies between a treatment, outcomes, and mediators across time. Existing approaches to dynamic causal mediation analysis are limited to regular measurement intervals, simple parametric models, and disregard long-range mediator--outcome interactions. To address these limitations, we propose a non-parametric mediator--outcome model where the mediator is assumed to be a temporal point process that interacts with the outcome process. With this model, we estimate the direct and indirect effects of an external intervention on the outcome, showing how each of these affects the whole future trajectory. We demonstrate on semi-synthetic data that our method can accurately estimate direct and indirect effects. On real-world healthcare data, our model infers clinically meaningful direct and indirect effect trajectories for blood glucose after a surgery. Caglar Hizli, St John, Anne Juuti, Tuure Saarinen, Kirsi Pietiläinen, Pekka Marttinen |
NeurIPS | 6 |
| 2023 | HAPNEST: efficient, large-scale generation and evaluation of synthetic datasets for genotypes and phenotypesabstractMOTIVATION: Existing methods for simulating synthetic genotype and phenotype datasets have limited scalability, constraining their usability for large-scale analyses. Moreover, a systematic approach for evaluating synthetic data quality and a benchmark synthetic dataset for developing and evaluating methods for polygenic risk scores are lacking. RESULTS: We present HAPNEST, a novel approach for efficiently generating diverse individual-level genotypic and phenotypic data. In comparison to alternative methods, HAPNEST shows faster computational speed and a lower degree of relatedness with reference panels, while generating datasets that preserve key statistical properties of real data. These desirable synthetic data properties enabled us to generate 6.8 million common variants and nine phenotypes with varying degrees of heritability and polygenicity across 1 million individuals. We demonstrate how HAPNEST can facilitate biobank-scale analyses through the comparison of seven methods to generate polygenic risk scoring across multiple ancestry groups and different genetic architectures. AVAILABILITY AND IMPLEMENTATION: A synthetic dataset of 1 008 000 individuals and nine traits for 6.8 million common variants is available at https://www.ebi.ac.uk/biostudies/studies/S-BSST936. The HAPNEST software for generating synthetic datasets is available as Docker/Singularity containers and open source Julia and C code at https://github.com/intervene-EU-H2020/synthetic_data. Sophie Wharrie, Vishnu Raj, Remo Monti, Ying Wang 0074, Alicia Martin, Luke J. O'Connor, Samuel Kaski, Pekka Marttinen, Pier Francesco Palamara, Christoph Lippert, Andrea Ganna |
Bioinform. | 10 |
| 2023 | Quantifying Movement Behavior of Chronic Low Back Pain Patients in Virtual RealityabstractChronic low back pain (CLBP) is a globally common musculoskeletal problem. Measuring the sensation of pain and the effect of a treatment has always been a challenge for healthcare. Here, we study how the movement data, collected while using a virtual reality (VR) program, could be used as an objective measurement in patients with CLBP. A specific data collection method based on VR was developed and used with CLBP patients and healthy volunteers. We demonstrate that the movement data in VR can be used to classify individuals in these two groups with a high accuracy by using logistic regression. The most discriminative features are the duration of the movements and the total variation of movement velocity. Furthermore, we show that hidden Markov models can divide movement data into meaningful segments, which creates possibilities for defining even more detailed features, with potential to improve accuracy, when larger datasets become available in the future. Tommi Gröhn, Sammeli Liikkanen, Teppo Huttunen, Mika Mäkinen, Pasi Liljeberg, Pekka Marttinen |
ACM Trans. Comput. Heal. | 6 |
| 2023 | Deformation equivariant cross-modality image synthesis with paired non-aligned training dataabstractCross-modality image synthesis is an active research topic with multiple medical clinically relevant applications. Recently, methods allowing training with paired but misaligned data have started to emerge. However, no robust and well-performing methods applicable to a wide range of real world data sets exist. In this work, we propose a generic solution to the problem of cross-modality image synthesis with paired but non-aligned data by introducing new deformation equivariance encouraging loss functions. The method consists of joint training of an image synthesis network together with separate registration networks and allows adversarial training conditioned on the input even with misaligned data. The work lowers the bar for new clinical applications by allowing effortless training of cross-modality image synthesis networks for more difficult data sets. Joel Honkamaa, Sonja Koivukoski, Mira Valkonen, Leena Latonen, Pekka Ruusuvuori, Pekka Marttinen |
Medical Image Anal. | 7 |
| 2023 | Unsupervised feature selection based on variance-covariance subspace distanceabstractSubspace distance is an invaluable tool exploited in a wide range of feature selection methods. The power of subspace distance is that it can identify a representative subspace, including a group of features that can efficiently approximate the space of original features. On the other hand, employing intrinsic statistical information of data can play a significant role in a feature selection process. Nevertheless, most of the existing feature selection methods founded on the subspace distance are limited in properly fulfilling this objective. To pursue this void, we propose a framework that takes a subspace distance into account which is called "Variance-Covariance subspace distance". The approach gains advantages from the correlation of information included in the features of data, thus determines all the feature subsets whose corresponding Variance-Covariance matrix has the minimum norm property. Consequently, a novel, yet efficient unsupervised feature selection framework is introduced based on the Variance-Covariance distance to handle both the dimensionality reduction and subspace learning tasks. The proposed framework has the ability to exclude those features that have the least variance from the original feature set. Moreover, an efficient update algorithm is provided along with its associated convergence analysis to solve the optimization side of the proposed approach. An extensive number of experiments on nine benchmark datasets are also conducted to assess the performance of our method from which the results demonstrate its superiority over a variety of state-of-the-art unsupervised feature selection methods. The source code is available at https://github.com/SaeedKarami/VCSDFS. Saeed Karami, Farid Saberi Movahed, Prayag Tiwari, Pekka Marttinen, Sahar Vahdati |
Neural Networks | 4 |
| 2023 | Bayesian modeling of the impact of antibiotic resistance on the efficiency of MRSA decolonizationabstractMethicillin-resistant Staphylococcus aureus (MRSA) is a major cause of morbidity and mortality. Colonization by MRSA increases the risk of infection and transmission, underscoring the importance of decolonization efforts. However, success of these decolonization protocols varies, raising the possibility that some MRSA strains may be more persistent than others. Here, we studied how the persistence of MRSA colonization correlates with genomic presence of antibiotic resistance genes. Our analysis using a Bayesian mixed effects survival model found that genetic determinants of high-level resistance to mupirocin was strongly associated with failure of the decolonization protocol. However, we did not see a similar effect with genetic resistance to chlorhexidine or other antibiotics. Including strain-specific random effects improved the predictive performance, indicating that some strain characteristics other than resistance also contributed to persistence. Study subject-specific random effects did not improve the model. Our results highlight the need to consider the properties of the colonizing MRSA strain when deciding which treatments to include in the decolonization protocol. Fanni Ojala, Mohamad R. Abdul Sater, Loren G. Miller, James A. McKinnell, Mary K. Hayden, Susan S. Huang, Yonatan H. Grad, Pekka Marttinen |
PLoS Comput. Biol. | 8 |
| 2023 | Multitask Balanced and Recalibrated Network for Medical Code PredictionabstractHuman coders assign standardized medical codes to clinical documents generated during patients’ hospitalization, which is error prone and labor intensive. Automated medical coding approaches have been developed using machine learning methods, such as deep neural networks. Nevertheless, automated medical coding is still challenging because of complex code association, noise in lengthy documents, and the imbalanced class problem. We propose a novel neural network, called the Multitask Balanced and Recalibrated Neural Network, to solve these issues. Significantly, the multitask learning scheme shares the relationship knowledge between different coding branches to capture code association. A recalibrated aggregation module is developed by cascading convolutional blocks to extract high-level semantic features that mitigate the impact of noise in documents. Also, the cascaded structure of the recalibrated module can benefit learning from lengthy notes. To solve the imbalanced class problem, we deploy focal loss to redistribute the attention on low- and high-frequency medical codes. Experimental results show that our proposed model outperforms competitive baselines on a real-world clinical dataset called the Medical Information Mart for Intensive Care (MIMIC-III). Wei Sun 0046, Shaoxiong Ji, Erik Cambria, Pekka Marttinen |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2022 | Deconfounded Representation Similarity for Comparison of Neural NetworksabstractSimilarity metrics such as representational similarity analysis (RSA) and centered kernel alignment (CKA) have been used to understand neural networks by comparing their layer-wise representations. However, these metrics are confounded by the population structure of data items in the input space, leading to inconsistent conclusions about the \emph{functional} similarity between neural networks, such as spuriously high similarity of completely random neural networks and inconsistent domain relations in transfer learning. We introduce a simple and generally applicable fix to adjust for the confounder with covariate adjustment regression, which improves the ability of CKA and RSA to reveal functional similarity and also retains the intuitive invariance properties of the original similarity measures. We show that deconfounding the similarity metrics increases the resolution of detecting functionally similar neural networks across domains. Moreover, in real-world applications, deconfounding improves the consistency between CKA and domain similarity in transfer learning, and increases the correlation between CKA and model out-of-distribution accuracy similarity. Tianyu Cui, Pekka Marttinen, Samuel Kaski |
NeurIPS | 3 |
| 2022 | Contextualized Graph Embeddings for Adverse Drug Event DetectionabstractAbstract An adverse drug event (ADE) is defined as an adverse reaction resulting from improper drug use, reported in various documents such as biomedical literature, drug reviews, and user posts on social media. The recent advances in natural language processing techniques have facilitated automated ADE detection from documents. However, the contextualized information and relations among text pieces are less explored. This paper investigates contextualized language models and heterogeneous graph representations. It builds a contextualized graph embedding model for adverse drug event detection. We employ different convolutional graph neural networks and pre-trained contextualized embeddings as the building blocks. Experimental results show that our methods can improve the performance by comparing recent ADE detection models, suggesting that a text graph can capture causal relationships and dependency between different entities in a document. Ya Gao 0005, Shaoxiong Ji, Tongxuan Zhang, Prayag Tiwari, Pekka Marttinen |
ECML/PKDD (2) | 5 |
| 2022 | COVIDNet: An Automatic Architecture for COVID-19 Detection With Deep Learning From Chest X-Ray ImagesabstractUp to now, the coronavirus disease 2019 (COVID-19) has been sweeping across all over the world, which has affected individual’s lives in an overwhelming way. To fight efficiently against the COVID-19, radiography and radiology images are used by clinicians in hospitals. This article presents an integrated framework, named COVIDNet, for classifying COVID-19 patients and healthy controls. Specifically, ResNet (i.e., ResNet-18 and ResNet-50) is adopted as a backbone network to extract the discriminative features first. Second, the spatial pyramid pooling (SPP) layer is adopted to capture the middle-level features from the features of ResNet. To learn the high-level features, the NetVLAD layer is used to aggregate the features representation from middle-level features. The context gating (CG) mechanism is adopted to further learn the high-level features for predicting the COVID-19 patients or not. Finally, extensive experiments are conducted on the collected database, showing the excellent performance of the proposed integrated architecture, with the sensitivity up to 97% and specificity of 99.5% of the ResNet-18, and with the sensitivity up to 99% and specificity of 99.4% of the ResNet-50. Prayag Tiwari, Xiuying Shi, Pekka Marttinen, Neeraj Kumar 0001 |
IEEE Internet Things J. | 5 |
| 2022 | A Survey on Knowledge Graphs: Representation, Acquisition, and ApplicationsabstractHuman knowledge provides a formal understanding of the world. Knowledge graphs that represent structural relations between entities have become an increasingly popular research direction toward cognition and human-level intelligence. In this survey, we provide a comprehensive review of the knowledge graph covering overall research topics about: 1) knowledge graph representation learning; 2) knowledge acquisition and completion; 3) temporal knowledge graph; and 4) knowledge-aware applications and summarize recent breakthroughs and perspective directions to facilitate future research. We propose a full-view categorization and new taxonomies on these topics. Knowledge graph embedding is organized from four aspects of representation space, scoring function, encoding models, and auxiliary information. For knowledge acquisition, especially knowledge graph completion, embedding methods, path inference, and logical rule reasoning are reviewed. We further explore several emerging topics, including metarelational learning, commonsense reasoning, and temporal knowledge graphs. To facilitate future research on knowledge graphs, we also provide a curated collection of data sets and open-source libraries on different tasks. In the end, we have a thorough outlook on several promising research directions. Shaoxiong Ji, Shirui Pan, Erik Cambria, Pekka Marttinen, Philip S. Yu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | A Critical Look at the Consistency of Causal Estimation with Deep Latent Variable ModelsabstractUsing deep latent variable models in causal inference has attracted considerable interest recently, but an essential open question is their ability to yield consistent causal estimates. While they have demonstrated promising results and theory exists on some simple model formulations, we also know that causal effects are not even identifiable in general with latent variables. We investigate this gap between theory and empirical results with analytical considerations and extensive experiments under multiple synthetic and real-world data sets, using the causal effect variational autoencoder (CEVAE) as a case study. While CEVAE seems to work reliably under some simple scenarios, it does not estimate the causal effect correctly with a misspecified latent variable or a complex data distribution, as opposed to its original motivation. Hence, our results show that more attention should be paid to ensuring the correctness of causal estimates with deep latent variable models. Severi Rissanen, Pekka Marttinen |
NeurIPS | 2 |
| 2021 | Multitask Recalibrated Aggregation Network for Medical Code PredictionabstractAbstract Medical coding translates professionally written medical reports into standardized codes, which is an essential part of medical information systems and health insurance reimbursement. Manual coding by trained human coders is time-consuming and error-prone. Thus, automated coding algorithms have been developed, building especially on the recent advances in machine learning and deep neural networks. To solve the challenges of encoding lengthy and noisy clinical documents and capturing code associations, we propose a multitask recalibrated aggregation network. In particular, multitask learning shares information across different coding schemes and captures the dependencies between different medical codes. Feature recalibration and aggregation in shared modules enhance representation learning for lengthy notes. Experiments with a real-world MIMIC-III dataset show significantly improved predictive performance. Wei Sun 0046, Shaoxiong Ji, Erik Cambria, Pekka Marttinen |
ECML/PKDD (4) | 4 |
| 2021 | Errors-in-Variables Modeling of Personalized Treatment-Response TrajectoriesabstractEstimating the impact of a treatment on a given response is needed in many biomedical applications. However, methodology is lacking for the case when the response is a continuous temporal curve, treatment covariates suffer extensively from measurement error, and even the exact timing of the treatments is unknown. We introduce a novel method for this challenging scenario. We model personalized treatment-response curves as a combination of parametric response functions, hierarchically sharing information across individuals, and a sparse Gaussian process for the baseline trend. Importantly, our model accounts for errors not only in treatment covariates, but also in treatment timings, a problem arising in practice for example when data on treatments are based on user self-reporting. We validate our model with simulated and real patient data, and show that in a challenging application of estimating the impact of diet on continuous blood glucose measurements, accounting for measurement error significantly improves estimation and prediction accuracy. Guangyi Zhang 0001, Reza A. Ashrafi, Anne Juuti, Kirsi Pietiläinen, Pekka Marttinen |
IEEE J. Biomed. Health Informatics | 5 |
| 2020 | Learning Global Pairwise Interactions with Bayesian Neural NetworksabstractEstimating global pairwise interaction effects, i.e., the difference between the joint effect and the sum of marginal effects of two input features, with uncertainty properly quantified, is centrally important in science applications. We propose a non-parametric probabilistic method for detecting interaction effects of unknown form. First, the relationship between the features and the output is modelled using a Bayesian neural network, capable of representing complex interactions and principled uncertainty. Second, interaction effects and their uncertainty are estimated from the trained model. For the second step, we propose an intuitive global interaction measure: Bayesian Group Expected Hessian (GEH), which aggregates information of local interactions as captured by the Hessian. GEH provides a natural trade-off between type I and type II error and, moreover, comes with theoretical guarantees ensuring that the estimated interaction effects and their uncertainty can be improved by training a more accurate BNN. The method empirically outperforms available non-probabilistic alternatives on simulated and real-world data. Finally, we demonstrate its ability to detect interpretable interactions between higher-level features (at deeper layers of the neural network). Tianyu Cui, Pekka Marttinen, Samuel Kaski |
ECAI | 2 |
| 2020 | Batch simulations and uncertainty quantification in Gaussian process surrogate approximate Bayesian computationabstractThe computational efficiency of approximate Bayesian computation (ABC) has been improved by using surrogate models such as Gaussian processes (GP). In one such promising framework the discrepancy between the simulated and observed data is modelled with a GP which is further used to form a model-based estimator for the intractable posterior. In this article we improve this approach in several ways. We develop batch-sequential Bayesian experimental design strategies to parallellise the expensive simulations. In earlier work only sequential strategies have been used. Current surrogate-based ABC methods also do not fully account the uncertainty due to the limited budget of simulations as they output only a point estimate of the ABC posterior. We propose a numerical method to fully quantify the uncertainty in, for example, ABC posterior moments. We also provide some new analysis on the GP modelling assumptions in the resulting improved framework called Bayesian ABC and discuss its connection to Bayesian quadrature (BQ) and Bayesian optimisation (BO). Experiments with toy and real-world simulation models demonstrate advantages of the proposed techniques. Marko Järvenpää, Aki Vehtari, Pekka Marttinen |
UAI | 3 |
| 2019 | Modelling G×E with historical weather information improves genomic prediction in new environmentsabstractMOTIVATION: Interaction between the genotype and the environment (G×E) has a strong impact on the yield of major crop plants. Although influential, taking G×E explicitly into account in plant breeding has remained difficult. Recently G×E has been predicted from environmental and genomic covariates, but existing works have not shown that generalization to new environments and years without access to in-season data is possible and practical applicability remains unclear. Using data from a Barley breeding programme in Finland, we construct an in silico experiment to study the viability of G×E prediction under practical constraints. RESULTS: We show that the response to the environment of a new generation of untested Barley cultivars can be predicted in new locations and years using genomic data, machine learning and historical weather observations for the new locations. Our results highlight the need for models of G×E: non-linear effects clearly dominate linear ones, and the interaction between the soil type and daily rain is identified as the main driver for G×E for Barley in Finland. Our study implies that genomic selection can be used to capture the yield potential in G×E effects for future growth seasons, providing a possible means to achieve yield improvements, needed for feeding the growing population. AVAILABILITY AND IMPLEMENTATION: The data accompanied by the method code (http://research.cs.aalto.fi/pml/software/gxe/bioinformatics_codes.zip) is available in the form of kernels to allow reproducing the results. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jussi Gillberg, Pekka Marttinen, Hiroshi Mamitsuka, Samuel Kaski |
Bioinform. | 2 |
| 2019 | A Bayesian model of acquisition and clearance of bacterial colonization incorporating within-host variationabstractBacterial populations that colonize a host can play important roles in host health, including serving as a reservoir that transmits to other hosts and from which invasive strains emerge, thus emphasizing the importance of understanding rates of acquisition and clearance of colonizing populations. Studies of colonization dynamics have been based on assessment of whether serial samples represent a single population or distinct colonization events. With the use of whole genome sequencing to determine genetic distance between isolates, a common solution to estimate acquisition and clearance rates has been to assume a fixed genetic distance threshold below which isolates are considered to represent the same strain. However, this approach is often inadequate to account for the diversity of the underlying within-host evolving population, the time intervals between consecutive measurements, and the uncertainty in the estimated acquisition and clearance rates. Here, we present a fully Bayesian model that provides probabilities of whether two strains should be considered the same, allowing us to determine bacterial clearance and acquisition from genomes sampled over time. Our method explicitly models the within-host variation using population genetic simulation, and the inference is done using a combination of Approximate Bayesian Computation (ABC) and Markov Chain Monte Carlo (MCMC). We validate the method with multiple carefully conducted simulations and demonstrate its use in practice by analyzing a collection of methicillin resistant Staphylococcus aureus (MRSA) isolates from a large recently completed longitudinal clinical study. An R-code implementation of the method is freely available at: https://github.com/mjarvenpaa/bacterial-colonization-model. Marko Järvenpää, Mohamad R. Abdul Sater, Georgia K. Lagoudas, Paul C. Blainey, Loren G. Miller, James A. McKinnell, Susan S. Huang, Yonatan H. Grad, Pekka Marttinen |
PLoS Comput. Biol. | 9 |
| 2018 | Bacmeta: simulator for genomic evolution in bacterial metapopulationsabstractSummary: The advent of genomic data from densely sampled bacterial populations has created a need for flexible simulators by which models and hypotheses can be efficiently investigated in the light of empirical observations. Bacmeta provides fast stochastic simulation of neutral evolution within a large collection of interconnected bacterial populations with completely adjustable connectivity network. Stochastic events of mutations, recombinations, insertions/deletions, migrations and micro-epidemics can be simulated in discrete non-overlapping generations with a Wright-Fisher model that operates on explicit sequence data of any desired genome length. Each model component, including locus, bacterial strain, population and ultimately the whole metapopulation, is efficiently simulated using C++ objects and detailed metadata from each level can be acquired. The software can be executed in a cluster environment using simple textual input files, enabling, e.g. large-scale simulations and likelihood-free inference. Availability and implementation: Bacmeta is implemented with C++ for Linux, Mac and Windows. It is available at https://bitbucket.org/aleksisipola/bacmeta under the BSD 3-clause license. Supplementary information: Supplementary data are available at Bioinformatics online. Aleksi Sipola, Pekka Marttinen, Jukka Corander |
Bioinform. | 2 |
| 2018 | Improving genomics-based predictions for precision medicine through active elicitation of expert knowledgeabstractMotivation: Precision medicine requires the ability to predict the efficacies of different treatments for a given individual using high-dimensional genomic measurements. However, identifying predictive features remains a challenge when the sample size is small. Incorporating expert knowledge offers a promising approach to improve predictions, but collecting such knowledge is laborious if the number of candidate features is very large. Results: We introduce a probabilistic framework to incorporate expert feedback about the impact of genomic measurements on the outcome of interest and present a novel approach to collect the feedback efficiently, based on Bayesian experimental design. The new approach outperformed other recent alternatives in two medical applications: prediction of metabolic traits and prediction of sensitivity of cancer cells to different drugs, both using genomic features as predictors. Furthermore, the intelligent approach to collect feedback reduced the workload of the expert to approximately 11%, compared to a baseline approach. Availability and implementation: Source code implementing the introduced computational methods is freely available at https://github.com/AaltoPML/knowledge-elicitation-for-precision-medicine. Supplementary information: Supplementary data are available at Bioinformatics online. Iiris Sundin, Tomi Peltola, Luana Micallef, Homayun Afrabandpey, Marta Soare, Muntasir Mamun Majumder, Pedram Daee, Chen He 0003, Baris Serim, Aki S. Havulinna, Caroline Heckman, Giulio Jacucci, Pekka Marttinen, Samuel Kaski |
Bioinform. | 13 |
| 2018 | ELFI: Engine for Likelihood-Free InferenceabstractEngine for Likelihood-Free Inference (ELFI) is a Python software library for performing likelihood-free inference (LFI). ELFI provides a convenient syntax for arranging components in LFI, such as priors, simulators, summaries or distances, to a network called ELFI graph. The components can be implemented in a wide variety of languages. The stand-alone ELFI graph can be used with any of the available inference methods without modifications. A central method implemented in ELFI is Bayesian Optimization for Likelihood-Free Inference (BOLFI), which has recently been shown to accelerate likelihood-free inference up to several orders of magnitude by surrogate-modelling the distance. ELFI also has an inbuilt support for output data storing for reuse and analysis, and supports parallelization of computation from multiple cores up to a cluster environment. ELFI is designed to be extensible and provides interfaces for widening its functionality. This makes the adding of new inference methods to ELFI straightforward and automatically compatible with the inbuilt features. Jarno Lintusaari, Henri Vuollekoski, Antti Kangasrääsiö, Kusti Skytén, Marko Järvenpää, Pekka Marttinen, Michael U. Gutmann, Aki Vehtari, Jukka Corander, Samuel Kaski |
J. Mach. Learn. Res. | 6 |
| 2017 | Interactive Elicitation of Knowledge on Feature Relevance Improves Predictions in Small Data SetsabstractProviding accurate predictions is challenging for machine learning algorithms when the number of features is larger than the number of samples in the data. Prior knowledge can improve machine learning models by indicating relevant variables and parameter values. Yet, this prior knowledge is often tacit and only available from domain experts. We present a novel approach that uses interactive visualization to elicit the tacit prior knowledge and uses it to improve the accuracy of prediction models. The main component of our approach is a user model that models the domain expert's knowledge of the relevance of different features for a prediction task. In particular, based on the expert's earlier input, the user model guides the selection of the features on which to elicit user's knowledge next. The results of a controlled user study show that the user model significantly improves prior knowledge elicitation and prediction accuracy, when predicting the relative citation counts of scientific documents in a specific domain. Luana Micallef, Iiris Sundin, Pekka Marttinen, Muhammad Ammad-ud-din, Tomi Peltola, Marta Soare, Giulio Jacucci, Samuel Kaski |
IUI | 3 |
| 2017 | biMM: efficient estimation of genetic variances and covariances for cohorts with high-dimensional phenotype measurementsabstractSUMMARY: Genetic research utilizes a decomposition of trait variances and covariances into genetic and environmental parts. Our software package biMM is a computationally efficient implementation of a bivariate linear mixed model for settings where hundreds of traits have been measured on partially overlapping sets of individuals. AVAILABILITY AND IMPLEMENTATION: Implementation in R freely available at www.iki.fi/mpirinen . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Matti Pirinen, Christian Benner, Pekka Marttinen, Marjo-Riitta Järvelin, Manuel A. Rivas, Samuli Ripatti |
Bioinform. | 3 |
| 2017 | Speciation trajectories in recombining bacterial speciesabstractIt is generally agreed that bacterial diversity can be classified into genetically and ecologically cohesive units, but what produces such variation is a topic of intensive research. Recombination may maintain coherent species of frequently recombining bacteria, but the emergence of distinct clusters within a recombining species, and the impact of habitat structure in this process are not well described, limiting our understanding of how new species are created. Here we present a model of bacterial evolution in overlapping habitat space. We show that the amount of habitat overlap determines the outcome for a pair of clusters, which may range from fast clonal divergence with little interaction between the clusters to a stationary population structure, where different clusters maintain an equilibrium distance between each other for an indefinite time. We fit our model to two data sets. In Streptococcus pneumoniae, we find a genomically and ecologically distinct subset, held at a relatively constant genetic distance from the majority of the population through frequent recombination with it, while in Campylobacter jejuni, we find a minority population we predict will continue to diverge at a higher rate. This approach may predict and define speciation trajectories in multiple bacterial species. Pekka Marttinen, William P. Hanage |
PLoS Comput. Biol. | 1 |
| 2016 | metaCCA: summary statistics-based multivariate meta-analysis of genome-wide association studies using canonical correlation analysisabstractMOTIVATION: A dominant approach to genetic association studies is to perform univariate tests between genotype-phenotype pairs. However, analyzing related traits together increases statistical power, and certain complex associations become detectable only when several variants are tested jointly. Currently, modest sample sizes of individual cohorts, and restricted availability of individual-level genotype-phenotype data across the cohorts limit conducting multivariate tests. RESULTS: We introduce metaCCA, a computational framework for summary statistics-based analysis of a single or multiple studies that allows multivariate representation of both genotype and phenotype. It extends the statistical technique of canonical correlation analysis to the setting where original individual-level records are not available, and employs a covariance shrinkage algorithm to achieve robustness.Multivariate meta-analysis of two Finnish studies of nuclear magnetic resonance metabolomics by metaCCA, using standard univariate output from the program SNPTEST, shows an excellent agreement with the pooled individual-level analysis of original data. Motivated by strong multivariate signals in the lipid genes tested, we envision that multivariate association testing using metaCCA has a great potential to provide novel insights from already published summary statistics from high-throughput phenotyping technologies. AVAILABILITY AND IMPLEMENTATION: Code is available at https://github.com/aalto-ics-kepaco CONTACTS: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Anna Cichonska, Juho Rousu, Pekka Marttinen, Antti J. Kangas, Pasi Soininen, Terho Lehtimäki, Olli T. Raitakari, Marjo-Riitta Järvelin, Veikko Salomaa, Mika Ala-Korpela, Samuli Ripatti, Matti Pirinen |
Bioinform. | 3 |
| 2016 | Multiple Output Regression with Latent NoiseabstractIn high-dimensional data, structured noise caused by observed and unobserved factors affecting multiple target variables simultaneously, imposes a serious challenge for modeling, by masking the often weak signal. Therefore, (1) explaining away the structured noise in multiple-output regression is of paramount importance. Additionally, (2) assumptions about the correlation structure of the regression weights are needed. We note that both can be formulated in a natural way in a latent variable model, in which both the interesting signal and the noise are mediated through the same latent factors. Under this assumption, the signal model then borrows strength from the noise model by encouraging similar effects on correlated targets. We introduce a hyperparameter for the latent signal-to-noise ratio which turns out to be important for modelling weak signals, and an ordered infinite-dimensional shrinkage prior that resolves the rotational unidentifiability in reduced-rank regression models. Simulations and prediction experiments with metabolite, gene expression, FMRI measurement, and macroeconomic time series data show that our model equals or exceeds the state-of-the-art performance and, in particular, outperforms the standard approach of assuming independent noise and signal models. Jussi Gillberg, Pekka Marttinen, Matti Pirinen, Antti J. Kangas, Pasi Soininen, Mehreen Ali, Aki S. Havulinna, Marjo-Riitta Järvelin, Mika Ala-Korpela, Samuel Kaski |
J. Mach. Learn. Res. | 2 |
| 2014 | Assessing multivariate gene-metabolome associations with rare variants using Bayesian reduced rank regressionabstractMOTIVATION: A typical genome-wide association study searches for associations between single nucleotide polymorphisms (SNPs) and a univariate phenotype. However, there is a growing interest to investigate associations between genomics data and multivariate phenotypes, for example, in gene expression or metabolomics studies. A common approach is to perform a univariate test between each genotype-phenotype pair, and then to apply a stringent significance cutoff to account for the large number of tests performed. However, this approach has limited ability to uncover dependencies involving multiple variables. Another trend in the current genetics is the investigation of the impact of rare variants on the phenotype, where the standard methods often fail owing to lack of power when the minor allele is present in only a limited number of individuals. RESULTS: We propose a new statistical approach based on Bayesian reduced rank regression to assess the impact of multiple SNPs on a high-dimensional phenotype. Because of the method's ability to combine information over multiple SNPs and phenotypes, it is particularly suitable for detecting associations involving rare variants. We demonstrate the potential of our method and compare it with alternatives using the Northern Finland Birth Cohort with 4702 individuals, for whom genome-wide SNP data along with lipoprotein profiles comprising 74 traits are available. We discovered two genes (XRCC4 and MTHFD2L) without previously reported associations, which replicated in a combined analysis of two additional cohorts: 2390 individuals from the Cardiovascular Risk in Young Finns study and 3659 individuals from the FINRISK study. AVAILABILITY AND IMPLEMENTATION: R-code freely available for download at http://users.ics.aalto.fi/pemartti/gene_metabolome/. Pekka Marttinen, Matti Pirinen, Antti-Pekka Sarin, Jussi Gillberg, Johannes Kettunen, Ida Surakka, Antti J. Kangas, Pasi Soininen, Paul F. O'Reilly, Marika Kaakinen, Mika Kähönen, Terho Lehtimäki, Mika Ala-Korpela, Olli T. Raitakari, Veikko Salomaa, Marjo-Riitta Järvelin, Samuli Ripatti, Samuel Kaski |
Bioinform. | 1 |
| 2010 | Efficient Bayesian approach for multilocus association mapping including gene-gene interactionsabstractBACKGROUND: since the introduction of large-scale genotyping methods that can be utilized in genome-wide association (GWA) studies for deciphering complex diseases, statistical genetics has been posed with a tremendous challenge of how to most appropriately analyze such data. A plethora of advanced model-based methods for genetic mapping of traits has been available for more than 10 years in animal and plant breeding. However, most such methods are computationally intractable in the context of genome-wide studies. Therefore, it is hardly surprising that GWA analyses have in practice been dominated by simple statistical tests concerned with a single marker locus at a time, while the more advanced approaches have appeared only relatively recently in the biomedical and statistical literature. RESULTS: we introduce a novel Bayesian modeling framework for association mapping which enables the detection of multiple loci and their interactions that influence a dichotomous phenotype of interest. The method is shown to perform well in a simulation study when compared to widely used standard alternatives and its computational complexity is typically considerably smaller than that of a maximum likelihood based approach. We also discuss in detail the sensitivity of the Bayesian inferences with respect to the choice of prior distributions in the GWA context. CONCLUSIONS: our results show that the Bayesian model averaging approach which explicitly considers gene-gene interactions may improve the detection of disease associated genetic markers in two respects: first, by providing better estimates of the locations of the causal loci; second, by reducing the number of false positives. The benefits are most apparent when the interacting genes exhibit no main effects. However, our findings also illustrate that such an approach is somewhat sensitive to the prior distribution assigned on the model structure. Pekka Marttinen, Jukka Corander |
BMC Bioinform. | 1 |
| 2009 | Bayesian clustering and feature selection for cancer tissue samplesabstractBACKGROUND: The versatility of DNA copy number amplifications for profiling and categorization of various tissue samples has been widely acknowledged in the biomedical literature. For instance, this type of measurement techniques provides possibilities for exploring sets of cancerous tissues to identify novel subtypes. The previously utilized statistical approaches to various kinds of analyses include traditional algorithmic techniques for clustering and dimension reduction, such as independent and principal component analyses, hierarchical clustering, as well as model-based clustering using maximum likelihood estimation for latent class models. RESULTS: While purely algorithmic methods are usually easily applicable, their suboptimal performance and limitations in making formal inference have been thoroughly discussed in the statistical literature. Here we introduce a Bayesian model-based approach to simultaneous identification of underlying tissue groups and the informative amplifications. The model-based approach provides the possibility of using formal inference to determine the number of groups from the data, in contrast to the ad hoc methods often exploited for similar purposes. The model also automatically recognizes the chromosomal areas that are relevant for the clustering. CONCLUSION: Validatory analyses of simulated data and a large database of DNA copy number amplifications in human neoplasms are used to illustrate the potential of our approach. Our software implementation BASTA for performing Bayesian statistical tissue profiling is freely available for academic purposes at (http://web.abo.fi/fak/mnf/mate/jc/software/basta.html). Pekka Marttinen, Samuel Myllykangas, Jukka Corander |
BMC Bioinform. | 1 |
| 2009 | Robust extraction of functional signals from gene set analysis using a generalized threshold free scoring functionabstractBACKGROUND: A central task in contemporary biosciences is the identification of biological processes showing response in genome-wide differential gene expression experiments. Two types of analysis are common. Either, one generates an ordered list based on the differential expression values of the probed genes and examines the tail areas of the list for over-representation of various functional classes. Alternatively, one monitors the average differential expression level of genes belonging to a given functional class. So far these two types of method have not been combined. RESULTS: We introduce a scoring function, Gene Set Z-score (GSZ), for the analysis of functional class over-representation that combines two previous analysis methods. GSZ encompasses popular functions such as correlation, hypergeometric test, Max-Mean and Random Sets as limiting cases. GSZ is stable against changes in class size as well as across different positions of the analysed gene list in tests with randomized data. GSZ shows the best overall performance in a detailed comparison to popular functions using artificial data. Likewise, GSZ stands out in a cross-validation of methods using split real data. A comparison of empirical p-values further shows a strong difference in favour of GSZ, which clearly reports better p-values for top classes than the other methods. Furthermore, GSZ detects relevant biological themes that are missed by the other methods. These observations also hold when comparing GSZ with popular program packages. CONCLUSION: GSZ and improved versions of earlier methods are a useful contribution to the analysis of differential gene expression. The methods and supplementary material are available from the website http://ekhidna.biocenter.helsinki.fi/users/petri/public/GSZ/GSZscore.html. Petri Törönen, Pauli J. Ojala, Pekka Marttinen, Liisa Holm |
BMC Bioinform. | 3 |
| 2009 | Bayesian learning of graphical vector autoregressions with unequal lag-lengths
Pekka Marttinen, Jukka Corander |
Mach. Learn. | 1 |
| 2009 | Bayesian Clustering of Fuzzy Feature Vectors Using a Quasi-Likelihood ApproachabstractBayesian model-based classifiers, both unsupervised and supervised, have been studied extensively and their value and versatility have been demonstrated on a wide spectrum of applications within science and engineering. A majority of the classifiers are built on the assumption of intrinsic discreteness of the considered data features or on the discretization of them prior to the modeling. On the other hand, Gaussian mixture classifiers have also been utilized to a large extent for continuous features in the Bayesian framework. Often the primary reason for discretization in the classification context is the simplification of the analytical and numerical properties of the models. However, the discretization can be problematic due to its \textit{ad hoc} nature and the decreased statistical power to detect the correct classes in the resulting procedure. We introduce an unsupervised classification approach for fuzzy feature vectors that utilizes a discrete model structure while preserving the continuous characteristics of data. This is achieved by replacing the ordinary likelihood by a binomial quasi-likelihood to yield an analytical expression for the posterior probability of a given clustering solution. The resulting model can be justified from an information-theoretic perspective. Our method is shown to yield highly accurate clusterings for challenging synthetic and empirical data sets. Pekka Marttinen, Jing Tang 0002, Bernard De Baets, Peter Dawyndt, Jukka Corander |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2008 | Enhanced Bayesian modelling in BAPS software for learning genetic structures of populationsabstractBACKGROUND: During the most recent decade many Bayesian statistical models and software for answering questions related to the genetic structure underlying population samples have appeared in the scientific literature. Most of these methods utilize molecular markers for the inferences, while some are also capable of handling DNA sequence data. In a number of earlier works, we have introduced an array of statistical methods for population genetic inference that are implemented in the software BAPS. However, the complexity of biological problems related to genetic structure analysis keeps increasing such that in many cases the current methods may provide either inappropriate or insufficient solutions. RESULTS: We discuss the necessity of enhancing the statistical approaches to face the challenges posed by the ever-increasing amounts of molecular data generated by scientists over a wide range of research areas and introduce an array of new statistical tools implemented in the most recent version of BAPS. With these methods it is possible, e.g., to fit genetic mixture models using user-specified numbers of clusters and to estimate levels of admixture under a genetic linkage model. Also, alleles representing a different ancestry compared to the average observed genomic positions can be tracked for the sampled individuals, and a priori specified hypotheses about genetic population structure can be directly compared using Bayes' theorem. In general, we have improved further the computational characteristics of the algorithms behind the methods implemented in BAPS facilitating the analyses of large and complex datasets. In particular, analysis of a single dataset can now be spread over multiple computers using a script interface to the software. CONCLUSION: The Bayesian modelling methods introduced in this article represent an array of enhanced tools for learning the genetic structure of populations. Their implementations in the BAPS software are designed to meet the increasing need for analyzing large-scale population genetics data. The software is freely downloadable for Windows, Linux and Mac OS X systems at http://web.abo.fi/fak/mnf//mate/jc/software/baps.html. Jukka Corander, Pekka Marttinen, Jukka Sirén, Jing Tang 0002 |
BMC Bioinform. | 2 |
| 2008 | Bayesian modeling of recombination events in bacterial populationsabstractBACKGROUND: We consider the discovery of recombinant segments jointly with their origins within multilocus DNA sequences from bacteria representing heterogeneous populations of fairly closely related species. The currently available methods for recombination detection capable of probabilistic characterization of uncertainty have a limited applicability in practice as the number of strains in a data set increases. RESULTS: We introduce a Bayesian spatial structural model representing the continuum of origins over sites within the observed sequences, including a probabilistic characterization of uncertainty related to the origin of any particular site. To enable a statistically accurate and practically feasible approach to the analysis of large-scale data sets representing a single genus, we have developed a novel software tool (BRAT, Bayesian Recombination Tracker) implementing the model and the corresponding learning algorithm, which is capable of identifying the posterior optimal structure and to estimate the marginal posterior probabilities of putative origins over the sites. CONCLUSION: A multitude of challenging simulation scenarios and an analysis of real data from seven housekeeping genes of 120 strains of genus Burkholderia are used to illustrate the possibilities offered by our approach. The software is freely available for download at URL http://web.abo.fi/fak/mnf//mate/jc/software/brat.html. Pekka Marttinen, Adam Baldwin, William P. Hanage, Chris Dowson, Eshwar Mahenthiralingam, Jukka Corander |
BMC Bioinform. | 1 |
| 2006 | Bayesian search of functionally divergent protein subgroups and their function specific residuesabstractMOTIVATION: The rapid increase in the amount of protein sequence data has created a need for an automated identification of evolutionarily related subgroups from large datasets. The existing methods typically require a priori specification of the number of putative groups, which defines the resolution of the classification solution. RESULTS: We introduce a Bayesian model-based approach to simultaneous identification of evolutionary groups and conserved parts of the protein sequences. The model-based approach provides an intuitive and efficient way of determining the number of groups from the sequence data, in contrast to the ad hoc methods often exploited for similar purposes. Our model recognizes the areas in the sequences that are relevant for the clustering and regards other areas as noise. We have implemented the method using a fast stochastic optimization algorithm which yields a clustering associated with the estimated maximum posterior probability. The method has been shown to have high specificity and sensitivity in simulated and real clustering tasks. With real datasets the method also highlights the residues close to the active site. AVAILABILITY: Software 'kPax' is available at http://www.rni.helsinki.fi/jic/softa.html Pekka Marttinen, Jukka Corander, Petri Törönen, Liisa Holm |
Bioinform. | 1 |
| 2004 | BAPS 2: enhanced possibilities for the analysis of genetic population structureabstractUNLABELLED: Bayesian statistical methods based on simulation techniques have recently been shown to provide powerful tools for the analysis of genetic population structure. We have previously developed a Markov chain Monte Carlo (MCMC) algorithm for characterizing genetically divergent groups based on molecular markers and geographical sampling design of the dataset. However, for large-scale datasets such algorithms may get stuck to local maxima in the parameter space. Therefore, we have modified our earlier algorithm to support multiple parallel MCMC chains, with enhanced features that enable considerably faster and more reliable estimation compared to the earlier version of the algorithm. We consider also a hierarchical tree representation, from which a Bayesian model-averaged structure estimate can be extracted. The algorithm is implemented in a computer program that features a user-friendly interface and built-in graphics. The enhanced features are illustrated by analyses of simulated data and an extensive human molecular dataset. AVAILABILITY: Freely available at http://www.rni.helsinki.fi/~jic/bapspage.html. Jukka Corander, Patrik Waldmann, Pekka Marttinen, Mikko J. Sillanpää |
Bioinform. | 3 |