Arto Klami

dblp:21/5316 · DBLP profile ↗
← Back
61ranked-venue papers
11as first author
22since 2021 · last 2025
0000-0002-7950-1355ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 45 · 10 first-author · 17 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
YearPublicationVenuePosition
2025 On the Importance of Representation in Imitating Human-Like Gameplay
abstract
Being able to synthesize human-like game play is highly useful for automating play-testing, generating naturally behaving bots, and for assisting game design in general. However, for complex games that require long-term planning this remains a major challenge. Imitation learning (IL) algorithms offer one approach for learning to replicate human-like behavior by training models that imitate the behavior observed in human demonstrations, and in fact much of the research in IL is done in context of games due to fast feedback loops and relatively easy access to the required data. One recent approach employing IL is Video PreTraining (VPT) that leverages representation models trained on massive unlabeled video collection of humans playing a complex game, specifically Minecraft. Despite promising results on a complex task, in this work we empirically demonstrate a failure mode of VPT as a representation learner in Minecraft, showing how the pretrained model is insufficient for providing useful representations for new tasks. We then explain how the fundamental challenge remains even in a substantially simplified game environment. We argue that for many practitioners finetuning the representation models for the tasks of interest is unfeasible, making the overall approach of limited interest for use in game design applications.
Ville Tanskanen, Arto Klami, Ville Hautamäki
CoG2
2025 Stochastic variance-reduced Gaussian variational inference on the Bures-Wasserstein manifold
abstract
Optimization in the Bures-Wasserstein space has been gaining popularity in the machine learning community since it draws connections between variational inference and Wasserstein gradient flows. The variational inference objective function of Kullback–Leibler divergence can be written as the sum of the negative entropy and the potential energy, making forward-backward Euler the method of choice. Notably, the backward step admits a closed-form solution in this case, facilitating the practicality of the scheme. However, the forward step is not exact since the Bures-Wasserstein gradient of the potential energy involves "intractable" expectations. Recent approaches propose using the Monte Carlo method -- in practice a single-sample estimator -- to approximate these terms, resulting in high variance and poor performance. We propose a novel variance-reduced estimator based on the principle of control variates. We theoretically show that this estimator has a smaller variance than the Monte-Carlo estimator in scenarios of interest. We also prove that variance reduction helps improve the optimization bounds of the current analysis. We demonstrate that the proposed estimator gains order-of-magnitude improvements over the previous Bures-Wasserstein methods.
Hoang Phuc Hau Luu, Hanlin Yu, Bernardo Williams, Marcelo Hartmann, Arto Klami
ICLR5
2025 Density Ratio Estimation with Conditional Probability Paths
abstract
Density ratio estimation in high dimensions can be reframed as integrating a certain quantity, the time score, over probability paths which interpolate between the two densities. In practice, the time score has to be estimated based on samples from the two densities. However, existing methods for this problem remain computationally expensive and can yield inaccurate estimates. Inspired by recent advances in generative modeling, we introduce a novel framework for time score estimation, based on a conditioning variable. Choosing the conditioning variable judiciously enables a closed-form objective function. We demonstrate that, compared to previous approaches, our approach results in faster learning of the time score and competitive or better estimation accuracies of the density ratio on challenging tasks. Furthermore, we establish theoretical guarantees on the error of the estimated density ratio.
Hanlin Yu, Arto Klami, Aapo Hyvärinen, Anna Korba, Omar Chehab
ICML2
2025 Geodesic Slice Sampler for Multimodal Distributions with Strong Curvature
abstract
Traditional Markov Chain Monte Carlo sampling methods often struggle with sharp curvatures, intricate geometries, and multimodal distributions. Slice sampling can resolve local exploration inefficiency issues, and Riemannian geometries help with sharp curvatures. Recent extensions enable slice sampling on Riemannian manifolds, but they are restricted to cases where geodesics are available in a closed form. We propose a method that generalizes Hit-and-Run slice sampling to more general geometries tailored to the target distribution, by approximating geodesics as solutions to differential equations. Our approach enables the exploration of the regions with strong curvature and rapid transitions between modes in multimodal distributions. We demonstrate the advantages of the approach over challenging sampling problems.
Bernardo Williams, Hanlin Yu, Hoang Phuc Hau Luu, Georgios Arvanitidis, Arto Klami
UAI5
2025 Estimating expert prior knowledge from optimization trajectories
abstract
A recurring task in research is iterative optimization of a process that can be evaluated only by conducting an experiment. Powerful algorithms for assisting this process exist, but they largely ignore the valuable knowledge of expert scientists. We consider a problem within this general scope, not aiming to automate the optimization but instead studying how to infer tacit expert knowledge. This complements the current literature focusing on how such information is used in the optimization process, paying little attention on how the information is obtained. We consider a new formulation where the expertise is inferred by passively observing a human solving an optimization problem, without requiring explicit elicitation techniques. Our solution leverages concepts from Bayesian optimization (BO) commonly used for automating the optimization, but now these tools are used as a theoretical model for the user behavior instead. We assume the expert solves the task approximately in the same manner as a BO algorithm would and solve what kind of prior knowledge about the target function is consistent with the sequence of choices they made. We introduce the problem and a concrete solution, and show that the recovered priors match the true priors in controlled simulated studies. We also empirically evaluate the robustness of the method against violations of the modeling assumptions and demonstrate it on real user data.
Ville Tanskanen, Petrus Mikkola, Aras Erarslan, Arto Klami
Neurocomputing4
2024 Riemannian Laplace Approximation with the Fisher Metric
abstract
Laplace’s method approximates a target density with a Gaussian distribution at its mode. It is computationally efficient and asymptotically exact for Bayesian inference due to the Bernstein-von Mises theorem, but for complex targets and finite-data posteriors it is often too crude an approximation. A recent generalization of the Laplace Approximation transforms the Gaussian approximation according to a chosen Riemannian geometry providing a richer approximation family, while still retaining computational efficiency. However, as shown here, its properties depend heavily on the chosen metric, indeed the metric adopted in previous work results in approximations that are overly narrow as well as being biased even at the limit of infinite data. We correct this shortcoming by developing the approximation family further, deriving two alternative variants that are exact at the limit of infinite data, extending the theoretical analysis of the method, and demonstrating practical improvements in a range of experiments.
Hanlin Yu, Marcelo Hartmann, Bernardo Williams, Mark A. Girolami, Arto Klami
AISTATS5
2024 Subsystem Discovery in High-Dimensional Time-Series Using Masked Autoencoders
abstract
Deep neural networks are increasingly used for time series tasks, yet they often struggle to interpretably model high-dimensional data. In this context, we consider the task of learning easy to understand connections between time-series variables, and organizing them into subsystems, directly from observed data. Our approach reconstructs multivariate time-series with a masked autoencoder, where all information between individual variables is mediated by a learned adjacency matrix. This intuitive pairwise relationship enables grouping of variables without prior knowledge of cluster quantity or size, and is particularly useful for analyzing complex sensor systems with unknown structural interdependencies. Our method simultaneously learns a useful signal representation and aids in understanding the underlying processes. We show that we can learn the correct subsystems from simulated data, and demonstrate identification of plausible subsystem structure from high-dimensional real-world data. In addition, we show that the model retains high predictive performance.
Teemu Sarapisto, Haoyu Wei, Keijo Heljanko, Arto Klami, Laura Ruotsalainen
ECAI4
2024 Non-geodesically-convex optimization in the Wasserstein space
abstract
We study a class of optimization problems in the Wasserstein space (the space of probability measures) where the objective function is nonconvex along generalized geodesics. Specifically, the objective exhibits some difference-of-convex structure along these geodesics. The setting also encompasses sampling problems where the logarithm of the target distribution is difference-of-convex. We derive multiple convergence insights for a novel semi Forward-Backward Euler scheme under several nonconvex (and possibly nonsmooth) regimes. Notably, the semi Forward-Backward Euler is just a slight modification of the Forward-Backward Euler whose convergence is---to our knowledge---still unknown in our very general non-geodesically-convex setting.
Hoang Phuc Hau Luu, Hanlin Yu, Bernardo Williams, Petrus Mikkola, Marcelo Hartmann, Kai Puolamäki, Arto Klami
NeurIPS7
2024 Preferential Normalizing Flows
abstract
Eliciting a high-dimensional probability distribution from an expert via noisy judgments is notoriously challenging, yet useful for many applications, such as prior elicitation and reward modeling. We introduce a method for eliciting the expert's belief density as a normalizing flow based solely on preferential questions such as comparing or ranking alternatives. This allows eliciting in principle arbitrarily flexible densities, but flow estimation is susceptible to the challenge of collapsing or diverging probability mass that makes it difficult in practice. We tackle this problem by introducing a novel functional prior for the flow, motivated by a decision-theoretic argument, and show empirically that the belief density can be inferred as the function-space maximum a posteriori estimate. We demonstrate our method by eliciting multivariate belief densities of simulated experts, including the prior belief of a general-purpose large language model over a real-world dataset.
Petrus Mikkola, Luigi Acerbi, Arto Klami
NeurIPS3
2024 Quantifying uncertainty of uplift: Trees and T-learners
abstract
Uplift modeling refers to the task of estimating the causal effect of a treatment on an individual, also known as the conditional average treatment effect. However, uplift models do not usually provide uncertainty estimates of the predictions. We explain why estimating uncertainty of the treatment effect is particularly important in many common use cases and we show how epistemic uncertainty of the uplift estimates can be quantified for T-learners and trees. We tested the methods on three empirical datasets and evaluated them on a simulated dataset. We found that high uncertainty might be the result of both modeling choices and properties of the data. Sometimes there is not enough data or the data is simply not rich enough to identify the treatment effect well resulting in high uncertainty. In addition, our results suggest that one commonly used dataset might not be suitable for benchmarking.
Otto Nyberg, Arto Klami
Neurocomputing2
2023 Transformed Gaussian Processes for Characterizing a Model's Discrepancy
Aurélien Nioche, Ville Tanskanen, Marcelo Hartmann, Arto Klami
ACML4
2023 Estimating the Contamination Factor's Distribution in Unsupervised Anomaly Detection
abstract
Anomaly detection methods identify examples that do not follow the expected behaviour, typically in an unsupervised fashion, by assigning real-valued anomaly scores to the examples based on various heuristics. These scores need to be transformed into actual predictions by thresholding so that the proportion of examples marked as anomalies equals the expected proportion of anomalies, called contamination factor. Unfortunately, there are no good methods for estimating the contamination factor itself. We address this need from a Bayesian perspective, introducing a method for estimating the posterior distribution of the contamination factor for a given unlabeled dataset. We leverage several anomaly detectors to capture the basic notion of anomalousness and estimate the contamination using a specific mixture formulation. Empirically on 22 datasets, we show that the estimated distribution is well-calibrated and that setting the threshold using the posterior mean improves the detectors' performance over several alternative methods.
Lorenzo Perini, Paul C. Bürkner, Arto Klami
ICML3
2023 Extending Interpolation Consistency Training for Unsupervised Domain Adaptation
abstract
Interpolation consistency training (ICT) is a semi-supervised learning method that encourages predictions of interpolated samples to be consistent with the interpolation of predictions of the corresponding original samples. It has achieved highly impressive results on semi-supervised learning benchmarks, but has not been evaluated in domain adaptation settings where the distributions of labeled and unlabeled data are different. We extend the ICT principle for domain adaptation tasks, by combining ICT with a gradient reversal mechanism that accounts for the domain shift in an adversarial manner. We show that ICT alone is not sufficient for handling the distribution shift and even deteriorates the performance, but the proposed method achieves good performance on visual domain adaptation benchmarks.
Shayan Gharib, Arto Klami
IJCNN2
2023 Exploring uplift modeling with high class imbalance
abstract
Abstract Uplift modeling refers to individual level causal inference. Existing research on the topic ignores one prevalent and important aspect: high class imbalance. For instance in online environments uplift modeling is used to optimally target ads and discounts, but very few users ever end up clicking an ad or buying. One common approach to deal with imbalance in classification is by undersampling the dataset. In this work, we show how undersampling can be extended to uplift modeling. We propose four undersampling methods for uplift modeling. We compare the proposed methods empirically and show when some methods have a tendency to break down. One key observation is that accounting for the imbalance is particularly important for uplift random forests, which explains the poor performance of the model in earlier works. Undersampling is also crucial for class-variable transformation based models.
Otto Nyberg, Arto Klami
Data Min. Knowl. Discov.2
2023 Prior Specification for Bayesian Matrix Factorization via Prior Predictive Matching
abstract
The behavior of many Bayesian models used in machine learning critically depends on the choice of prior distributions, controlled by some hyperparameters typically selected through Bayesian optimization or cross-validation. This requires repeated, costly, posterior inference. We provide an alternative for selecting good priors without carrying out posterior inference, building on the prior predictive distribution that marginalizes the model parameters. We estimate virtual statistics for data generated by the prior predictive distribution and then optimize over the hyperparameters to learn those for which the virtual statistics match the target values provided by the user or estimated from (a subset of) the observed data. We apply the principle for probabilistic matrix factorization, for which good solutions for prior selection have been missing. We show that for Poisson factorization models we can analytically determine the hyperparameters, including the number of factors, that best replicate the target statistics, and we empirically study the sensitivity of the approach for the model mismatch. We also present a model-independent procedure that determines the hyperparameters for general models by stochastic optimization and demonstrate this extension in the context of hierarchical matrix factorization models.
Eliezer S. Silva, Tomasz Kusmierczyk, Marcelo Hartmann, Arto Klami
J. Mach. Learn. Res.4
2022 Lagrangian manifold Monte Carlo on Monge patches
abstract
The efficiency of Markov Chain Monte Carlo (MCMC) depends on how the underlying geometry of the problem is taken into account. For distributions with strongly varying curvature, Riemannian metrics help in efficient exploration of the target distribution. Unfortunately, they have significant computational overhead due to e.g. repeated inversion of the metric tensor, and current geometric MCMC methods using the Fisher information matrix to induce the manifold are in practice slow. We propose a new alternative Riemannian metric for MCMC, by embedding the target distribution into a higher-dimensional Euclidean space as a Monge patch, thus using the induced metric determined by direct geometric reasoning. Our metric only requires first-order gradient information and has fast inverse and determinants, and allows reducing the computational complexity of individual iterations from cubic to quadratic in the problem dimensionality. We demonstrate how Lagrangian Monte Carlo in this metric efficiently explores the target distributions.
Marcelo Hartmann, Mark A. Girolami, Arto Klami
AISTATS3
2022 How Suitable Is Your Naturalistic Dataset for Theory-based User Modeling?
abstract
Theory-based, or “white-box,” models come with a major benefit that makes them appealing for deployment in user modeling: their parameters are interpretable. However, most theory-based models have been developed in controlled settings, in which researchers determine the experimental design. In contrast, real-world application of these models demands setups that are beyond developer control. In non-experimental, naturalistic settings, the tasks with which users are presented may be very limited, and it is not clear that model parameters can be reliably inferred. This paper describes a technique for assessing whether a naturalistic dataset is suitable for use with a theory-based model. The proposed parameter recovery technique can warn against possible over-confidence in inferred model parameters. This technique also can be used to study conditions under which parameter inference is feasible. The method is demonstrated for two models of decision-making under risk with naturalistic data from a turn-based game.
Aini Putkonen, Aurélien Nioche, Ville Tanskanen, Arto Klami, Antti Oulasvirta
UMAP4
2021 Uplift Modeling with High Class Imbalance
abstract
Uplift modeling refers to estimating the causal effect of a treatment on an individual observation, used for instance to identify customers worth targeting with a discount in e-commerce. We introduce a simple yet effective undersampling strategy for dealing with the prevalent problem of high class imbalance (low conversion rate) in such applications. Our strategy is agnostic to the base learners and produces a 6.5% improvement over the best published benchmark for the largest public uplift data which incidentally exhibits high class imbalance. We also introduce a new metric on calibration for uplift modeling and present a strategy to improve the calibration of the proposed method.
Otto Nyberg, Tomasz Kusmierczyk, Arto Klami
ACML3
2021 Modeling Risky Choices in Unknown Environments
abstract
Decision-theoretic models explain human behavior in choice problems involving uncertainty, in terms of individual tendencies such as risk aversion. However, many classical models of risk require knowing the distribution of possible outcomes (rewards) for all options, limiting their applicability outside of controlled experiments. We study the task of learning such models in contexts where the modeler does not know the distributions but instead can only observe the choices and their outcomes for a user familiar with the decision problems, for example a skilled player playing a digital game. We propose a framework combining two separate components, one for modeling the unknown decision-making environment and another for the risk behavior. By using environment models capable of learning distributions we are able to infer classical models of decision-making under risk from observations of the user’s choices and outcomes alone, and we also demonstrate alternative models for predictive purposes. We validate the approach on artificial data and demonstrate a practical use case in modeling risk attitudes of professional esports teams.
Ville Tanskanen, Chang Rajani, Homayun Afrabandpey, Aini Putkonen, Aurélien Nioche, Arto Klami
ACML6
2021 Reliably Calibrated Isotonic Regression
Otto Nyberg, Arto Klami
PAKDD (1)2
2021 Snapshot hyperspectral imaging using wide dilation networks
abstract
Abstract Hyperspectral (HS) cameras record the spectrum at multiple wavelengths for each pixel in an image, and are used, e.g., for quality control and agricultural remote sensing. We introduce a fast, cost-efficient and mobile method of taking HS images using a regular digital camera equipped with a passive diffraction grating filter, using machine learning for constructing the HS image. The grating distorts the image by effectively mapping the spectral information into spatial dislocations, which we convert into a HS image by a convolutional neural network utilizing novel wide dilation convolutions that accurately model optical properties of diffraction. We demonstrate high-quality HS reconstruction using a model trained on only 271 pairs of diffraction grating and ground truth HS images.
Mikko E. Toivonen, Chang Rajani, Arto Klami
Mach. Vis. Appl.3
2021 Multiscale Cloud Detection in Remote Sensing Images Using a Dual Convolutional Neural Network
abstract
Semantic segmentation by convolutional neural networks (CNN) has advanced the state of the art in pixel-level classification of remote sensing images. However, processing large images typically requires analyzing the image in small patches, and hence, features that have a large spatial extent still cause challenges in tasks, such as cloud masking. To support a wider scale of spatial features while simultaneously reducing computational requirements for large satellite images, we propose an architecture of two cascaded CNN model components successively processing undersampled and full-resolution images. The first component distinguishes between patches in the inner cloud area from patches at the cloud’s boundary region. For the cloud-ambiguous edge patches requiring further segmentation, the framework then delegates computation to a fine-grained model component. We apply the architecture to a cloud detection data set of complete Sentinel-2 multispectral images, approximately annotated for minimal false negatives in a land-use application. On this specific task and data, we achieve a 16% relative improvement in pixel accuracy over a CNN baseline based on patching.
Markku Luotamo, Sari Metsämäki, Arto Klami
IEEE Trans. Geosci. Remote. Sens.3
2020 Correcting Predictions for Approximate Bayesian Inference
abstract
Bayesian models quantify uncertainty and facilitate optimal decision-making in downstream applications. For most models, however, practitioners are forced to use approximate inference techniques that lead to sub-optimal decisions due to incorrect posterior predictive distributions. We present a novel approach that corrects for inaccuracies in posterior inference by altering the decision-making process. We train a separate model to make optimal decisions under the approximate posterior, combining interpretable Bayesian modeling with optimization of direct predictive accuracy in a principled fashion. The solution is generally applicable as a plug-in module for predictive decision-making for arbitrary probabilistic programs, irrespective of the posterior inference strategy. We demonstrate the approach empirically in several problems, confirming its potential.
Tomasz Kusmierczyk, Joseph Sakaya, Arto Klami
AAAI3
2020 Flexible Prior Elicitation via the Prior Predictive Distribution
abstract
The prior distribution for the unknown model parameters plays a crucial role in the process of statistical inference based on Bayesian methods. However, specifying suitable priors is often difficult even when detailed prior knowledge is available in principle. The challenge is to express quantitative information in the form of a probability distribution. Prior elicitation addresses this question by extracting subjective information from an expert and transforming it into a valid prior. Most existing methods, however, require information to be provided on the unobservable parameters, whose effect on the data generating process is often complicated and hard to understand. We propose an alternative approach that only requires knowledge about the observable outcomes - knowledge which is often much easier for experts to provide. Building upon a principled statistical framework, our approach utilizes the prior predictive distribution implied by the model to automatically transform experts judgements about plausible outcome values to suitable priors on the parameters. We also provide computational strategies to perform inference and guidelines to facilitate practical use.
Marcelo Hartmann, Georgi Agiashvili, Paul C. Bürkner, Arto Klami
UAI4
2020 Sensor Placement for Spatial Gaussian Processes with Integral Observations
abstract
Gaussian processes (GP) are a natural tool for estimating unknown functions, typically based on a collection of point-wise observations. Interestingly, the GP formalism can be used also with observations that are integrals of the unknown function along some known trajectories, which makes GPs a promising technique for inverse problems in a wide range of physical sensing problems. However, in many real world applications collecting data is laborious and time consuming. We provide tools for optimizing sensor locations for GPs using integral observations, extending both model-based and geometric strategies for GP sensor placement.We demonstrate the techniques in ultrasonic detection of fouling in closed pipes.
Krista Longi, Chang Rajani, Tom Sillanpää, Joni Mäkinen, Timo Rauhala, Ari Salmi, Edward Hæggström, Arto Klami
UAI8
2019 Low-Rank Approximations of Second-Order Document Representations
abstract
Document embeddings, created with methods ranging from simple heuristics to statistical and deep models, are widely applicable.Bagof-vectors models for documents include the mean and quadratic approaches (Torki, 2018).We present evidence that quadratic statistics alone, without the mean information, can offer superior accuracy, fast document comparison, and compact document representations.In matching news articles to their comment threads, low-rank representations of only 3-4 times the size of the mean vector give most accurate matching, and in standard sentence comparison tasks, results are state of the art despite faster computation.Similarity measures are discussed, and the Frobenius product implicit in the proposed method is contrasted to Wasserstein or Bures metric from the transportation theory.We also shortly demonstrate matching of unordered word lists to documents, to measure topicality or sentiment of documents.
Jarkko Lagus, Janne Sinkkonen, Arto Klami
CoNLL3
2019 Variational Bayesian Decision-making for Continuous Utilities
abstract
Bayesian decision theory outlines a rigorous framework for making optimal decisions based on maximizing expected utility over a model posterior. However, practitioners often do not have access to the full posterior and resort to approximate inference strategies. In such cases, taking the eventual decision-making task into account while performing the inference allows for calibrating the posterior approximation to maximize the utility. We present an automatic pipeline that co-opts continuous utilities into variational inference algorithms to account for decision-making. We provide practical strategies for approximating and maximizing the gain, and empirically demonstrate consistent improvement when calibrating approximations for specific utilities.
Tomasz Kusmierczyk, Joseph Sakaya, Arto Klami
NeurIPS3
2018 On Controlling the Size of Clusters in Probabilistic Clustering
abstract
Classical model-based partitional clustering algorithms, such ask-means or mixture of Gaussians, provide only loose and indirect control over the size of the resulting clusters. In this work, we present a family of probabilistic clustering models that can be steered towards clusters of desired size by providing a prior distribution over the possible sizes, allowing the analyst to fine-tune exploratory analysis or to produce clusters of suitable size for future down-stream processing.Our formulation supports arbitrary multimodal prior distributions, generalizing the previous work on clustering algorithms searching for clusters of equal size or algorithms designed for the microclustering task of finding small clusters. We provide practical methods for solving the problem, using integer programming for making the cluster assignments, and demonstrate that we can also automatically infer the number of clusters.
Aditya Jitta, Arto Klami
AAAI2
2018 Lambert Matrix Factorization
Arto Klami, Jarkko Lagus, Joseph Sakaya
ECML/PKDD (2)1
2018 Transfer-Learning Methods in Programming Course Outcome Prediction
abstract
The computing education research literature contains a wide variety of methods that can be used to identify students who are either at risk of failing their studies or who could benefit from additional challenges. Many of these are based on machine-learning models that learn to make predictions based on previously observed data. However, in educational contexts, differences between courses set huge challenges for the generalizability of these methods. For example, traditional machine-learning methods assume identical distribution in all data—in our terms, traditional machine-learning methods assume that all teaching contexts are alike. In practice, data collected from different courses can be very different as a variety of factors may change, including grading, materials, teaching approach, and the students. Transfer-learning methodologies have been created to address this challenge. They relax the strict assumption of identical distribution for training and test data. Some similarity between the contexts is still needed for efficient learning. In this work, we review the concept of transfer learning especially for the purpose of predicting the outcome of an introductory programming course and contrast the results with those from traditional machine-learning methods. The methods are evaluated using data collected in situ from two separate introductory programming courses. We empirically show that transfer-learning methods are able to improve the predictions, especially in cases with limited amount of training data, for example, when making early predictions for a new context. The difference in predictive power is, however, rather subtle, and traditional machine-learning models can be sufficiently accurate assuming the contexts are closely related and the features describing the student activity are carefully chosen to be insensitive to the fine differences.
Jarkko Lagus, Krista Longi, Arto Klami, Arto Hellas
ACM Trans. Comput. Educ.3
2017 Semi-supervised Convolutional Neural Networks for Identifying Wi-Fi Interference Sources
abstract
We present a convolutional neural network for identifying radio frequency devices from signal data, in order to detect possible interference sources for wireless local area networks. Collecting training data for this problem is particularly challenging due to a high number of possible interfering devices, difficulty in obtaining precise timings, and the need to measure the devices in varying conditions. To overcome this challenge we focus on semi-supervised learning, aiming to minimize the need for reliable training samples while utilizing larger amounts of unsupervised labels to improve the accuracy. In particular, we propose a novel structured extension of the pseudo-label technique to take advantage of temporal continuity in the data and show that already a few seconds of training data for each device is sufficient for highly accurate recognition.
Krista Longi, Teemu Pulkkinen, Arto Klami
ACML3
2017 Importance Sampled Stochastic Optimization for Variational Inference
Joseph Sakaya, Arto Klami
UAI2
2017 Partially hidden Markov models for privacy-preserving modeling of indoor trajectories
Aditya Jitta, Arto Klami
Neurocomputing2
2016 Typing Patterns and Authentication in Practical Programming Exams
abstract
In traditional programming courses, students have usually been at least partly graded using pen and paper exams. One of the problems related to such exams is that they only partially connect to the practice conducted within such courses. Testing students in a more practical environment has been constrained due to the limited resources that are needed, for example, for authentication.
Juho Leinonen 0001, Krista Longi, Arto Klami, Alireza Ahadi, Arto Vihavainen
ITiCSE3
2016 Automatic Inference of Programming Performance and Experience from Typing Patterns
abstract
Studies on retention and success in introductory programming course have suggested that previous programming experience contributes to students' course outcomes. If such background information could be automatically distilled from students' working process, additional guidance and support mechanisms could be provided even to those, who do not wish to disclose such information. In this study, we explore methods for automatically distinguishing novice programmers from more experienced programmers using fine-grained source code snapshot data. We approach the issue by partially replicating a previous study that used students' keystroke latencies as a proxy to introductory programming course outcomes, and follow this by an exploration of machine learning methods to separate those students with little to no previous programming experience from those with more experience. Our results confirm that students' keystroke latencies can be used as a metric for measuring course outcomes. At the same time, our results show that students programming experience can be identified to some extent from keystroke latency data, which means that such data has potential as a source of information for customizing the students' learning experience.
Juho Leinonen 0001, Krista Longi, Arto Klami, Arto Vihavainen
SIGCSE3
2016 Probabilistic Size-constrained Microclustering
Arto Klami, Aditya Jitta
UAI1
2016 Using regression makes extraction of shared variation in multiple datasets easy
Jussi Korpela, Andreas Henelius, Lauri Ahonen, Arto Klami, Kai Puolamäki
Data Min. Knowl. Discov.4
2015 Latent feature regression for multivariate count data
abstract
We consider the problem of regression on multivariate count data and present a Gibbs sampler for a latent feature regression model suitable for both under- and overdispersed response variables. The model learns count-valued latent features conditional on arbitrary covariates, modeling them as negative binomial variables, and maps them into the dependent count-valued observations using a Dirichlet-multinomial distribution. From another viewpoint, the model can be seen as a generalization of a specific topic model for scenarios where we are interested in generating the actual counts of observations and not just their relative frequencies and co-occurrences. The model is demonstrated on a smart traffic application where the task is to predict public transportation volume for unknown locations based on a characterization of the close-by services and venues.
Arto Klami, Abhishek Tripathi, Johannes Sirola, Lauri Väre, Frédéric Roulland
AISTATS1
2015 IntentStreams: Smart Parallel Search Streams for Branching Exploratory Search
abstract
The user's understanding of information needs and the information available in the data collection can evolve during an exploratory search session. Search systems tailored for well-defined narrow search tasks may be suboptimal for exploratory search where the user can sequentially refine the expressions of her information needs and explore alternative search directions. A major challenge for exploratory search systems design is how to support such behavior and expose the user to relevant yet novel information that can be difficult to discover by using conventional query formulation techniques. We introduce IntentStreams, a system for exploratory search that provides interactive query refinement mechanisms and parallel visualization of search streams. The system models each search stream via an intent model allowing rapid user feedback. The user interface allows swift initiation of alternative and parallel search streams by direct manipulation that does not require typing. A study with 13 participants shows that IntentStreams provides better support for branching behavior compared to a conventional search system.
Salvatore Andolina, Khalil Klouche, Jaakko Peltonen, Mohammad E. Hoque, Tuukka Ruotsalo, Diogo Cabral, Arto Klami, Dorota Glowacka, Patrik Floréen, Giulio Jacucci
IUI7
2015 Group Factor Analysis
abstract
Factor analysis (FA) provides linear factors that describe the relationships between individual variables of a data set. We extend this classical formulation into linear factors that describe the relationships between groups of variables, where each group represents either a set of related variables or a data set. The model also naturally extends canonical correlation analysis to more than two sets, in a way that is more flexible than previous extensions. Our solution is formulated as a variational inference of a latent variable model with structural sparsity, and it consists of two hierarchical levels: 1) the higher level models the relationships between the groups and 2) the lower models the observed variables given the higher level. We show that the resulting solution solves the group factor analysis (GFA) problem accurately, outperforming alternative FA-based solutions as well as more straightforward implementations of GFA. The method is demonstrated on two life science data sets, one on brain activation and the other on systems biology, illustrating its applicability to the analysis of different types of high-dimensional data sources.
Arto Klami, Seppo Virtanen, Eemeli Leppäaho, Samuel Kaski
IEEE Trans. Neural Networks Learn. Syst.1
2014 Polya-gamma augmentations for factor models
Arto Klami
ACML1
2014 Multi-task and multi-view learning of user state
Melih Kandemir, Akos Vetek, Mehmet Gönen, Arto Klami, Samuel Kaski
Neurocomputing4
2013 Bayesian Canonical correlation analysis
Arto Klami, Seppo Virtanen, Samuel Kaski
J. Mach. Learn. Res.1
2013 Bayesian object matching
Arto Klami
Mach. Learn.1
2012 Unsupervised Inference of Auditory Attention from Biosensors
Melih Kandemir, Arto Klami, Akos Vetek, Samuel Kaski
ECML/PKDD (2)2
2012 Factorized Multi-Modal Topic Model
Seppo Virtanen, Yangqing Jia, Arto Klami, Trevor Darrell
UAI3
2011 Bayesian CCA via Group Sparsity
Seppo Virtanen, Arto Klami, Samuel Kaski
ICML2
2011 Matching samples of multiple views
Abhishek Tripathi, Arto Klami, Matej Oresic, Samuel Kaski
Data Min. Knowl. Discov.2
2010 Variational Bayesian Mixture of Robust CCA Models
Jaakko Viinikanoja, Arto Klami, Samuel Kaski
ECML/PKDD (3)2
2010 Bayesian exponential family projections for coupled data sources
Arto Klami, Seppo Virtanen, Samuel Kaski
UAI1
2010 Infinite factorization of multiple non-parametric views
Simon Rogers, Arto Klami, Janne Sinkkonen, Mark A. Girolami, Samuel Kaski
Mach. Learn.2
2009 Fast dependent components for fMRI analysis
abstract
Canonical correlation analysis (CCA) can be used to find correlating projections of two datasets with co-occurring samples. Instead of correlation, we would typically want to find more general dependencies, measured by mutual information. Variants of CCA based on non-parametric estimation of mutual information have been proposed previously; they outperform traditional CCA for non-Gaussian data but require infeasible amounts of computation for already quite modest sample sizes. We introduce a novel variant that uses a semi parametric estimate leading to a considerably faster algorithm. We apply the method on searching for statistical dependencies between multi-sensory stimuli and functional magnetic resonance imaging (fMRI) of brain activity- in contrast to using regression on either of them.
Eerika Savia, Arto Klami, Samuel Kaski
ICASSP2
2009 Using dependencies to pair samples for multi-view learning
abstract
Several data analysis tools such as (kernel) canonical correlation analysis and various multi-view learning methods require paired observations in two data sets. We study the problem of inferring such pairing for data sets with no known one-to-one pairing. The pairing is found by an iterative algorithm that alternates between searching for feature representations that reveal statistical dependencies between the data sets, and finding the best pairs for the samples. The method is applied on pairing probe sets of two different microarray platforms.
Abhishek Tripathi, Arto Klami, Samuel Kaski
ICASSP2
2009 GaZIR: gaze-based zooming interface for image retrieval
abstract
We introduce GaZIR, a gaze-based interface for browsing and searching for images. The system computes on-line predictions of relevance of images based on implicit feedback, and when the user zooms in, the images predicted to be the most relevant are brought out. The key novelty is that the relevance feedback is inferred from implicit cues obtained in real-time from the gaze pattern, using an estimator learned during a separate training phase. The natural zooming interface can be connected to any content-based information retrieval engine operating on user feedback. We show with experiments on one engine that there is sufficient amount of information in the gaze patterns to make the estimated relevance feedback a viable choice to complement or even replace explicit feedback by pointing-and-clicking.
László Kozma 0002, Arto Klami, Samuel Kaski
ICMI2
2008 Simple integrative preprocessing preserves what is shared in data sources
abstract
BACKGROUND: Bioinformatics data analysis toolbox needs general-purpose, fast and easily interpretable preprocessing tools that perform data integration during exploratory data analysis. Our focus is on vector-valued data sources, each consisting of measurements of the same entity but on different variables, and on tasks where source-specific variation is considered noisy or not interesting. Principal components analysis of all sources combined together is an obvious choice if it is not important to distinguish between data source-specific and shared variation. Canonical Correlation Analysis (CCA) focuses on mutual dependencies and discards source-specific "noise" but it produces a separate set of components for each source. RESULTS: It turns out that components given by CCA can be combined easily to produce a linear and hence fast and easily interpretable feature extraction method. The method fuses together several sources, such that the properties they share are preserved. Source-specific variation is discarded as uninteresting. We give the details and implement them in a software tool. The method is demonstrated on gene expression measurements in three case studies: classification of cell cycle regulated genes in yeast, identification of differentially expressed genes in leukemia, and defining stress response in yeast. The software package is available at http://www.cis.hut.fi/projects/mi/software/drCCA/. CONCLUSION: We introduced a method for the task of data fusion for exploratory data analysis, when statistical dependencies between the sources and not within a source are interesting. The method uses canonical correlation analysis in a new way for dimensionality reduction, and inherits its good properties of being simple, fast, and easily interpretable as a linear projection.
Abhishek Tripathi, Arto Klami, Samuel Kaski
BMC Bioinform.2
2008 Probabilistic approach to detecting dependencies between data sets
Arto Klami, Samuel Kaski
Neurocomputing1
2007 Local dependent components
abstract
We introduce a mixture of probabilistic canonical correlation analyzers model for analyzing local correlations, or more generally mutual statistical dependencies, in cooccurring data pairs. The model extends the traditional canonical correlation analysis and its probabilistic interpretation in three main ways. First, a full Bayesian treatment enables analysis of small samples (large p, small n, a crucial problem in bioinformatics, for instance), and rigorous estimation of the degree of dependency and independency. Secondly, the mixture formulation generalizes the method from global linearity to the more reasonable assumption of different kinds of dependencies for different kinds of data. As a third novel extension the method decomposes the variation in the data into shared and data set-specific components. 1.
Arto Klami, Samuel Kaski
ICML1
2005 Non-parametric dependent components
abstract
Canonical correlation analysis (CCA) is equivalent to finding mutual information-maximizing projections for normally distributed data. We remove the restriction of normality by non-parametric estimation, and formulate the problem of finding dependent components with a connection to Bayes factors. The method is applied for characterizing yeast stress by finding what is in common in several different stress conditions.
Arto Klami, Samuel Kaski
ICASSP (5)1
2005 Discriminative clustering
Samuel Kaski, Janne Sinkkonen, Arto Klami
Neurocomputing3
2004 Improved learning of Riemannian metrics for exploratory analysis
Jaakko Peltonen, Arto Klami, Samuel Kaski
Neural Networks2
2002 Learning More Accurate Metrics for Self-Organizing Maps
Jaakko Peltonen, Arto Klami, Samuel Kaski
ICANN2