Kai Puolamäki

dblp:71/3034 · DBLP profile ↗
← Back
32ranked-venue papers in the field
6as first author
5since 2021 · last 2026
0000-0003-1819-1047ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 26 (3 first)Database Systems & Data Management · 3 (1 first)Other / Interdisciplinary · 2 (1 first)Information Retrieval & Web Search · 1 (1 first)
YearPublicationVenuePosition
2026 CODAC: Constraint-based Deep Active Clustering
abstract
Abstract Constraint-based Deep Active Clustering (CODAC) integrates actively selected pairwise constraints into deep representation learning to efficiently improve existing cluster structures, even under tight query budgets. CODAC encodes the constraint information into the embedding so that the learned representation can generalize to unconstrained data, leading to a rapid improvement of the clustering quality even on large datasets. CODAC makes minimal assumptions regarding the data and can be combined with a wide variety of deep clustering models. It does not require the number of clusters to be known a priori and is even effective if the initial estimate is badly misspecified. Across diverse image, text, and tabular datasets, CODAC consistently attains higher cluster quality with fewer queries than the previous state-of-the-art, and can substantially improve the clustering quality with just 100–200 queries compared to the deep clustering baselines.
Anri Patron, Sandra Gilhuber, Kai Puolamäki, Collin Leiber
Data Min. Knowl. Discov.3
2024 SLIPMAP: Fast and Robust Manifold Visualisation for Explainable AI
abstract
Abstract We propose a new supervised manifold visualisation method, slipmap, that finds local explanations for complex black-box supervised learning methods and creates a two-dimensional embedding of the data items such that data items with similar local explanations are embedded nearby. This work extends and improves our earlier algorithm and addresses its shortcomings: poor scalability, inability to make predictions, and a tendency to find patterns in noise. We present our visualisation problem and provide an efficient GPU-optimised library to solve it. We experimentally verify that slipmap is fast and robust to noise, provides explanations that are on the level or better than the other local explanation methods, and are usable in practice.
Anton Björklund, Lauri Seppäläinen, Kai Puolamäki
IDA (2)3
2022 SLISEMAP: Combining Supervised Dimensionality Reduction with Local Explanations
abstract
Abstract We introduce a Python library, called slisemap, that contains a supervised dimensionality reduction method that can be used for global explanation of black box regression or classification models. slisemap takes a data matrix and predictions from a black box model as input, and outputs a (typically) two-dimensional embedding, such that the black box model can be approximated, to a good fidelity, by the same interpretable white box model for points with similar embeddings. The library includes basic visualisation tools and extensive documentation, making it easy to get started and obtain useful insights. The slisemap library is published on GitHub and PyPI under an open source license.
Anton Björklund, Jarmo Mäkelä, Kai Puolamäki
ECML/PKDD (6)3
2022 Robust regression via error tolerance
abstract
Abstract Real-world datasets are often characterised by outliers; data items that do not follow the same structure as the rest of the data. These outliers might negatively influence modelling of the data. In data analysis it is, therefore, important to consider methods that are robust to outliers. In this paper we develop a robust regression method that finds the largest subset of data items that can be approximated using a sparse linear model to a given precision. We show that this can yield the best possible robustness to outliers. However, this problem is NP-hard and to solve it we present an efficient approximation algorithm, termed SLISE. Our method extends existing state-of-the-art robust regression methods, especially in terms of speed on high-dimensional datasets. We demonstrate our method by applying it to both synthetic and real-world regression problems.
Anton Björklund, Andreas Henelius, Emilia Oikarinen, Kimmo Kallonen, Kai Puolamäki
Data Min. Knowl. Discov.5
2021 Detecting virtual concept drift of regressors without ground truth values
abstract
Abstract Regression analysis is a standard supervised machine learning method used to model an outcome variable in terms of a set of predictor variables. In most real-world applications the true value of the outcome variable we want to predict is unknown outside the training data, i.e., the ground truth is unknown. Phenomena such as overfitting and concept drift make it difficult to directly observe when the estimate from a model potentially is wrong. In this paper we present an efficient framework for estimating the generalization error of regression functions, applicable to any family of regression functions when the ground truth is unknown. We present a theoretical derivation of the framework and empirically evaluate its strengths and limitations. We find that it performs robustly and is useful for detecting concept drift in datasets in several real-world domains.
Emilia Oikarinen, Henri Tiittanen, Andreas Henelius, Kai Puolamäki
Data Min. Knowl. Discov.4
2020 Interactive visual data exploration with subjective feedback: an information-theoretic approach
abstract
Visual exploration of high-dimensional real-valued datasets is a fundamental task in exploratory data analysis (EDA). Existing methods use predefined criteria to choose the representation of data. There is a lack of methods that (i) elicit from the user what she has learned from the data and (ii) show patterns that she does not know yet. We construct a theoretical model where identified patterns can be input as knowledge to the system. The knowledge syntax here is intuitive, such as "this set of points forms a cluster", and requires no knowledge of maths. This background knowledge is used to find a Maximum Entropy distribution of the data, after which the system provides the user data projections in which the data and the Maximum Entropy distribution differ the most, hence showing the user aspects of the data that are maximally informative given the user's current knowledge. We provide an open source EDA system with tailored interactive visualizations to demonstrate these concepts. We study the performance of the system and present use cases on both synthetic and real data. We find that the model and the prototype system allow the user to learn information efficiently from various data sources and the system works sufficiently fast in practice. We conclude that the information theoretic approach to exploratory data analysis where patterns observed by a user are formalized as constraints provides a principled, intuitive, and efficient basis for constructing an EDA system.
Kai Puolamäki, Emilia Oikarinen, Bo Kang, Jefrey Lijffijt, Tijl De Bie
Data Min. Knowl. Discov.1
2020 A Constrained Randomization Approach to Interactive Visual Data Exploration with Subjective Feedback
abstract
Data visualization and iterative/interactive data mining are growing rapidly in attention, both in research as well as in industry. However, while there are a plethora of advanced data mining methods and lots of works in the field of visualization, integrated methods that combine advanced visualization and/or interaction with data mining techniques in a principled way are rare. We present a framework based on constrained randomization which lets users explore high-dimensional data via `subjectively informative' two-dimensional data visualizations. The user is presented with `interesting' projections, allowing users to express their observations using visual interactions that update a background model representing the user's belief state. This background model is then considered by a projection-finding algorithm employing data randomization to compute a new `interesting' projection. By providing users with information that contrasts with the background model, we maximize the chance that the user encounters striking new information present in the data. This process can be iterated until the user runs out of time or until the difference between the randomized and the real data is insignificant. We present two case studies, one controlled study on synthetic data and another on census data, using the proof-of-concept tool SIDE that demonstrates the presented framework.
Bo Kang, Kai Puolamäki, Jefrey Lijffijt, Tijl De Bie
IEEE Trans. Knowl. Data Eng.2
2019 Significance of Patterns in Data Visualisations
abstract
In this paper we consider the following important problem: when we explore data visually and observe patterns, how can we determine their statistical significance? Patterns observed in exploratory analysis are traditionally met with scepticism, since the hypotheses are formulated while viewing the data, rather than before doing so. In contrast to this belief, we show that it is, in fact, possible to evaluate the significance of patterns also during exploratory analysis, and that the knowledge of the analyst can be leveraged to improve statistical power by reducing the amount of simultaneous comparisons. We develop a principled framework for determining the statistical significance of visually observed patterns. Furthermore, we show how the significance of visual patterns observed during iterative data exploration can be determined. We perform an empirical investigation on real and synthetic tabular data and time series, using different test statistics and methods for generating surrogate data. We conclude that the proposed framework allows determining the significance of visual patterns during exploratory analysis.
Rafael Savvides, Andreas Henelius, Emilia Oikarinen, Kai Puolamäki
KDD4
2018 Subjectively Interesting Subgroup Discovery on Real-Valued Targets
abstract
Deriving insights from high-dimensional data is one of the core problems in data mining. The difficulty mainly stems from the large number of variable combinations to potentially consider. Hence, an obvious question is whether we can automate the search for interesting patterns. Here, we consider the setting where a user wants to learn as efficiently as possible about real-valued attributes. We introduce a method to find subgroups in the data that are maximally informative (in the Information Theoretic sense) with respect to one or more real-valued target attributes. The succinct subgroup descriptions are in terms of arbitrarily-typed description attributes. The approach is based on the Subjective Interestingness framework FORSIED to use prior knowledge when mining most informative patterns.
Jefrey Lijffijt, Bo Kang, Wouter Duivesteijn, Kai Puolamäki, Emilia Oikarinen, Tijl De Bie
ICDE4
2018 Interactive Visual Data Exploration with Subjective Feedback: An Information-Theoretic Approach
Kai Puolamäki, Emilia Oikarinen, Bo Kang, Jefrey Lijffijt, Tijl De Bie
ICDE1
2018 Tiler: Software for Human-Guided Data Exploration
Andreas Henelius, Emilia Oikarinen, Kai Puolamäki
ECML/PKDD (3)3
2017 Multivariate Confidence Intervals
abstract
Confidence intervals are a popular way to visualize and analyze data distributions. Unlike p-values, they can convey information both about statistical significance as well as effect size. However, very little work exists on applying confidence intervals to multivariate data. In this paper we define confidence intervals for multivariate data that extend the one-dimensional definition in a natural way. In our definition every variable is associated with its own confidence interval as usual, but a data vector can be outside of a few of these, and still be considered to be within the confidence area. We analyze the problem and show that the resulting confidence areas retain the good qualities of their one-dimensional counterparts: they are informative and easy to interpret. Furthermore, we show that the problem of finding multivariate confidence intervals is hard, but provide efficient approximate algorithms to solve the problem.
Jussi Korpela, Emilia Oikarinen, Kai Puolamäki, Antti Ukkonen
SDM3
2016 Semigeometric Tiling of Event Sequences
Andreas Henelius, Isak Karlsson, Panagiotis Papapetrou, Antti Ukkonen, Kai Puolamäki
ECML/PKDD (1)5
2016 A Tool for Subjective and Interactive Visual Data Exploration
Bo Kang, Kai Puolamäki, Jefrey Lijffijt, Tijl De Bie
ECML/PKDD (3)2
2016 Interactive Visual Data Exploration with Subjective Feedback
Kai Puolamäki, Bo Kang, Jefrey Lijffijt, Tijl De Bie
ECML/PKDD (2)1
2016 Using regression makes extraction of shared variation in multiple datasets easy
Jussi Korpela, Andreas Henelius, Lauri Ahonen, Arto Klami, Kai Puolamäki
Data Min. Knowl. Discov.5
2015 Size matters: choosing the most informative set of window lengths for mining patterns in event sequences
Jefrey Lijffijt, Panagiotis Papapetrou, Kai Puolamäki
Data Min. Knowl. Discov.3
2014 A peek into the black box: exploring classifiers by randomization
Andreas Henelius, Kai Puolamäki, Henrik Boström, Lars Asker, Panagiotis Papapetrou
Data Min. Knowl. Discov.2
2014 Confidence bands for time series data
Jussi Korpela, Kai Puolamäki, Aristides Gionis
Data Min. Knowl. Discov.2
2014 A statistical significance testing approach to mining the most informative set of patterns
Jefrey Lijffijt, Panagiotis Papapetrou, Kai Puolamäki
Data Min. Knowl. Discov.3
2013 Explaining Interval Sequences by Randomization
Andreas Henelius, Jussi Korpela, Kai Puolamäki
ECML/PKDD (1)3
2012 Size Matters: Finding the Most Informative Set of Window Lengths
Jefrey Lijffijt, Panagiotis Papapetrou, Kai Puolamäki
ECML/PKDD (2)3
2011 Analyzing Word Frequencies in Large Text Corpora Using Inter-arrival Times and Bootstrapping
Jefrey Lijffijt, Panagiotis Papapetrou, Kai Puolamäki, Heikki Mannila
ECML/PKDD (2)3
2009 Bayesian Solutions to the Label Switching Problem
Kai Puolamäki, Samuel Kaski
IDA1
2009 Two-Way Grouping by One-Way Topic Models
Eerika Savia, Kai Puolamäki, Samuel Kaski
IDA2
2009 Tell me something I don't know: randomization strategies for iterative data mining
abstract
There is a wide variety of data mining methods available, and it is generally useful in exploratory data analysis to use many different methods for the same dataset. This, however, leads to the problem of whether the results found by one method are a reflection of the phenomenon shown by the results of another method, or whether the results depict in some sense unrelated properties of the data. For example, using clustering can give indication of a clear cluster structure, and computing correlations between variables can show that there are many significant correlations in the data. However, it can be the case that the correlations are actually determined by the cluster structure.
Sami Hanhijärvi, Markus Ojala, Niko Vuokko, Kai Puolamäki, Nikolaj Tatti, Heikki Mannila
KDD4
2009 Randomization Techniques for Graphs
abstract
Mining graph data is an active research area. Several data mining methods and algorithms have been proposed to identify structures from graphs; still, the evaluation of those results is lacking. Within the framework of statistical hypothesis testing, we focus in this paper on randomization techniques for unweighted undirected graphs. Randomization is an important approach to assess the statistical significance of data mining results. Given an input graph, our randomization method will sample data from the class of graphs that share certain structural properties with the input graph. Here we describe three alternative algorithms based on local edge swapping and Metropolis sampling. We test our framework with various graph data sets and mining algorithms for two applications, namely graph clustering and frequent subgraph mining. 1
Sami Hanhijärvi, Gemma C. Garriga, Kai Puolamäki
SDM3
2009 A randomized approximation algorithm for computing bucket orders
Antti Ukkonen, Kai Puolamäki, Aristides Gionis, Heikki Mannila
Inf. Process. Lett.2
2008 An approximation ratio for biclustering
Kai Puolamäki, Sami Hanhijärvi, Gemma C. Garriga
Inf. Process. Lett.1
2006 Algorithms for discovering bucket orders from data
abstract
Ordering and ranking items of different types are important tasks in various applications, such as query processing and scientific data mining. A total order for the items can be misleading, since there are groups of items that have practically equal ranks.We consider bucket orders, i.e., total orders with ties. They can be used to capture the essential order information without overfitting the data: they form a useful concept class between total orders and arbitrary partial orders. We address the question of finding a bucket order for a set of items, given pairwise precedence information between the items. We also discuss methods for computing the pairwise precedence data.We describe simple and efficient algorithms for finding good bucket orders. Several of the algorithms have a provable approximation guarantee, and they scale well to large datasets. We provide experimental results on artificial and a real data that show the usefulness of bucket orders and demonstrate the accuracy and efficiency of the algorithms.
Aristides Gionis, Heikki Mannila, Kai Puolamäki, Antti Ukkonen
KDD3
2005 On Discriminative Joint Density Modeling
Jarkko Salojärvi, Kai Puolamäki, Samuel Kaski
ECML2
2005 Combining eye movements and collaborative filtering for proactive information retrieval
abstract
We study a new task, proactive information retrieval by combining implicit relevance feedback and collaborative filtering. We have constructed a controlled experimental setting, a prototype application, in which the users try to find interesting scientific articles by browsing their titles. Implicit feedback is inferred from eye movement signals, with discriminative hidden Markov models estimated from existing data in which explicit relevance feedback is available. Collaborative filtering is carried out using the User Rating Profile model, a state-of-the-art probabilistic latent variable model, computed using Markov Chain Monte Carlo techniques. For new document titles the prediction accuracy with eye movements, collaborative filtering, and their combination was significantly better than by chance. The best prediction accuracy still leaves room for improvement but shows that proactive information retrieval and combination of many sources of relevance feedback is feasible.
Kai Puolamäki, Jarkko Salojärvi, Eerika Savia, Jaana Simola, Samuel Kaski
SIGIR1