Jaakko Peltonen

dblp:80/791 · DBLP profile ↗
← Back
72ranked-venue papers
15as first author
18since 2021 · last 2026
0000-0003-3485-8585ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 39 · 9 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 14 · 4 since 2021Human-computer interaction and ubiquitous computing · 13 · 3 first-author · 4 since 2021Theory of computation · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 The many facets of fairness in recommender systems: Consumers, providers and items
abstract
Autonomous decision-making systems, particularly recommender systems, have received increasing attention concerning fairness, i.e., if all stakeholders affected by such a system are treated equally as a result of the recommendations. Existing approaches primarily focus on fairness between two stakeholders – consumers and providers or consumers and items – treating providers and items as the same entity. However, we argue for the treatment of providers and items as distinct stakeholders to offer more comprehensive models of fairness in recommender systems. To this end, we propose a fairness-aware recommender system, CIPFRS, designed to optimize fairness across all three key stakeholders: consumers, providers, and items. We examine consumer fairness regarding their level of interaction with the system; high and low-activity users should be treated equally. Further, all providers should have an equal opportunity for their products to be recommended. Finally, we propose an approach to implement item fairness in each provider’s inventory. We report an extensive evaluation of the proposed solution through three datasets, demonstrating that considering all three stakeholders yields improved recommendations while minimizing bias.
Reza Shafiloo, Maria Stratigi, Jaakko Peltonen, Thomas Olsson 0002, Kostas Stefanidis
Inf. Syst.3
2026 Multisided fairness under limited item availability in recommender systems
abstract
Recommender systems often aim to serve multiple stakeholders, such as consumers and providers, each with distinct fairness expectations. While fairness-aware recommendation has gained increasing attention, most existing methods assume unlimited item availability. However, real-world scenarios often involve limited supply, where only a small number of item copies can be allocated. This creates new challenges in balancing fair exposure, equitable access, and relevance. In this paper, we propose a multisided fairness-aware recommendation framework designed for settings with limited item availability. Our approach explicitly models fairness both across stakeholders, ensuring consumers and providers are treated equitably compared to their peers, and within consumer–provider relationships, ensuring stakeholders treat their counterparts fairly. We formalize these as inter- and intra-stakeholder fairness and introduce evaluation metrics that measure treatment consistency under supply constraints. To address the allocation challenge, we develop a novel algorithm that assigns limited items while jointly optimizing for fairness and relevance. We evaluate our method on real-world datasets from Amazon and Goodreads, showing that it mitigates bias toward highly active users and dominant providers. Compared to conventional recommendation algorithms, our approach reduces fairness disparities by up to 80 %, underscoring the importance of fairness-aware design in real-world, resource-constrained recommendation scenarios.
Reza Shafiloo, Maria Stratigi, Jaakko Peltonen, Kostas Stefanidis
Inf. Sci.3
2025 Constrained Non-negative Matrix Factorization for Guided Topic Modeling of Minority Topics
abstract
Topic models often fail to capture lowprevalence, domain-critical themes-so-called minority topics-such as mental health themes in online comments.While some existing methods can incorporate domain knowledge such as expected topical content, methods allowing guidance may require overly detailed expected topics, hindering the discovery of topic divisions and variation.We propose a topic modeling solution via a specially constrained NMF.We incorporate a seed word list characterizing minority content of interest, but we do not require experts to pre-specify their division across minority topics.Through prevalence constraints on minority topics and seed word content across topics, we learn distinct data-driven minority topics as well as majority topics.The constrained NMF is fitted via Karush-Kuhn-Tucker (KKT) conditions with multiplicative updates.We outperform several baselines on synthetic data in terms of topic purity, normalized mutual information, and also evaluate topic quality using Jensen-Shannon divergence (JSD).We conduct a case study on YouTube vlog comments, analyzing viewer discussion of mental health content; our model successfully identifies and reveals this domainrelevant minority content.
Seyedeh Fatemeh Ebrahimi, Jaakko Peltonen
EMNLP2
2025 Nonnegative Matrix Factorization for Joint Clustering and Topic Modeling with Minority Topics
Seyedeh Fatemeh Ebrahimi, Jaakko Peltonen
ICONIP (1)2
2024 TraQuLA: Transparent Question Answering Over RDF Through Linguistic Analysis
Elizaveta Zimina, Kalervo Järvelin, Jaakko Peltonen, Aarne Ranta, Jyrki Nummenmaa
ICWE3
2023 Polarized Pills vs. Gaming Thrills: Empirical Exploration of r/TheRedPill and r/TheBluePill Users in r/gaming
Giacomo Lauritano, Valeria Marina Borodi, Chien Lu, Jaakko Peltonen
DiGRA4
2023 Human-Environment Relationships in Alba: A Typological Analysis of Player Engagement in Steam Reviews
Chien Lu, Giacomo Lauritano, Timo Nummenmaa, Jaakko Peltonen
DiGRA4
2023 Fair Neighbor Embedding
abstract
We consider fairness in dimensionality reduction. Nonlinear dimensionality reduction yields low dimensional representations that let users visualize and explore high-dimensional data. However, traditional dimensionality reduction may yield biased visualizations overemphasizing relationships of societal phenomena to sensitive attributes or protected groups. We introduce a framework of fair neighbor embedding, the Fair Neighbor Retrieval Visualizer, which formulates fair nonlinear dimensionality reduction as an information retrieval task whose performance and fairness are quantified by information retrieval criteria. The method optimizes low-dimensional embeddings that preserve high-dimensional data neighborhoods without yielding biased association of such neighborhoods to protected groups. In experiments the method yields fair visualizations outperforming previous methods.
Jaakko Peltonen, Timo Nummenmaa, Jyrki Nummenmaa
ICML1
2023 Faster Edge-Path Bundling through Graph Spanners
abstract
Abstract Edge‐Path bundling is a recent edge bundling approach that does not incur ambiguities caused by bundling disconnected edges together. Although the approach produces less ambiguous bundlings, it suffers from high computational cost. In this paper, we present a new Edge‐Path bundling approach that increases the computational speed of the algorithm without reducing the quality of the bundling. First, we demonstrate that biconnected components can be processed separately in an Edge‐Path bundling of a graph without changing the result. Then, we present a new edge bundling algorithm that is based on observing and exploiting a strong relationship between Edge‐Path bundling and graph spanners. Although the worst case complexity of the approach is the same as of the original Edge‐Path bundling algorithm, we conduct experiments to demonstrate that the new approach is 5–256 times faster than Edge‐Path bundling depending on the dataset, which brings its practical running time more in line with traditional edge bundling algorithms.
Markus Wallinger, Daniel Archambault, David Auber, Martin Nöllenburg, Jaakko Peltonen
Comput. Graph. Forum5
2022 CryptoKitties vs. Axie Infinity: Computational Analysis of NFT Game Reddit Discussions
Chien Lu, Giacomo Lauritano, Jaakko Peltonen
ArtsIT3
2022 Supervised dimensionality reduction technique accounting for soft classes
abstract
Exploratory visual analysis of multidimensional labeled data is challenging.Multidimensional Projections for labeled data attempt to separate classes while preserving neighborhoods.In this work, we consider the case where instances are assigned multiple labels with probabilities or weights: for example, the output of a probabilistic classifier, fuzzy membership functions in fuzzy logic, or the share of votes for each candidate in an election.We propose a new technique to better preserve neighborhoods of such data.Our experiments show improved qualitative results compared to unsupervised, and existing dimensionality reduction techniques.* The work of SM and SL has been realized with the participation of INES.2S.The work of DD has been
Sorina Mustatea, Michaël Aupetit 0001, Jaakko Peltonen, Sylvain Lespinats, Denys Dutykh
ESANN3
2022 Gaussian Copula Embeddings
abstract
Learning latent vector representations via embedding models has been shown promising in machine learning. However, most of the embedding models are still limited to a single type of observation data. We propose a Gaussian copula embedding model to learn latent vector representations of items in a heterogeneous data setting. The proposed model can effectively incorporate different types of observed data and, at the same time, yield robust embeddings. We demonstrate the proposed model can effectively learn in many different scenarios, outperforming competing models in modeling quality and task performance.
Chien Lu, Jaakko Peltonen
NeurIPS2
2022 Nonparametric exponential family graph embeddings for multiple representation learning
abstract
In graph data, each node often serves multiple functionalities. However, most graph embedding models assume that each node can only possess one representation. We address this issue by proposing a nonparametric graph embedding model. The model allows each node to learn multiple representations where they are needed to represent the complexity of random walks in the graph. It extends the Exponential family graph embedding model with two nonparametric prior settings, the Dirichlet process and the uniform process. The model combines the ability of Exponential family graph embedding to take the number of occurrences of context nodes into account with nonparametric priors giving it the flexibility to learn more than one latent representation for each node. The learned embeddings outperform other state of the art approaches in link prediction and node classification tasks.
Chien Lu, Jaakko Peltonen, Timo Nummenmaa, Jyrki Nummenmaa
UAI2
2022 Using parsed and annotated corpora to analyze parliamentarians' talk in Finland
abstract
Abstract We present a search system for grammatically analyzed corpora of Finnish parliamentary records and interviews with former parliamentarians, annotated with metadata of talk structure and involved parliamentarians, and discuss their use through carefully chosen digital humanities case studies. We first introduce the construction, contents, and principles of use of the corpora. Then we discuss the application of the search system and the corpora to study how politicians talk about power, how ideological terms are used in political speech, and how to identify narratives in the data. All case studies stem from questions in the humanities and the social sciences, but rely on the grammatically parsed corpora in both identifying and quantifying passages of interest. Finally, the paper discusses the role of natural language processing methods for questions in the (digital) humanities. It makes the claim that a digital humanities inquiry of parliamentary speech and interviews with politicians cannot only rely on computational humanities modeling, but needs to accommodate a range of perspectives starting with simple searches, quantitative exploration, and ending with modeling. Furthermore, the digital humanities need a more thorough discussion about how the utilization of tools from information science and technologies alter the research questions posed in the humanities.
Mykola Andrushchenko, Kirsi Sandberg, Risto Turunen, Jani Marjanen, Mari Hatavara, Jussi Kurunmäki, Timo Nummenmaa, Matti Hyvärinen, Kari Teräs, Jaakko Peltonen, Jyrki Nummenmaa
J. Assoc. Inf. Sci. Technol.10
2022 Multicriteria Optimization for Dynamic Demers Cartograms
abstract
Cartograms are popular for visualizing numerical data for administrative regions in thematic maps. When there are multiple data values per region (over time or from different datasets) shown as animated or juxtaposed cartograms, preserving the viewer's mental map in terms of stability between multiple cartograms is another important criterion alongside traditional cartogram criteria such as maintaining adjacencies. We present a method to compute stable stable Demers cartograms, where each region is shown as a square scaled proportionally to the given numerical data and similar data yield similar cartograms. We enforce orthogonal separation constraints using linear programming, and measure quality in terms of keeping adjacent regions close (cartogram quality) and using similar positions for a region between the different data values (stability). Our method guarantees the ability to connect most lost adjacencies with minimal-length planar orthogonal polylines. Experiments show that our method yields good quality and stability on multiple quality criteria.
Soeren Terziadis, Max Sondag, Wouter Meulemans, Stephen G. Kobourov, Jaakko Peltonen, Martin Nöllenburg
IEEE Trans. Vis. Comput. Graph.5
2022 Edge-Path Bundling: A Less Ambiguous Edge Bundling Approach
abstract
Edge bundling techniques cluster edges with similar attributes (i.e. similarity in direction and proximity) together to reduce the visual clutter. All edge bundling techniques to date implicitly or explicitly cluster groups of individual edges, or parts of them, together based on these attributes. These clusters can result in ambiguous connections that do not exist in the data. Confluent drawings of networks do not have these ambiguities, but require the layout to be computed as part of the bundling process. We devise a new bundling method, Edge-Path bundling, to simplify edge clutter while greatly reducing ambiguities compared to previous bundling techniques. Edge-Path bundling takes a layout as input and clusters each edge along a weighted, shortest path to limit its deviation from a straight line. Edge-Path bundling does not incur independent edge ambiguities typically seen in all edge bundling methods, and the level of bundling can be tuned through shortest path distances, Euclidean distances, and combinations of the two. Also, directed edge bundling naturally emerges from the model. Through metric evaluations, we demonstrate the advantages of Edge-Path bundling over other techniques.
Markus Wallinger, Daniel Archambault, David Auber, Martin Nöllenburg, Jaakko Peltonen
IEEE Trans. Vis. Comput. Graph.5
2021 Cross-structural Factor-topic Model: Document Analysis with Sophisticated Covariates
abstract
Modern text data is increasingly gathered in situations where it is paired with a high-dimensional collection of covariates: then both the text, the covariates, and their relationships are of interest to analyze. Despite the growing amount of such data, current topic models are unable to take into account large amounts of covariates successfully: they fail to model structure among covariates and distort findings of both text and covariates. This paper presents a solution: a novel factor-topic model that enables researchers to analyze latent structure in both text and sophisticated document-level covariates collectively. The key innovation is that besides learning the underlying topical structure, the model also learns the underlying factorial structure from the covariates and the interactions between the two structures. A set of tailored variational inference algorithms for efficient computation are provided. Experiments on three different datasets show the model outperforms comparable topic models in the ability to predict held-out document content. Two case studies focusing on Finnish parliamentary election candidates and game players on Steam demonstrate the model discovers semantically meaningful topics, factors, and their interactions. The model both outperforms state-of-the-art models in predictive accuracy and offers new factor-topic insights beyond other topic models.
Chien Lu, Jaakko Peltonen, Timo Nummenmaa, Jyrki Nummenmaa, Kalervo Jäarvelin
ACML2
2021 Directing and Combining Multiple Queries for Exploratory Search by Visual Interactive Intent Modeling
Jonathan Strahl, Jaakko Peltonen, Patrik Floréen
INTERACT (3)2
2020 Enhancing Nearest Neighbor Based Entropy Estimator for High Dimensional Distributions via Bootstrapping Local Ellipsoid
abstract
An ellipsoid-based, improved kNN entropy estimator based on random samples of distribution for high dimensionality is developed. We argue that the inaccuracy of the classical kNN estimator in high dimensional spaces results from the local uniformity assumption and the proposed method mitigates the local uniformity assumption by two crucial extensions, a local ellipsoid-based volume correction and a correction acceptance testing procedure. Relevant theoretical contributions are provided and several experiments from simple to complicated cases have shown that the proposed estimator can effectively reduce the bias especially in high dimensionalities, outperforming current state of the art alternative estimators.
Chien Lu, Jaakko Peltonen
AAAI2
2020 Scalable Probabilistic Matrix Factorization with Graph-Based Priors
abstract
In matrix factorization, available graph side-information may not be well suited for the matrix completion problem, having edges that disagree with the latent-feature relations learnt from the incomplete data matrix. We show that removing these contested edges improves prediction accuracy and scalability. We identify the contested edges through a highly-efficient graphical lasso approximation. The identification and removal of contested edges adds no computational complexity to state-of-the-art graph-regularized matrix factorization, remaining linear with respect to the number of non-zeros. Computational load even decreases proportional to the number of edges removed. Formulating a probabilistic generative model and using expectation maximization to extend graph-regularised alternating least squares (GRALS) guarantees convergence. Rich simulated experiments illustrate the desired properties of the resulting algorithm. On real data experiments we demonstrate improved prediction accuracy with fewer graph edges (empirical evidence that graph side-information is often inaccurate). A 300 thousand dimensional graph with three million edges (Yahoo music side-information) can be analyzed in under ten minutes on a standard laptop computer demonstrating the efficiency of our graph update.
Jonathan Strahl, Jaakko Peltonen, Hiroshi Mamitsuka, Samuel Kaski
AAAI2
2020 The World Is Your Playground: A Bibliometric and Text Mining Analysis of Location-Based Game Research
Chien Lu, Elina Koskinen, Dale Leorke, Timo Nummenmaa, Jaakko Peltonen
ArtsIT5
2020 Probabilistic Dynamic Non-negative Group Factor Model for Multi-source Text Mining
abstract
Nonnegative matrix factorization (NMF) is a popular approach to model data, however, most models are unable to flexibly take into account multiple matrices across sources and time or apply only to integer-valued data. We introduce a probabilistic, Gaussian Process-based, more inclusive NMF-based model which jointly analyzes nonnegative data such as text data word content from multiple sources in a temporal dynamic manner. The model collectively models observed matrix data, source-wise latent variables, and their dependencies and temporal evolution with a full-fledged hierarchical approach including flexible nonparametric temporal dynamics. Experiments on simulated data and real data show the model out-performs, comparable models. A case study on social media and news demonstrates the model discovers semantically meaningful topical factors and their evolution
Chien Lu, Jaakko Peltonen, Jyrki Nummenmaa, Kalervo Järvelin
CIKM2
2020 Patches and Player Community Perceptions: Analysis of No Man's Sky Steam Reviews
Chien Lu, Xiaozhou Li 0002, Timo Nummenmaa, Zheying Zhang, Jaakko Peltonen
DiGRA5
2020 Steering Distortions to Preserve Classes and Neighbors in Supervised Dimensionality Reduction
abstract
Nonlinear dimensionality reduction of high-dimensional data is challenging as the low-dimensional embedding will necessarily contain distortions, and it can be hard to determine which distortions are the most important to avoid. When annotation of data into known relevant classes is available, it can be used to guide the embedding to avoid distortions that worsen class separation. The supervised mapping method introduced in the present paper, called ClassNeRV, proposes an original stress function that takes class annotation into account and evaluates embedding quality both in terms of false neighbors and missed neighbors. ClassNeRV shares the theoretical framework of a family of methods descended from Stochastic Neighbor Embedding (SNE). Our approach has a key advantage over previous ones: in the literature supervised methods often emphasize class separation at the price of distorting the data neighbors' structure; conversely, unsupervised methods provide better preservation of structure at the price of often mixing classes. Experiments show that ClassNeRV can preserve both neighbor structure and class separation, outperforming nine state of the art alternatives.
Benoît Colange, Jaakko Peltonen, Michaël Aupetit 0001, Denys Dutykh, Sylvain Lespinats
NeurIPS2
2019 Game postmortems vs. developer Reddit AMAs: computational analysis of developer communication
abstract
Postmortems and Reddit Ask Me Anything (AMA) threads represent communications of game developers through two different channels about their game development experiences, culture, processes, and practices. We carry out a quantitative text mining based comprehensive analysis of online available postmortems and AMA threads from game developers over multiple years. We find and analyze underlying topics from the postmortems and AMAs as well as their variation among the data sources and over time. The analysis is done based on structural topic modeling, a probabilistic modeling technique for text mining. The extracted topics reveal differing and common interests as well as their evolution of prevalence over time in the two text sources. We have found that postmortems put more emphasis on detail-oriented development aspects as well as technically-oriented game design problems whereas AMAs feature a wider variety of discussion topics that are related to a more general game development process, game-play and game-play experience related game design. The prevalences of the topics also evolve differently over time in postmortems versus AMAs.
Chien Lu, Jaakko Peltonen, Timo Nummenmaa
FDG2
2019 Computing Stable Demers Cartograms
Soeren Terziadis, Max Sondag, Wouter Meulemans, Markus Chimani, Stephen G. Kobourov, Jaakko Peltonen, Martin Nöllenburg
GD6
2018 Author Tree-Structured Hierarchical Dirichlet Process
Md. Hijbul Alam, Jaakko Peltonen, Jyrki Nummenmaa, Kalervo Järvelin
DS2
2018 Rethinking Summarization and Storytelling for Modern Social Multimedia
Stevan Rudinac, Tat-Seng Chua, Nicolás E. Díaz Ferreyra, Gerald Friedland, Tatjana Gornostaja, Benoit Huet, Rianne Kaptein, Krister Lindén, Marie-Francine Moens, Jaakko Peltonen, Miriam Redi, Markus Schedl, David A. Shamma, Alan F. Smeaton, Lexing Xie
MMM (1)10
2018 Interactive Intent Modeling for Exploratory Search
abstract
Exploratory search requires the system to assist the user in comprehending the information space and expressing evolving search intents for iterative exploration and retrieval of information. We introduce interactive intent modeling, a technique that models a user’s evolving search intents and visualizes them as keywords for interaction. The user can provide feedback on the keywords, from which the system learns and visualizes an improved intent estimate and retrieves information. We report experiments comparing variants of a system implementing interactive intent modeling to a control system. Data comprising search logs, interaction logs, essay answers, and questionnaires indicate significant improvements in task performance, information retrieval performance over the session, information comprehension performance, and user experience. The improvements in retrieval effectiveness can be attributed to the intent modeling and the effect on users’ task performance, breadth of information comprehension, and user experience are shown to be dependent on a richer visualization. Our results demonstrate the utility of combining interactive modeling of search intentions with interactive visualization of the models that can benefit both directing the exploratory search process and making sense of the information space. Our findings can help design personalized systems that support exploratory information seeking and discovery of novel information.
Tuukka Ruotsalo, Jaakko Peltonen, Manuel J. A. Eugster, Dorota Glowacka, Patrik Floréen, Petri Myllymäki, Giulio Jacucci, Samuel Kaski
ACM Trans. Inf. Syst.2
2017 Topic-Relevance Map: Visualization for Improving Search Result Comprehension
abstract
We introduce topic-relevance map, an interactive search result visualization that assists rapid information comprehension across a large ranked set of results. The topic-relevance map visualizes a topical overview of the search result space as keywords with respect to two essential information retrieval measures: relevance and topical similarity. Non-linear dimensionality reduction is used to embed high-dimensional keyword representations of search result data into angles on a radial layout. Relevance of keywords is estimated by a ranking method and visualized as radiuses on the radial layout. As a result, similar keywords are modeled by nearby points, dissimilar keywords are modeled by distant points, more relevant keywords are closer to the center of the radial display, and less relevant keywords are distant from the center of the radial display. We evaluated the effect of the topic-relevance map in a search result comprehension task where 24 participants were summarizing search results and produced a conceptualization of the result space. The results show that topic-relevance map significantly improves participants' comprehension capability compared to a conventional ranked list presentation.
Jaakko Peltonen, Kseniia Belorustceva, Tuukka Ruotsalo
IUI1
2017 Negative Relevance Feedback for Exploratory Search with Visual Interactive Intent Modeling
abstract
In difficult information seeking tasks, the majority of top-ranked documents for an initial query may be non-relevant, and negative relevance feedback may then help find relevant documents. Traditional negative relevance feedback has been studied on document results; we introduce a system and interface for negative feedback in a novel exploratory search setting, where continuous-valued feedback is directly given to keyword features of an inferred probabilistic user intent model. The introduced system allows both positive and negative feedback directly on an interactive visual interface, by letting the user manipulate keywords on an optimized visualization of modeled user intent. Feedback on the interactive intent model lets the user direct the search: Relevance of keywords is estimated from feedback by Bayesian inference, influence of feedback is increased by a novel propagation step, documents are retrieved by likelihoods of relevant versus non-relevant intents, and the most relevant keywords (having the highest upper confidence bounds of relevance) and the most non-relevant ones (having the smallest lower confidence bounds of relevance) are shown as options for further feedback. We carry out task-based information seeking experiments with real users on difficult real tasks; we compare the system to the nearest state of the art baseline allowing positive feedback only, and show negative feedback significantly improves the quality of retrieved information and user satisfaction for difficult tasks.
Jaakko Peltonen, Jonathan Strahl, Patrik Floréen
IUI1
2017 Visualizing activity traces to support collaborative literature searching
abstract
Following recent advances in visual interfaces for search, we investigate how to visualize activity traces to support collaborative information-seeking. We implemented a prototype system that visualizes traces of three types of activities (queries typed, articles bookmarked, and interested keywords) on top of a recent visual search system. The current interface of the system provides an interactive keyword visualization to support exploratory search. We designed two icons to visualize user interactions with the keywords and articles. We also implemented a time line to explicitly display the issued queries and additional details about the interactions. We conducted a longitudinal user study to evaluate the usability of these visualizations. We found that the visualization of traces of issued queries and bookmarked articles help the users to better monitor the progress of the collaborators, increase collaboration, and reassure their findings. Traces of interested keywords support sub-topic identification. Furthermore, activity traces of collaborators help the users to learn about the system features. The findings and the novel collaborative system provide valuable insights and tools into future research on designing interfaces for collaborative information-seeking.
Kumaripaba Athukorala, Luana Micallef, Chao An, Aki Reijonen, Jaakko Peltonen, Tuukka Ruotsalo, Giulio Jacucci
VINCI5
2017 What you see is what you can change: Human-centered machine learning by interactive visualization
Dominik Sacha, Michael Sedlmair, Leishi Zhang, John A. Lee 0001, Jaakko Peltonen, Daniel Weiskopf, Stephen C. North, Daniel A. Keim
Neurocomputing5
2017 Visual Interaction with Dimensionality Reduction: A Structured Literature Analysis
abstract
Dimensionality Reduction (DR) is a core building block in visualizing multidimensional data. For DR techniques to be useful in exploratory data analysis, they need to be adapted to human needs and domain-specific problems, ideally, interactively, and on-the-fly. Many visual analytics systems have already demonstrated the benefits of tightly integrating DR with interactive visualizations. Nevertheless, a general, structured understanding of this integration is missing. To address this, we systematically studied the visual analytics and visualization literature to investigate how analysts interact with automatic DR techniques. The results reveal seven common interaction scenarios that are amenable to interactive control such as specifying algorithmic constraints, selecting relevant features, or choosing among several DR algorithms. We investigate specific implementations of visual analysis systems integrating DR, and analyze ways that other machine learning methods have been combined with DR. Summarizing the results in a "human in the loop" process model provides a general lens for the evaluation of visual interactive DR systems. We apply the proposed model to study and classify several systems previously described in the literature, and to derive future research opportunities.
Dominik Sacha, Leishi Zhang, Michael Sedlmair, John A. Lee 0001, Jaakko Peltonen, Daniel Weiskopf, Stephen C. North, Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.5
2016 Peacock Bundles: Bundle Coloring for Graphs with Globality-Locality Trade-Off
Jaakko Peltonen, Ziyuan Lin
GD1
2015 Majorization-Minimization for Manifold Embedding
abstract
Nonlinear dimensionality reduction by manifold embedding has become a popular and powerful approach both for visualization and as preprocessing for predictive tasks, but more efficient optimization algorithms are still crucially needed. Majorization-Minimization (MM) is a promising approach that monotonically decreases the cost function, but it remains unknown how to tightly majorize the manifold embedding objective functions such that the resulting MM algorithms are efficient and robust. We propose a new MM procedure that yields fast MM algorithms for a wide variety of manifold embedding problems. In our majorization step, two parts of the cost function are respectively upper bounded by quadratic and Lipschitz surrogates, and the resulting upper bound can be minimized in closed form. For cost functions amenable to such QL-majorization, the MM yields monotonic improvement and is efficient: in experiments the newly developed MM algorithms outperform five state-of-the-art optimization approaches in manifold embedding tasks.
Zhirong Yang, Jaakko Peltonen, Samuel Kaski
AISTATS2
2015 IntentStreams: Smart Parallel Search Streams for Branching Exploratory Search
abstract
The user's understanding of information needs and the information available in the data collection can evolve during an exploratory search session. Search systems tailored for well-defined narrow search tasks may be suboptimal for exploratory search where the user can sequentially refine the expressions of her information needs and explore alternative search directions. A major challenge for exploratory search systems design is how to support such behavior and expose the user to relevant yet novel information that can be difficult to discover by using conventional query formulation techniques. We introduce IntentStreams, a system for exploratory search that provides interactive query refinement mechanisms and parallel visualization of search streams. The system models each search stream via an intent model allowing rapid user feedback. The user interface allows swift initiation of alternative and parallel search streams by direct manipulation that does not require typing. A study with 13 participants shows that IntentStreams provides better support for branching behavior compared to a conventional search system.
Salvatore Andolina, Khalil Klouche, Jaakko Peltonen, Mohammad E. Hoque, Tuukka Ruotsalo, Diogo Cabral, Arto Klami, Dorota Glowacka, Patrik Floréen, Giulio Jacucci
IUI3
2015 SciNet: Interactive Intent Modeling for Information Discovery
abstract
Current search engines offer limited assistance for exploration and information discovery in complex search tasks. Instead, users are distracted by the need to focus their cognitive efforts on finding navigation cues, rather than selecting relevant information. Interactive intent modeling enhances the human information exploration capacity through computational modeling, visualized for interaction. Interactive intent modeling has been shown to increase task-level information seeking performance by up to 100%. In this demonstration, we showcase SciNet, a system implementing interactive intent modeling on top of a scientific article database of over 60 million documents.
Tuukka Ruotsalo, Jaakko Peltonen, Manuel J. A. Eugster, Dorota Glowacka, Aki Reijonen, Giulio Jacucci, Petri Myllymäki, Samuel Kaski
SIGIR2
2015 User Model in a Box: Cross-System User Model Transfer for Resolving Cold Start Problems
Chirayu Wongchokprasitti, Jaakko Peltonen, Tuukka Ruotsalo, Payel Bandyopadhyay, Giulio Jacucci, Peter Brusilovsky
UMAP2
2015 Information retrieval approach to meta-visualization
Jaakko Peltonen, Ziyuan Lin
Mach. Learn.1
2014 Optimal Neighborhood Preserving Visualization by Maximum Satisfiability
abstract
We present a novel approach to low-dimensional neighbor embedding for visualization, based on formulating an information retrieval based neighborhood preservation cost function as Maximum satisfiability on a discretized output display. The method has a rigorous interpretation as optimal visualization based on the cost function. Unlike previous low-dimensional neighbor embedding methods, our formulation is guaranteed to yield globally optimal visualizations, and does so reasonably fast. Unlike previous manifold learning methods yielding global optima of their cost functions, our cost function and method are designed for low-dimensional visualization where evaluation and minimization of visualization errors are crucial. Our method performs well in experiments, yielding clean embeddings of datasets where a state-of-the-art comparison method yields poor arrangements. In a real-world case study for semi-supervised WLAN signal mapping in buildings we outperform state-of-the-art methods.
Kerstin Bunte, Matti Järvisalo, Jeremias Berg, Petri Myllymäki, Jaakko Peltonen, Samuel Kaski
AAAI5
2014 Optimization Equivalence of Divergences Improves Neighbor Embedding
abstract
Visualization methods that arrange data objects in 2D or 3D layouts have followed two main schools, methods oriented for graph layout and methods oriented for vectorial embedding. We show the two previously separate approaches are tied by an optimization equivalence, making it possible to relate methods from the two approaches and to build new methods that take the best of both worlds. In detail, we prove a theorem of optimization equivalences between beta- and gamma-, as well as alpha- and Renyi-divergences through a connection scalar. Through the equivalences we represent several nonlinear dimensionality reduction and graph drawing methods in a generalized stochastic neighbor embedding setting, where information divergences are minimized between similarities in input and output spaces, and the optimal connection scalar provides a natural choice for the tradeoff between attractive and repulsive forces. We give two examples of developing new visualization methods through the equivalences: 1) We develop weighted symmetric stochastic neighbor embedding (ws-SNE) from Elastic Embedding and analyze its benefits, good performance for both vectorial and network data; in experiments ws-SNE has good performance across data sets of different types, whereas comparison methods fail for some of the data sets; 2) we develop a gamma-divergence version of a PolyLog layout method; the new method is scale invariant in the output space and makes it possible to efficiently use large-scale smoothed neighborhoods.
Zhirong Yang, Jaakko Peltonen, Samuel Kaski
ICML2
2014 Optimizing Spatial and Temporal Reuse inWireless Networks by Decentralized Partially Observable Markov Decision Processes
abstract
The performance of medium access control (MAC) depends on both spatial locations and traffic patterns of wireless agents. In contrast to conventional MAC policies, we propose a MAC solution that adapts to the prevailing spatial and temporal opportunities. The proposed solution is based on a decentralized partially observable Markov decision process (DEC-POMDP), which is able to handle wireless network dynamics described by a Markov model. A DEC-POMDP takes both sensor noise and partial observations into account, and yields MAC policies that are optimal for the network dynamics model. The DEC-POMDP MAC policies can be optimized for a freely chosen goal, such as maximal throughput or minimal latency, with the same algorithm. We make approximate optimization efficient by exploiting problem structure: the policies are optimized by a factored DEC-POMDP method, yielding highly compact state machine representations for MAC policies. Experiments show that our approach yields higher throughput and lower latency than CSMA/CA based comparison methods adapted to the current wireless network configuration.
Joni Pajarinen, Ari Hottinen, Jaakko Peltonen
IEEE Trans. Mob. Comput.3
2013 Information Retrieval Perspective to Meta-visualization
abstract
In visual data exploration with scatter plots, no single plot is sufficient to analyze complicated high-dimensional data sets. Given numerous visualizations created with different features or methods, meta-visualization is needed to analyze the visualizations together. We solve \emphhow to arrange numerous visualizations onto a meta-visualization display, so that their similarities and differences can be analyzed. We introduce a machine learning approach to optimize the meta-visualization, based on an information retrieval perspective: two visualizations are similar if the analyst would retrieve similar neighborhoods between data samples from either visualization. Based on the approach, we introduce a nonlinear embedding method for meta-visualization: it optimizes locations of visualizations on a display, so that visualizations giving similar information about data are close to each other.
Jaakko Peltonen, Ziyuan Lin
ACML1
2013 Directing exploratory search with interactive intent modeling
abstract
We introduce interactive intent modeling, where the user directs exploratory search by providing feedback for estimates of search intents. The estimated intents are visualized for interaction on an Intent Radar, a novel visual interface that organizes intents onto a radial layout where relevant intents are close to the center of the visualization and similar intents have similar angles. The user can give feedback on the visualized intents, from which the system learns and visualizes improved intent estimates. We systematically evaluated the effect of the interactive intent modeling in a mixed-method task-based information seeking setting with 30 users, where we compared two interface variants for interactive intent modeling, namely intent radar and a simpler list-based interface, to a conventional search system. The results show that interactive intent modeling significantly improves users' task performance and the quality of retrieved information.
Tuukka Ruotsalo, Jaakko Peltonen, Manuel J. A. Eugster, Dorota Glowacka, Ksenia Konyushkova, Kumaripaba Athukorala, Ilkka Kosunen, Aki Reijonen, Petri Myllymäki, Giulio Jacucci, Samuel Kaski
CIKM2
2013 Scalable Optimization of Neighbor Embedding for Visualization
abstract
Neighbor embedding (NE) methods have found their use in data visualization but are limited in big data analysis tasks due to their O(n^2) complexity for n data samples. We demonstrate that the obvious approach of subsampling produces inferior results and propose a generic approximated optimization technique that reduces the NE optimization cost to O(n log n). The technique is based on realizing that in visualization the embedding space is necessarily very low-dimensional (2D or 3D), and hence efficient approximations developed for n-body force calculations can be applied. In gradient-based NE algorithms the gradient for an individual point decomposes into “forces” exerted by the other points. The contributions of close-by points need to be computed individually but far-away points can be approximated by their “center of mass”, rapidly computable by applying a recursive decomposition of the visualization space into quadrants. The new algorithm brings a significant speed-up for medium-size data, and brings “big data” within reach of visualization.
Zhirong Yang, Jaakko Peltonen, Samuel Kaski
ICML (2)2
2013 Expectation Maximization for Average Reward Decentralized POMDPs
Joni Pajarinen, Jaakko Peltonen
ECML/PKDD (1)2
2013 Transfer learning using a nonparametric sparse topic model
Ali Faisal, Jussi Gillberg, Gayle Leen, Jaakko Peltonen
Neurocomputing4
2012 Sparse Nonparametric Topic Model for Transfer Learning
Ali Faisal, Jussi Gillberg, Jaakko Peltonen, Gayle Leen, Samuel Kaski
ESANN3
2012 Machine learning for signal processing 2010
Jaakko Peltonen, Tapani Raiko, Samuel Kaski
Neurocomputing1
2012 Focused multi-task learning in a Gaussian process framework
Gayle Leen, Jaakko Peltonen, Samuel Kaski
Mach. Learn.2
2011 Efficient Planning for Factored Infinite-Horizon DEC-POMDPs
Joni Pajarinen, Jaakko Peltonen
IJCAI2
2011 Periodic Finite State Controllers for Efficient POMDP and DEC-POMDP Planning
abstract
Applications such as robot control and wireless communication require planning under uncertainty. Partially observable Markov decision processes (POMDPs) plan policies for single agents under uncertainty and their decentralized versions (DEC-POMDPs) find a policy for multiple agents. The policy in infinite-horizon POMDP and DEC-POMDP problems has been represented as finite state controllers (FSCs). We introduce a novel class of periodic FSCs, composed of layers connected only to the previous and next layer. Our periodic FSC method finds a deterministic finite-horizon policy and converts it to an initial periodic infinite-horizon policy. This policy is optimized by a new infinite-horizon algorithm to yield deterministic periodic policies, and by a new expectation maximization algorithm to yield stochastic periodic policies. Our method yields better results than earlier planning methods and can compute larger solutions than with regular FSCs.
Joni Pajarinen, Jaakko Peltonen
NIPS2
2011 Focused Multi-task Learning Using Gaussian Processes
Gayle Leen, Jaakko Peltonen, Samuel Kaski
ECML/PKDD (2)2
2011 Fault tolerant machine learning for nanoscale cognitive radio
Joni Pajarinen, Jaakko Peltonen, Mikko A. Uusitalo
Neurocomputing2
2010 An information retrieval perspective on visualization of gene expression data with ontological annotation
abstract
High-dimensional data are often visualized by dimensionality reduction methods whose goals are not directly related to visualization. We use a recent formalization of visualization as information retrieval and apply that formalism to data with structured annotations: we analyze gene expression data with annotations from the Gene Ontology (GO). We show that using the GO information in visualization yields better retrieval with respect to known ontological relationships and allows discovery of data properties not explained by the ontology.
Jaakko Peltonen, Helena Aidos, Nils Gehlenborg, Alvis Brazma, Samuel Kaski
ICASSP1
2010 Efficient Planning in Large POMDPs through Policy Graph Based Factorized Approximations
Joni Pajarinen, Jaakko Peltonen, Ari Hottinen, Mikko A. Uusitalo
ECML/PKDD (3)2
2010 Relevant subtask learning by constrained mixture models
abstract
We introduce relevant subtask learning, a new learning problem which is a variant of multi-task learning. The goal is to build a classifier for a task-of-interest for which we have too few training samples. We additionally have “supplementary data” c
Jaakko Peltonen, Yusuf Yaslan, Samuel Kaski
Intell. Data Anal.1
2010 Information Retrieval Perspective to Nonlinear Dimensionality Reduction for Data Visualization
Jarkko Venna, Jaakko Peltonen, Kristian Nybo, Helena Aidos, Samuel Kaski
J. Mach. Learn. Res.2
2009 Supervised nonlinear dimensionality reduction by Neighbor Retrieval
abstract
Many recent works have combined two machine learning topics, learning of supervised distance metrics and manifold embedding methods, into supervised nonlinear dimensionality reduction methods. We show that a combination of an early metric learning method and a recent unsupervised dimensionality reduction method empirically outperforms previous methods. In our method, the Riemannian distance metric measures local change of class distributions, and the dimensionality reduction method makes a rigorous tradeoff between precision and recall in retrieving similar data points based on the reduced-dimensional display. The resulting supervised visualizations are good for finding (sets of) similar data samples that have similar class distributions.
Jaakko Peltonen, Helena Aidos, Samuel Kaski
ICASSP1
2009 Latent state models of primary user behavior for opportunistic spectrum access
abstract
Opportunistic spectrum access, where cognitive radio devices detect available unused radio channels and exploit them for communication, avoiding collisions with existing users of the channels, is a central topic of research for future wireless communication. When each device has limited resources to sense which channels are available, the task becomes a reinforcement learning problem that has been studied with partially observable Markov decision processes (POMDPs). However, current POMDP solutions are based on simplistic representations where channels are simply on/off (transmitting or idle). We show that more complicated Markov models where on/off states are part of complicated behavior of the channel owner (primary user) yield better POMDPs achieving more successful transmissions and less collisions.
Joni Pajarinen, Jaakko Peltonen, Mikko A. Uusitalo, Ari Hottinen
PIMRC2
2007 Learning from Relevant Tasks Only
Samuel Kaski, Jaakko Peltonen
ECML2
2007 Methods for estimating human endogenous retrovirus activities from EST databases
abstract
BACKGROUND: Human endogenous retroviruses (HERVs) are surviving traces of ancient retrovirus infections and now reside within the human DNA. Recently HERV expression has been detected in both normal tissues and diseased patients. However, the activities (expression levels) of individual HERV sequences are mostly unknown. RESULTS: We introduce a generative mixture model, based on Hidden Markov Models, for estimating the activities of the individual HERV sequences from EST (expressed sequence tag) databases. We use the model to estimate the relative activities of 181 HERVs. We also empirically justify a faster heuristic method for HERV activity estimation and use it to estimate the activities of 2450 HERVs. The majority of the HERV activities were previously unknown. CONCLUSION: (i) Our methods estimate activity accurately based on experiments on simulated data. (ii) Our estimate on real data shows that 7% of the HERVs are active. The active ones are spread unevenly into HERV groups and relatively uniformly in terms of estimated age. HERVs with the retroviral env gene are more often active than HERVs without env. Few of the active HERVs have open reading frames for retroviral proteins.
Merja Oja, Jaakko Peltonen, Jonas Blomberg, Samuel Kaski
BMC Bioinform.2
2005 Discriminative components of data
abstract
A simple probabilistic model is introduced to generalize classical linear discriminant analysis (LDA) in finding components that are informative of or relevant for data classes. The components maximize the predictability of the class distribution which is asymptotically equivalent to 1) maximizing mutual information with the classes, and 2) finding principal components in the so-called learning or Fisher metrics. The Fisher metric measures only distances that are relevant to the classes, that is, distances that cause changes in the class distribution. The components have applications in data exploration, visualization, and dimensionality reduction. In empirical experiments, the method outperformed, in addition to more classical methods, a Renyi entropy-based alternative while having essentially equivalent computational cost.
Jaakko Peltonen, Samuel Kaski
IEEE Trans. Neural Networks1
2004 Sequential information bottleneck for finite data
abstract
The sequential information bottleneck (sIB) algorithm clusters co-occurrence data such as text documents vs. words. We introduce a variant that models sparse co-occurrence data by a generative process. This turns the objective function of sIB, mutual information, into a Bayes factor, while keeping it intact asymptotically, for non-sparse data. Experimental performance of the new algorithm is comparable to the original sIB for large data sets, and better for smaller, sparse sets.
Jaakko Peltonen, Janne Sinkkonen, Samuel Kaski
ICML1
2004 Improved learning of Riemannian metrics for exploratory analysis
Jaakko Peltonen, Arto Klami, Samuel Kaski
Neural Networks1
2003 Visualizations for Assessing Convergence and Mixing of MCMC
Jarkko Venna, Samuel Kaski, Jaakko Peltonen
ECML3
2003 Informative Discriminant Analysis
Samuel Kaski, Jaakko Peltonen
ICML2
2003 TimeMachine Oulu - Dynamic Creation of Cultural-Spatio-Temporal Models as a Mobile Service
Jaakko Peltonen, Mark Ollila, Timo Ojala
Mobile HCI1
2002 Learning More Accurate Metrics for Self-Organizing Maps
Jaakko Peltonen, Arto Klami, Samuel Kaski
ICANN1
2001 Data Visualization and Analysis with Self-Organizing Maps in Learning Metrics
Samuel Kaski, Janne Sinkkonen, Jaakko Peltonen
DaWaK3
2001 Bankruptcy analysis with self-organizing maps in learning metrics
abstract
We introduce a method for deriving a metric, locally based on the Fisher information matrix, into the data space. A self-organizing map (SOM) is computed in the new metric to explore financial statements of enterprises. The metric measures local distances in terms of changes in the distribution of an auxiliary random variable that reflects what is important in the data. In this paper the variable indicates bankruptcy within the next few years. The conditional density of the auxiliary variable is first estimated, and the change in the estimate resulting from local displacements in the primary data space is measured using the Fisher information matrix. When a self-organizing map is computed in the new metric it still visualizes the data space in a topology-preserving fashion, but represents the (local) directions in which the probability of bankruptcy changes the most.
Samuel Kaski, Janne Sinkkonen, Jaakko Peltonen
IEEE Trans. Neural Networks3