Jyrki Nummenmaa

dblp:58/1177 · DBLP profile ↗
← Back
38ranked-venue papers
3as first author
9since 2021 · last 2024
0000-0002-7476-7840ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 25 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 15 · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Theory of computation · 3 · 1 first-authorSecurity and privacy · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 TraQuLA: Transparent Question Answering Over RDF Through Linguistic Analysis
Elizaveta Zimina, Kalervo Järvelin, Jaakko Peltonen, Aarne Ranta, Jyrki Nummenmaa
ICWE5
2023 Fair Neighbor Embedding
abstract
We consider fairness in dimensionality reduction. Nonlinear dimensionality reduction yields low dimensional representations that let users visualize and explore high-dimensional data. However, traditional dimensionality reduction may yield biased visualizations overemphasizing relationships of societal phenomena to sensitive attributes or protected groups. We introduce a framework of fair neighbor embedding, the Fair Neighbor Retrieval Visualizer, which formulates fair nonlinear dimensionality reduction as an information retrieval task whose performance and fairness are quantified by information retrieval criteria. The method optimizes low-dimensional embeddings that preserve high-dimensional data neighborhoods without yielding biased association of such neighborhoods to protected groups. In experiments the method yields fair visualizations outperforming previous methods.
Jaakko Peltonen, Timo Nummenmaa, Jyrki Nummenmaa
ICML4
2022 Nonparametric exponential family graph embeddings for multiple representation learning
abstract
In graph data, each node often serves multiple functionalities. However, most graph embedding models assume that each node can only possess one representation. We address this issue by proposing a nonparametric graph embedding model. The model allows each node to learn multiple representations where they are needed to represent the complexity of random walks in the graph. It extends the Exponential family graph embedding model with two nonparametric prior settings, the Dirichlet process and the uniform process. The model combines the ability of Exponential family graph embedding to take the number of occurrences of context nodes into account with nonparametric priors giving it the flexibility to learn more than one latent representation for each node. The learned embeddings outperform other state of the art approaches in link prediction and node classification tasks.
Chien Lu, Jaakko Peltonen, Timo Nummenmaa, Jyrki Nummenmaa
UAI4
2022 Using parsed and annotated corpora to analyze parliamentarians' talk in Finland
abstract
Abstract We present a search system for grammatically analyzed corpora of Finnish parliamentary records and interviews with former parliamentarians, annotated with metadata of talk structure and involved parliamentarians, and discuss their use through carefully chosen digital humanities case studies. We first introduce the construction, contents, and principles of use of the corpora. Then we discuss the application of the search system and the corpora to study how politicians talk about power, how ideological terms are used in political speech, and how to identify narratives in the data. All case studies stem from questions in the humanities and the social sciences, but rely on the grammatically parsed corpora in both identifying and quantifying passages of interest. Finally, the paper discusses the role of natural language processing methods for questions in the (digital) humanities. It makes the claim that a digital humanities inquiry of parliamentary speech and interviews with politicians cannot only rely on computational humanities modeling, but needs to accommodate a range of perspectives starting with simple searches, quantitative exploration, and ending with modeling. Furthermore, the digital humanities need a more thorough discussion about how the utilization of tools from information science and technologies alter the research questions posed in the humanities.
Mykola Andrushchenko, Kirsi Sandberg, Risto Turunen, Jani Marjanen, Mari Hatavara, Jussi Kurunmäki, Timo Nummenmaa, Matti Hyvärinen, Kari Teräs, Jaakko Peltonen, Jyrki Nummenmaa
J. Assoc. Inf. Sci. Technol.11
2022 Sequential group recommendations based on satisfaction and disagreement scores
abstract
Abstract Recently, group recommendations have gained much attention. Nevertheless, most approaches consider only one round of recommendations. However, in a real-life scenario, it is expected that the history of previous recommendations is exploited to tailor the recommendations towards meeting the needs of the group members. Such history should include not only which items the system suggested, but also the reaction of the members to these items. This work introduces the problem of sequential group recommendations, by exploiting the concept of satisfaction and disagreement. Satisfaction describes how well the group received the suggested items. Disagreement describes the satisfaction bias among the group members. We utilize these concepts in three new aggregation methods, SDAA, SIAA and Average+, designed to address the specific challenges introduced by sequential group recommendations. We experimentally show the effectiveness of our methods using big real datasets for both stable and ephemeral groups.
Maria Stratigi, Evaggelia Pitoura, Jyrki Nummenmaa, Kostas Stefanidis
J. Intell. Inf. Syst.3
2022 Efficient mining of concept-hierarchy aware distinguishing sequential patterns
Chengxin He, Lei Duan, Guozhu Dong, Jyrki Nummenmaa, Tingting Wang 0009, Tinghai Pang
Knowl. Based Syst.4
2021 Cross-structural Factor-topic Model: Document Analysis with Sophisticated Covariates
abstract
Modern text data is increasingly gathered in situations where it is paired with a high-dimensional collection of covariates: then both the text, the covariates, and their relationships are of interest to analyze. Despite the growing amount of such data, current topic models are unable to take into account large amounts of covariates successfully: they fail to model structure among covariates and distort findings of both text and covariates. This paper presents a solution: a novel factor-topic model that enables researchers to analyze latent structure in both text and sophisticated document-level covariates collectively. The key innovation is that besides learning the underlying topical structure, the model also learns the underlying factorial structure from the covariates and the interactions between the two structures. A set of tailored variational inference algorithms for efficient computation are provided. Experiments on three different datasets show the model outperforms comparable topic models in the ability to predict held-out document content. Two case studies focusing on Finnish parliamentary election candidates and game players on Steam demonstrate the model discovers semantically meaningful topics, factors, and their interactions. The model both outperforms state-of-the-art models in predictive accuracy and offers new factor-topic insights beyond other topic models.
Chien Lu, Jaakko Peltonen, Timo Nummenmaa, Jyrki Nummenmaa, Kalervo Jäarvelin
ACML4
2021 HMNet: Hybrid Matching Network for Few-Shot Link Prediction
Shan Xiao, Lei Duan, Guicai Xie, Renhao Li, Geng Deng, Jyrki Nummenmaa
DASFAA (1)7
2021 Efficient Mining of Outlying Sequential Behavior Patterns
Lei Duan, Guicai Xie, Longhai Li, Jyrki Nummenmaa
DASFAA (2)6
2020 Probabilistic Dynamic Non-negative Group Factor Model for Multi-source Text Mining
abstract
Nonnegative matrix factorization (NMF) is a popular approach to model data, however, most models are unable to flexibly take into account multiple matrices across sources and time or apply only to integer-valued data. We introduce a probabilistic, Gaussian Process-based, more inclusive NMF-based model which jointly analyzes nonnegative data such as text data word content from multiple sources in a temporal dynamic manner. The model collectively models observed matrix data, source-wise latent variables, and their dependencies and temporal evolution with a full-fledged hierarchical approach including flexible nonparametric temporal dynamics. Experiments on simulated data and real data show the model out-performs, comparable models. A case study on social media and news demonstrates the model discovers semantically meaningful topical factors and their evolution
Chien Lu, Jaakko Peltonen, Jyrki Nummenmaa, Kalervo Järvelin
CIKM3
2019 Discovering Relationship Patterns Among Associated Temporal Event Sequences
Lei Duan, Ruiqi Qin 0001, Jyrki Nummenmaa
DASFAA (1)6
2019 Incremental Blocking for Entity Resolution over Web Streaming Data
abstract
The widespread use of information systems has become a valuable source of semi-structured data. In this context, Entity Resolution (ER) emerges as a fundamental task to integrate multiple knowledge bases or identify similarities between data items (i.e., entities). Since ER is an inherently quadratic task, blocking techniques are often used to improve efficiency. Beyond the challenges related to the data volume and heterogeneity, blocking techniques also face two other challenges: streaming data and incremental processing. To address these challenges, we propose PRIME, a novel incremental schema-agnostic blocking technique that utilizes parallelism to enhance blocking efficiency. The proposed technique deals with streaming and incremental data using a distributed computational infrastructure. To improve efficiency, the technique avoids unnecessary comparisons and applies a time window strategy to prevent excessive memory consumption.
Tiago Brasileiro Araújo, Kostas Stefanidis, Carlos Eduardo S. Pires, Jyrki Nummenmaa, Thiago Pereira da Nóbrega
WI4
2019 Detecting measurement issues in SQL arithmetic expressions and aggregations
Peter Thanisch, Tapio Niemi, Jyrki Nummenmaa, Marko Niinimäki
Data Knowl. Eng.3
2018 A Player Behavior Model for Predicting Win-Loss Outcome in MOBA Games
Xuan Lan, Lei Duan, Ruiqi Qin 0001, Timo Nummenmaa, Jyrki Nummenmaa
ADMA6
2018 Author Tree-Structured Hierarchical Dirichlet Process
Md. Hijbul Alam, Jaakko Peltonen, Jyrki Nummenmaa, Kalervo Järvelin
DS3
2018 Bus-OLAP: A Data Management Model for Non-on-Time Events Query Over Bus Journey Data
abstract
Increasing the on-time rate of bus service can prompt the people’s willingness to travel by bus, which is an effective measure to mitigate the city traffic congestion. Performing queries on the bus arrival can be used to identify and analyze various kinds of non-on-time events that happened during the bus journey, which is helpful for detecting the factors of delaying events, and providing decision support for optimizing the bus schedules. We propose a data management model, called Bus-OLAP, for querying bus journey data, considering the characteristics of bus running and the scenarios of non-on-time analysis. While fulfilling typical requirements of bus journey data queries, Bus-OLAP not only provides a flexible way to manage the data and to implement multiple granularity data query and update, but it also supports distributed queries and computation. The experiments on real-world bus journey data verify that Bus-OLAP is effective and efficient.
Lei Duan, Tinghai Pang, Jyrki Nummenmaa, Jie Zuo, Changjie Tang
Data Sci. Eng.3
2017 Mining Top-k Distinguishing Temporal Sequential Patterns from Event Sequences
Lei Duan, Guozhu Dong, Jyrki Nummenmaa
DASFAA (2)4
2017 A Minimum Distortion: High Capacity Watermarking Technique for Relational Data
abstract
In this paper, a new multi-attribute and high capacity image-based watermarking technique for relational data is proposed. The embedding process causes low distortion into the data considering the usability restrictions defined over the marked relation. The conducted experiments show the high resilience of the proposed technique against tuple deletion and tuple addition attacks. An interesting trend of the extracted watermark is analyzed when, within certain limits, if the number of embedded marks is small, the watermark signal far from being compromised, discretely improves in the case of tuple addition attacks. According to the results, marking 13% of the attributes and under an attack of 100% of tuples addition, 96% of the watermark is extracted. Also, while previous techniques embed up to 61% of the watermark, under the same conditions, we guarantee to embed 99.96% of the marks.
Maikel L. Pérez Gort, Claudia Feregrino-Uribe, Jyrki Nummenmaa
IH&MMSec3
2016 Mining Distinguishing Customer Focus Sets for Online Shopping Decision Support
Lei Duan, Jyrki Nummenmaa, Guozhu Dong, Pan Qin
ADMA4
2015 Mining Itemset-based Distinguishing Sequential Patterns with Gap Constraint
Lei Duan, Guozhu Dong, Jyrki Nummenmaa, Changjie Tang
DASFAA (1)4
2015 Automatic verification of Dafny programs with traits
abstract
This paper describes the design of traits, abstract superclasses, in the verification-aware programming language Dafny. Although there is no inheritance among classes in Dafny, the traits make it possible to describe behavior common to several classes and to write code that abstracts over the particular classes involved. The design incorporates behavioral specifications for a trait's methods and functions, just like for classes in Dafny. The design has been implemented in the Dafny tool.
Reza Ahmadi, K. Rustan M. Leino, Jyrki Nummenmaa
FTfJP@ECOOP3
2014 Mining Frequent Closed Sequential Patterns with Non-user-defined Gap Constraints
Lei Duan, Jyrki Nummenmaa, Song Deng, Zhong-Qi Li, Changjie Tang
ADMA3
2014 Models for Mobile Application Maintenance Based on Update History
abstract
Good software development and particularly maintenance practices form an important factor for success in software business. If one wants to constantly produce new successful releases of the applications, a proper efficient software maintenance process is the key. In this work, we study data from mobile application maintenance to understand and conceptualize how mobile application maintenance takes place. Based on the data on release history, we deduce different mobile application maintenance models from the perspectives of maintenance scheduling and maintenance requirements.
Xiaozhou Li 0002, Zheying Zhang, Jyrki Nummenmaa
ENASE3
2014 Detecting summarizability in OLAP
Tapio Niemi, Marko Niinimäki, Peter Thanisch, Jyrki Nummenmaa
Data Knowl. Eng.4
2013 Cell-at-a-Time Approach to Lazy Evaluation of Dimensional Aggregations
Peter Thanisch, Jyrki Nummenmaa, Tapio Niemi, Marko Niinimäki
DaWaK2
2013 Decision-making in rights exporting: the integrated process
abstract
Rights exporting plays an essential role in the battle of fighting for Digital Rights Management (DRM) interoperability. The decision making process determines the results of rights exporting. In order to achieve optimal results in rights exporting, we leverage the process with algorithms for rights adaptation and rights decomposition. We also demonstrate how the proposed process can lead to improved results.
Wenhui Lu, Zheying Zhang, Jyrki Nummenmaa
MEDES3
2012 Deploying adaptation in rights exporting
abstract
The incompatibility of various Digital Rights Management (DRM) systems remain as an obstacle hampering user experience and business on DRM based services. To increase DRM interoperability, rights exporting currently seems a preferred solution. Based on an earlier DRM model [8,9], we extend and consolidate the concept of rights adaptation in rights exporting. We analyse different ways of adapting rights to be exported and present a general rights adaptation framework.
Wenhui Lu, Zheying Zhang, Jyrki Nummenmaa
CCNC3
2007 Ontologies with Semantic Web/Grid in Data Integration for OLAP
abstract
Traditionally, data used in OLAP (online analytical processing) have been limited to the contents of the data warehouse of a company. However, the needs for analysis are often more demanding and data are needed from different sources. In this article, we study how the semantics of data sources can be described to allow combining data from several sources into an OLAP cube. We apply Semantic Web technologies for defining an OWL/RDF ontology for OLAP data sources and OLAP cubes. These definitions are then utilised in OLAP cube formation by posing an OWL/RDF ontology-based query against them. We use Grid technologies to enhance the efficiency of processing and ensuring security. Our primary interest is in the cube construction (i.e., ETL process), and we assume that standard OLAP methods can be used for the actual analysis. Our tests show that the proposed approach can speed up the construction of an OLAP cube for ad hoc queries by supporting a high-level query language and reducing the amount of required data.
Tapio Niemi, Santtu Toivonen, Marko Niinimäki, Jyrki Nummenmaa
Int. J. Semantic Web Inf. Syst.4
2003 Normalising OLAP cubes for controlling sparsity
Tapio Niemi, Jyrki Nummenmaa, Peter Thanisch
Data Knowl. Eng.2
2002 Constructing an OLAP cube from distributed XML data
abstract
On-Line Analytical Processing (OLAP) is a powerful method for analysing large data warehouse data. Typically, the data for an OLAP database is collected from a set of data repositories such as e.g. operational databases. This data set is often huge, and it may not be known in advance what data is required and when to perform the desired data analysis tasks. Sometimes it may happen that some parts of the data are only needed occasionally. Therefore, keeping the OLAP database constantly up-to-date is not only a highly demanding task but it also may be overkill in practice.This suggests that in some applications it would be more feasible to form the OLAP cubes only when they are actually needed. We present such a system. As the data sources may well be heterogeneous, we propose an XML language for data collection. Our system also has a facility, where the user may pose a query against a universal OLAP cube using the MDX language. The query is analysed to determine which data is required for the desired OLAP cube.
Tapio Niemi, Marko Niinimäki, Jyrki Nummenmaa, Peter Thanisch
DOLAP3
2002 A Linear Time Special Case for MC Games
Timo Poranen, Jyrki Nummenmaa
Fundam. Informaticae2
2001 Constructing OLAP Cubes Based on Queries
abstract
An On-Line Analytical Processing (OLAP) user often follows a train of thought, posing a sequence of related queries against the data warehouse. Although their details are not known in advance, the general form of those queries is apparent beforehand. Thus, the user can outline the relevant portion of the data posing generalised queries against a cube representing the data warehouse.Since existing OLAP design methods are not suitable for non-professionals, we present a technique that automates cube design given the data warehouse, functional dependency information, and sample OLAP queries expressed in the general form. The method constructs complete but minimal cubes with low risks related to sparsity and incorrect aggregations. After the user has given queries, the system will suggest a cube design. The user can accept it or improve it by giving more queries. The method is also suitable for improving existing cubes using respective real MDX queries.
Tapio Niemi, Jyrki Nummenmaa, Peter Thanisch
DOLAP2
2000 Functional Dependencies in Controlling Sparsity of OLAP Cubes
Tapio Niemi, Jyrki Nummenmaa, Peter Thanisch
DaWaK2
2000 A Query Method Based on Intensional Concept Definition
Tapio Niemi, Jyrki Nummenmaa
EJC2
1994 Finding Compact Scheme Forests in Nested Normal Form is NP-Hard
Peter Thanisch, George Loizou, Jyrki Nummenmaa
Inf. Comput.3
1992 Constructing Compact Rectilinear Planar Layouts Using Canonical Representation of Planar Graphs
Jyrki Nummenmaa
Theor. Comput. Sci.1
1990 Constructing Layouts for ER-Diagrams from Visibility-Representations
Jyrki Nummenmaa, J. Tuomi
ER1
1990 Conjectures and Refutations in Database Design and Dependency Theory
Jyrki Nummenmaa, Peter Thanisch
ICDT1