Mateusz Przyborowski

dblp:234/2672 · DBLP profile ↗
← Back
3ranked-venue papers in the field
3as first author
2since 2021 · last 2022
0000-0002-7721-8433ORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 3 (3 first)
YearPublicationVenuePosition
2022 Large Scale Windowed Matching
abstract
Missing or invalid records in sales data are a common obstacle that can damage the overall effectiveness of market analysis. Completing the data on the basis of the records obtained so far can be formulated in means of a schema matching task. In this paper we present a machine learning based method for performing schema matching for transactional data. The analysis is based on a dataset of over 700.000 transactions from retail stores. We confront the proposed solution with manual and conventional approaches.
Mateusz Przyborowski, Krzysztof Ciebiera, Krzysztof Stencel
IEEE Big Data1
2022 Approximation of the expectation-maximization algorithm for Gaussian mixture models on big data
abstract
Gaussian mixture models are a very useful tool for modeling data distribution. While estimating parameters using the expectation-maximization algorithm, this approach does not scale well with big datasets, especially if it is necessary to prepare many models for the proper selection of metaparameters. In this article we present an approximation of the expectation-maximization algorithm obtained by merging crucial subsets of the dataset, that differ slightly in their effect on the expectation-maximization loss function, into information granules. Furthermore, application examples comparing new method with the classical approach are shown.
Mateusz Przyborowski, Dominik Slezak
IEEE Big Data1
2018 Toward Machine Learning on Granulated Data - a Case of Compact Autoencoder-based Representations of Satellite Images
abstract
We consider a problem of learning from compact representations of images for a purpose of object recognition and content-based image retrieval. We discuss a motivation for using compressed images in those tasks and indicate exemplary applications related to analysis on the data from satellites. Finally, we show some preliminary results of experiments conducted to demonstrate the impact of the image data granulation on the quality of classification. We empirically compare the performance of prediction models trained on original images, images compressed using autoencoders, and on images whose quality was lowered in order to reduce their size.
Mateusz Przyborowski, Tomasz Tajmajer, Lukasz Grad, Andrzej Janusz, Piotr Biczyk, Dominik Slezak
IEEE BigData1