Marek Grzegorowski

dblp:122/1797 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
5since 2021 · last 2023
0000-0003-4740-0725ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2023 Survival-Based Feature Extraction - Application in Supply Management for Dispersed Vending Machines
abstract
The outbreak of the COVID pandemic revealed that supply chains are not resilient to such a type of turmoil, and the food industry appeared to be particularly vulnerable. Meanwhile, customers expect uninterrupted deliveries and the products’ selection responding to their preferences. In this article, we discuss several topics related to supply management that allow preparing a delivery plan for a distributed network of vending machines, considering each location individually. The developed solution takes advantage of the state-of-the-art machine learning methods. However, it is human-centric and aligned with the concept of Industry 5.0. We present the conceptual and technological side of the solution with a particular emphasis on the developed feature extraction framework, which uses selected indicators from the survival analysis. We present an analysis of the real data confirming that the proposed approach copes well with high uncertainty in data, addressing the cold-start problem.
Marek Grzegorowski, Jaroslaw Litwin, Mateusz Wnuk, Mateusz Pabis, Lukasz Marcinowski
IEEE Trans. Ind. Informatics1
2022 Speeding Up Recommender Systems Using Association Rules
Eyad Kannout, Hung Son Nguyen, Marek Grzegorowski
ACIIDS (2)3
2022 Utilizing Frequent Pattern Mining for Solving Cold-Start Problem in Recommender Systems
abstract
Although several approaches have been proposed throughout the last decade to build recommender systems (RS), most of them suffer from the cold-start problem.This problem occurs when a new item hits the system or a new user signs up.It is generally recognized that the ability to handle cold users and items is one of the key success factors of any new recommender algorithm.This paper introduces a frequent pattern mining framework for recommender systems (FPRS) -a novel approach to address this challenging task.FPRS is a hybrid RS that incorporates collaborative and content-based recommendation algorithms and employs a frequent pattern (FP) growth algorithm.The article proposes several strategies to combine the generated frequent itemsets with content-based methods to mitigate the cold-start problem for both new users and new items.The performed empirical evaluation confirmed its usefulness.Furthermore, the developed solution can be easily combined with any other approach to build a recommender system and can be further extended to make up a complete and standalone RS.Index Terms-recommendation system, cold-start problem, frequent pattern mining, quality of recommendations.
Eyad Kannout, Michal Grodzki, Marek Grzegorowski
FedCSIS3
2022 Considering various aspects of models' quality in the ML pipeline - application in the logistics sector
abstract
The industrial machine learning applications today involve developing and deploying MLOps pipelines to ensure the versatile quality of forecasting models over an extended period, simultaneously assuring the model's accuracy, stability, short training time, and resilience.In this study, we present the ML pipeline conforming to all the abovementioned aspects of models' quality formulated as a constrained multi-objective optimization problem.We also provide the reference implementation on stateof-the-art methods for data preprocessing, feature extraction, dimensionality reduction, feature and instance selection, model fitting, and ensemble blending.The experimental study on the real data set from the logistics industry confirmed the qualities of the proposed approach, as the successful participation in an international data competition did.
Eyad Kannout, Michal Grodzki, Marek Grzegorowski
FedCSIS3
2022 Prescriptive Analytics for Optimization of FMCG Delivery Plans
Marek Grzegorowski, Andrzej Janusz, Stanislaw Lazewski, Maciej Swiechowski, Monika Jankowska
IPMU (2)1
2019 Cluster-size optimization within a cloud-based ETL framework for Big Data
abstract
The ability to analyze the available data is a valuable asset for any successful business, especially when the analysis yields meaningful knowledge. One of the key processes for acquiring such ability is the Extract-Transform-Load (ETL) process. For Big Data, ETL requires a significant effort and it is a very challenging task to be performed in a cost-effective way. There are quite a few examples in the literature that describe an architecture for cost-effective ETL but none of the available examples are complete enough and they are usually evaluated in narrow problem domains. The ones that are more general, require specific implementation details. In this paper we propose a cloud-based ETL framework where we use a general cluster-size optimization algorithm, while providing implementation details, and is able to perform the required job within a predefined, and thus known, time. We evaluated the algorithm by executing three scenarios regarding data aggregation during ETL: (i) ETL with no aggregation; (ii) aggregation based on predefined columns or time intervals; and (iii) aggregation within single user sessions spanning over arbitrary time intervals. The execution of the three ETL scenarios in a production setting showed that the cluster size could be optimized so it can process the required data volume within a predefined and thus, expected, latency. The scalability was evaluated on Amazon AWS Hadoop clusters by processing user logs collected with Kinesis streams with datasets ranging from 30 GB to 2.6 TB.
Eftim Zdravevski, Petre Lameski, Ace Dimitrievski, Marek Grzegorowski, Cas Apanowicz
IEEE BigData4
2019 Clash Royale Challenge: How to Select Training Decks for Win-rate Prediction
abstract
We summarize the sixth data mining competition organized at the Knowledge Pit platform in association with the Federated Conference on Computer Science and Information Systems series, titled Clash Royale Challenge: How to Select Training Decks for Win-rate Prediction.We outline the scope of this challenge and briefly present its results.We also discuss the problem of acquiring knowledge about new notions from video games through an active learning cycle.We explain how this task is related to the problem considered in the challenge and share results of experiments that we conducted to demonstrate usefulness of the active learning approach in practice.
Andrzej Janusz, Lukasz Grad, Marek Grzegorowski
FedCSIS3
2019 On resilient feature selection: Computational foundations of r-C-reducts
Marek Grzegorowski, Dominik Slezak
Inf. Sci.1
2018 A framework for learning and embedding multi-sensor forecasting models into a decision support system: A case study of methane concentration in coal mines
Dominik Slezak, Marek Grzegorowski, Andrzej Janusz, Michal Kozielski, Sinh Hoa Nguyen, Marek Sikora, Sebastian Stawicki, Lukasz Wróbel
Inf. Sci.2
2017 On the role of feature space granulation in feature selection processes
abstract
Information granulation plays an important role in the process of scaling up modern machine learning and knowledge discovery algorithms. By employing compact descriptions of granules - whereby granules are defined as collections of original data elements gathered together by means of their similarity, proximity or functionality - one can drastically accelerate computations and, moreover, make the results of those computations more meaningful for domain experts. In this paper, we summarize some of the feature space granulation approaches introduced by now. We discuss the meaning of similarity, proximity and functionality while considering the granules of physically existing or potentially derivable attributes. We also show several examples of utilization of the granulation structures defined over the feature spaces in the feature selection algorithms. As a case study, we consider the algorithms developed within the theory of rough sets, aimed at finding irreducible subsets of attributes that are sufficient to distinguish between the cases belonging to different target decision classes.
Marek Grzegorowski, Andrzej Janusz, Dominik Slezak, Marcin S. Szczuka
IEEE BigData1
2017 Predicting seismic events in coal mines based on underground sensor measurements
Andrzej Janusz, Marek Grzegorowski, Marcin Michalak 0001, Lukasz Wróbel, Marek Sikora, Dominik Slezak
Eng. Appl. Artif. Intell.2
2016 Massively Parallel Feature Extraction Framework Application in Predicting Dangerous Seismic Events
abstract
In this paper we introduce an automated mechanism for knowledge discovery from data streams.As a part of this work, we also present a new approach to the creation of classifiers ensemble based on a wide variety of models.Furthermore, we describe an innovative, highly scalable feature extraction and selection framework designed to work with the MapReduce programming model and the application of designed framework to build an ensemble of classifiers which takes into account both the quality and the diversity of individual models.The effectiveness of the solution has been verified through a participation in an open data mining competition which concerned the problem of predicting periods of increased seismic activity causing lifethreatening accidents in coal mines.The submitted solution obtained the highest AUC score of all the solutions uploaded by 106 participating research teams.
Marek Grzegorowski
FedCSIS1
2015 Window-based feature extraction framework for multi-sensor data: A posture recognition case study
abstract
The article introduces a novel mechanism for automatic extraction of features from streams of numerical data. It was originally designed for the purpose of processing multiple streams of readings generated by sensors in coal mines. The original research was conducted on methane concentration analysis in the DISESOR project. The article demonstrates an application of the elaborated mechanism for the case of tagging short series of readings from sensors that monitor activities and movements of firefighters during the action with labels corresponding to firefighter activities. The purpose of the experiment was to assess how the automatic feature extraction and construction of classifiers (without parameters tuning and without the use of classifier ensembles) can cope with the competition's task in comparison to other participants.
Marek Grzegorowski, Sebastian Stawicki
FedCSIS1