Dominik Slezak

dblp:53/3187 · DBLP profile ↗
← Back
39ranked-venue papers in the field
10as first author
11since 2021 · last 2025
0000-0003-2453-4974ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 17 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 12 (3 first)Other / Interdisciplinary · 5 (1 first)Database Systems & Data Management · 2 (2 first)Data Mining & Knowledge Discovery · 2 (1 first)Information Retrieval & Web Search · 1 (1 first)
YearPublicationVenuePosition
2025 Information Granulation for Hierarchical Feature Selection in Detection of Anomalies in IoT Devices
Lukasz Wawrowski, Konrad Chwelatiuk, Marcin Michalak 0001, Iwona Kostorz, Dominik Slezak, Piotr Biczyk, Blazej Adamczyk, Maksym Brzeczek
IEEE Big Data5
2024 Do Data Scientists Dream About Their Skills' Assessment? - Transforming a Competition Platform Into an Assessment Platform
abstract
We present a platform for automatic assessment of technical data science skills (hard skills) and competencies that help to apply those technical skills in practice (soft skills). The platform serves so-called assessment platform tasks that resemble data-focused tasks typical for online data science competitions. Actually, these tasks are designed based on international conference competitions that we have been organizing on our competition platform knowledgepit.ai. The main idea relies on our observation that the given person’s behavior during a competition (dynamics of submitting solutions, activity on competition forum, etc.) can be correlated with his/her soft competencies. Accordingly, our goal is to translate the specifics of international conference competitions into a framework that investigates the behaviors of data science job candidates, employees who want to improve their work in data science projects, students who wish to build their future professional careers on AI solutions, etc. We claim that such behaviors – if properly measured based on solving our assessment platform tasks – can effectively indicate both hard and soft types of competencies.
Dominik Slezak, Andrzej Janusz, Maciej Swiechowski, Agnieszka Chadzynska-Krasowska, Jacek Kaminski
IEEE Big Data1
2024 IEEE Big Data Cup 2024 Report: Predicting Chess Puzzle Difficulty at KnowledgePit.ai
abstract
We summarize the results of the IEEE BigData 2024 Cup: Predicting Chess Puzzle Difficulty – a data science competition organized at the knowledgepit.ai platform in association with the IEEE BigData 2024 conference. We describe the competition goal and tie it to existing research on human-computer interaction, focusing on task difficulty estimation and aligning human and AI behavior. We explain how we acquired and processed the data, separately for training and testing datasets. We review submitted solutions and evaluate their performance by comparing them to a simple benchmark. We explain how the achieved rating differences translate to user experience when solving chess puzzles. We further explore the concept of chess puzzle difficulty by replicating competition results with puzzle ratings obtained using chess bots only. We conclude with a summary of our findings and directions for future studies, as well as an invitation to the next edition of the competition in 2025.
Jan Zysko, Maciej Swiechowski, Sebastian Stawicki, Katarzyna Jagiela, Andrzej Janusz, Dominik Slezak
IEEE Big Data6
2023 IEEE BigData Cup 2023 Report: Object Recognition with Muon Tomography Using Cosmic Rays
abstract
We summarize the results of the IEEE BigData 2023 Cup: Object Recognition with Muon Tomography using Cosmic Rays - a data mining competition organized at the KnowledgePit.ai platform in association with the IEEE BigData 2023 conference. We describe the challenge at the heart of the competition task, as well as the data acquisition and preparation steps. We present the entire process of preparing experiments and subsequent data analysis for the purpose of recognizing X-rayed objects using muon tomography techniques. We conclude this analysis by presenting the baseline as well as the winning solution of the object segmentation algorithms for the research space reconstruction and object classification.
Mateusz Wnuk, Jan Dziuba, Andrzej Janusz, Dominik Slezak
IEEE Big Data4
2023 A practical study of methods for deriving insightful attribute importance rankings using decision bireducts
abstract
Subject matter experts (SMEs) often rely on attribute importance rankings to verify machine learning models, acquire insights into their outcomes, and gain a deeper understanding of the investigated phenomena. To further increase their usefulness, we introduce a new approach to the evaluation of attribute rankings produced by any machine learning method . As a real-world case study , we investigate the attribute importance scores produced using XGBoost and decision bireducts on the data gathered by an HR company, where the goal is to predict the willingness of candidates to change their job. For this task, XGBoost delivers accurate models but fails to identify many attributes that are important to SMEs. In comparison, decision bireducts lead to models that are easier to interpret and explore the data with a higher focus on the diversity of attributes. The ensembles of decision bireducts deliver comparable accuracy and their associated attribute rankings are more insightful than those of XGBoost.
Andrzej Janusz, Dominik Slezak, Sebastian Stawicki, Krzysztof Stencel
Inf. Sci.2
2022 Tensor-based Approach to Big Data Processing and Machine Learning
abstract
We present an approach to tensor compression and decomposition, as well as to a design of data processing algorithms on the top of them. Our implementation uses the popular scalable data processing framework Apache Parquet to effectively store the data. This library does not directly store tensors as native data types, but we slightly changed its implementation for our purpose using its specific data storage format and extending it with additional compression. We summarize the performance of tensor storage, as well as the effectiveness of multiple machine learning methods and their hyperparameter tuning.
Maciej Bartoszuk, Jaroslaw Litwin, Mateusz Wnuk, Dominik Slezak
IEEE Big Data4
2022 EVEAL - Expected Variance Estimation for Active Learning
abstract
Regression problems frequently occur in the surrounding world, therefore are unavoidable in real-world applications. However, to obtain a model with desired generalization performance, a vast amount of labels usually is required. In many scenarios, obtaining unlabelled data is relatively inexpensive, therefore active learning approaches may be used to reduce the needed annotation effort. Most uncertainty-based regression active learning algorithms use variance estimation of model predictions to choose informative samples. Those algorithms do not incorporate knowledge about the data distribution for the given task. In this paper, we propose a novel algorithm to incorporate information about data distribution and combine it with variance estimation as an informativeness function. Experiments conducted on four data sets show that the proposed approach outperforms standard variance-based sampling by a margin, and indicate its robustness.
Daniel Kaluza, Andrzej Janusz, Dominik Slezak
IEEE Big Data3
2022 Approximation of the expectation-maximization algorithm for Gaussian mixture models on big data
abstract
Gaussian mixture models are a very useful tool for modeling data distribution. While estimating parameters using the expectation-maximization algorithm, this approach does not scale well with big datasets, especially if it is necessary to prepare many models for the proper selection of metaparameters. In this article we present an approximation of the expectation-maximization algorithm obtained by merging crucial subsets of the dataset, that differ slightly in their effect on the expectation-maximization loss function, into information granules. Furthermore, application examples comparing new method with the classical approach are shown.
Mateusz Przyborowski, Dominik Slezak
IEEE Big Data2
2022 IEEE BigData Cup 2022 Report Privacy-preserving Matching of Encrypted Images
abstract
We summarize the results of IEEE BigData 2022 Cup: Privacy-preserving Matching of Encrypted Images - a data mining challenge organized at the KnowledgePit platform in association with the IEEE BigData 2022 conference. We describe the challenge in the hearth of the competition task, as well as the data acquisition and preparation steps. We also provide a brief overview of the top-performing solutions submitted by participants. Finally, we present results of the post-competition data analysis, in which we consider the similarity of solutions submitted by various teams in terms of their errors on the test data. We conclude this analysis by discussion of the significance and impact of the competition results on the underlying problem of constructing effective and efficient anonymization algorithms for monitoring in the DOOH advertising industry.
Marcin S. Szczuka, Andrzej Janusz, Boguslaw Cyganek, Jakub Grabek, Lukasz Przebinda, Andzelika Zalewska-Küpçü, Andrzej Bukala, Dominik Slezak
IEEE Big Data8
2022 Learning multimodal entity representations and their ensembles, with applications in a data-driven advisory framework for video game players
Andrzej Janusz, Daniel Kaluza, Maciej Matraszek, Lukasz Grad, Maciej Swiechowski, Dominik Slezak
Inf. Sci.6
2021 Predicting Victories in Video Games - IEEE BigData 2021 Cup Report
abstract
We summarize the results of IEEE BigData 2021 Cup: Predicting Victories in Video Games - a data mining challenge organized at the KnowledgePit platform in association with the IEEE BigData 2021 conference. We describe the competition task, as well as the data acquisition and preprocessing steps. We also provide a brief overview of the top-performing solutions submitted by participants. Finally, we present results of the post-competition data analysis, in which we consider the similarity of solutions submitted by various teams in terms of their errors on the test data. We conclude this analysis by demonstrating a method for constructing an ensemble of submitted solutions. Such an ensemble performs better than any of the individual solutions submitted during the competition.
Maciej Matraszek, Andrzej Janusz, Maciej Swiechowski, Dominik Slezak
IEEE BigData4
2020 Predicting Escalations in Customer Support: Analysis of Data Mining Challenge Results
abstract
We summarize IEEE Big Data Cup: Predicting Escalations in Customer Support - a data mining competition organized jointly by companies Information Builders and QED Software at the KnowledgePit platform, in the frame of the 2020 IEEE International Conference on Big Data. We discuss the motivation for organizing this event and highlight the factors that make it such a challenging topic. We describe the data provided to participants and formulate the competition task. We also provide an overview of competition results with a detailed analysis of a few selected solutions. Finally, we present a novel functionality of the KnowledgePit platform - an analytic module that allows organizers to investigate selected solutions using a convenient GUI and provides in-depth insights about their quality.
Andrzej Janusz, Guohua Hao, Daniel Kaluza, Tony Li, Robert Wojciechowski, Dominik Slezak
IEEE BigData6
2020 Reinventing Infobright's Concept of Rough Calculations on Granulated Tables for the Purpose of Accelerating Modern Data Processing Frameworks
abstract
We present an approach to data and information granulation known from the Infobright Community Edition (ICE) analytical database engine, now reimplemented within the two popular scalable data processing frameworks: Apache Parquet and ROOT. Both of these libraries, do not directly realize the idea of resolving queries based on rough-set-driven calculations on granulated data statistics, what was one of the biggest accelerators in ICE. We summarize the implementation of such level of operations and compare the performance of analytical SQL queries in ROOT, Parquet, and ICE.
Mateusz Wnuk, Sebastian Stawicki, Dominik Slezak
IEEE BigData3
2019 IEEE BigData 2019 Cup: Suspicious Network Event Recognition
abstract
“IEEE BigData 2019 Cup: Suspicious Network Event Recognition” was a data mining competition organized jointly by companies Security On-Demand and QED Software at the KnowledgePit online platform, in association with the IEEE BigData 2019 conference. The scope of this challenge referred to the notions of cybersecurity analytics and network alert evaluation. In this paper, we summarize the results of our competition. We explain how data sets had been prepared before it was possible to make them available to competition participants. We describe the baseline scoring models that we designed as a reference for participants, and we demonstrate how critical for their performance was the aspect of appropriate feature engineering. We also discuss the results of experiments conducted to verify the (un)suitability of deep recurrent neural networks in this particular case. In some sense, we show that there are no “perfect” machine learning approaches that could be applied equally successfully to every data science undertaking.
Andrzej Janusz, Daniel Kaluza, Agnieszka Chadzynska-Krasowska, Bartek Konarski, Joel Holland, Dominik Slezak
IEEE BigData6
2019 On resilient feature selection: Computational foundations of r-C-reducts
Marek Grzegorowski, Dominik Slezak
Inf. Sci.2
2018 Toward Machine Learning on Granulated Data - a Case of Compact Autoencoder-based Representations of Satellite Images
abstract
We consider a problem of learning from compact representations of images for a purpose of object recognition and content-based image retrieval. We discuss a motivation for using compressed images in those tasks and indicate exemplary applications related to analysis on the data from satellites. Finally, we show some preliminary results of experiments conducted to demonstrate the impact of the image data granulation on the quality of classification. We empirically compare the performance of prediction models trained on original images, images compressed using autoencoders, and on images whose quality was lowered in order to reduce their size.
Mateusz Przyborowski, Tomasz Tajmajer, Lukasz Grad, Andrzej Janusz, Piotr Biczyk, Dominik Slezak
IEEE BigData6
2018 Similarity-based Detection of Fertile Days at OvuFriend
abstract
We discuss recent AI-related developments at OvuFriend's online platform which is designed to assist families in overcoming infertility problems. One of functionalities of the platform is to detect fertile days basing on often incomplete and uncertain data provided by the users. Besides discussing the particular layers of the underlying OvuFriend's system architecture, we concentrate on one of the proposed fertile day detection models which is based on the idea of utilizing multivariate similarities between the current cycle and the past cycles available for the given user or for users who have a similar profile, with an additional self-checking procedure that enables the algorithm to neglect insufficiently reliable inputs.
Lukasz Sosnowski, Wojciech Chaber, Lukasz Milobedzki, Tomasz Penza, Jadwiga Sosnowska, Karol Zaleski, Joanna Fedorowicz, Iwona Szymusik, Dominik Slezak
IEEE BigData9
2018 How to Match Jobs and Candidates - A Recruitment Support System Based on Feature Engineering and Advanced Analytics
Andrzej Janusz, Sebastian Stawicki, Michal Drewniak, Krzysztof Ciebiera, Dominik Slezak, Krzysztof Stencel
IPMU (2)5
2018 SENSEI: An Intelligent Advisory System for the eSport Community and Casual Players
abstract
In this article, we describe the SENSEI system. It helps players to improve their skills in popular eSports games. We discuss the main goals of the system and explain the associated challenges. We also present its conceptual architecture which aims at enabling full automation of the data acquisition and analytic processes. The system is expected to provide in-depth analytics of players' performance and give practical advice regarding possible improvements. Thus its architecture allows players to provide feedback and manually label important concepts. Finally, we discuss our first case study - an advisory system for popular collectible card video games.
Andrzej Janusz, Dominik Slezak, Sebastian Stawicki, Krzysztof Stencel
WI2
2018 Grail: A Framework for Adaptive and Believable AI in Video Games
abstract
We describe a framework - called Grail - which aims at providing developers with tools for implementing AI in games. There is a whole variety of games and the role of AI in them can vary from case to case. Thus, the main challenge is to create a system allowing for meeting various design goals with relatively easy to use interfaces. We present the conceptual architecture of Grail and algorithms chosen by us to cover a wide spectrum of use cases: Planning, Utility System, Simplified Games with Tree Search and scripting. We believe that together they fulfill the requirements of a flexible AI engine.
Maciej Swiechowski, Dominik Slezak
WI2
2018 Bireducts with tolerance relations
María José Benítez-Caballero, Jesús Medina 0001, Eloísa Ramírez-Poussa, Dominik Slezak
Inf. Sci.4
2018 A framework for learning and embedding multi-sensor forecasting models into a decision support system: A case study of methane concentration in coal mines
Dominik Slezak, Marek Grzegorowski, Andrzej Janusz, Michal Kozielski, Sinh Hoa Nguyen, Marek Sikora, Sebastian Stawicki, Lukasz Wróbel
Inf. Sci.1
2018 A new approximate query engine based on intelligent capture and fast transformations of granulated data summaries
abstract
We outline the processes of intelligent creation and utilization of granulated data summaries in the engine aimed at fast approximate execution of analytical SQL statements. We discuss how to use the introduced engine for the purposes of ad-hoc data exploration over large and quickly increasing data collected in a heterogeneous or distributed fashion. We focus on mechanisms that transform input data summaries into result sets representing query outcomes. We also illustrate how our computational principles can be put together with other paradigms of scaling and harnessing data analytics.
Dominik Slezak, Rick Glick, Pawel Betlinski, Piotr Synak
J. Intell. Inf. Syst.1
2017 On the role of feature space granulation in feature selection processes
abstract
Information granulation plays an important role in the process of scaling up modern machine learning and knowledge discovery algorithms. By employing compact descriptions of granules - whereby granules are defined as collections of original data elements gathered together by means of their similarity, proximity or functionality - one can drastically accelerate computations and, moreover, make the results of those computations more meaningful for domain experts. In this paper, we summarize some of the feature space granulation approaches introduced by now. We discuss the meaning of similarity, proximity and functionality while considering the granules of physically existing or potentially derivable attributes. We also show several examples of utilization of the granulation structures defined over the feature spaces in the feature selection algorithms. As a case study, we consider the algorithms developed within the theory of rough sets, aimed at finding irreducible subsets of attributes that are sufficient to distinguish between the cases belonging to different target decision classes.
Marek Grzegorowski, Andrzej Janusz, Dominik Slezak, Marcin S. Szczuka
IEEE BigData3
2017 Scalable cyber-security analytics with a new summary-based approximate query engine
abstract
A growing need for scalable solutions for both machine learning and interactive analytics exists in the area of cyber-security. Machine learning aims at segmentation and classification of log events, which leads towards optimization of the threat monitoring processes. The tools for interactive analytics are required to resolve the uncertain cases, whereby machine learning algorithms are not able to provide a convincing outcome and human expertise is necessary. In this paper we focus on a case study of a security operations platform, whereby typical layers of information processing are integrated with a new database engine dedicated to approximate analytics. The engine makes it possible for the security experts to query massive log event data sets in a standard relational style. The query outputs are received orders of magnitude faster than any of the existing database solutions running with comparable resources and, in addition, they are sufficiently accurate to make the right decisions about suspicious corner cases. The engine internals are driven by the principles of information granulation and summary-based processing. They also refer to the ideas of data quantization, approximate computing, rough sets and probability propagation. In the paper we study how the engine's parameters can influence its performance within the considered environment. In addition to the results of experiments conducted on large data sets, we also discuss some of our high level design decisions including the choice of an approximate query result accuracy measure that should reflect the specifics of the considered threat monitoring operations.
Dominik Slezak, Agnieszka Chadzynska-Krasowska, Joel Holland, Piotr Synak, Rick Glick, Marcin Perkowski
IEEE BigData1
2015 Granular modeling with fuzzy comparators
abstract
We present an overview of an approach to solving various real-life tasks related to Computational Intelligence by means of modeling with information granules. The particular methods of building the model are based on networks of fuzzy comparators. We demonstrate that comparator networks are a powerful and versatile tool suitable for applications. The introduction of methodology is accompanied with brief presentation of its formal basis. We also list existing and prospective, practical applications of the described approach.
Lukasz Sosnowski, Marcin S. Szczuka, Dominik Slezak
IEEE BigData3
2014 Implementing algorithms of rough set theory and fuzzy rough set theory in the R package "RoughSets"
Lala Septem Riza, Andrzej Janusz, Christoph Bergmeir, Chris Cornelis, Francisco Herrera, Dominik Slezak, José Manuel Benítez 0001
Inf. Sci.6
2014 Processing and mining complex data streams
Jerzy Stefanowski, Alfredo Cuzzocrea, Dominik Slezak
Inf. Sci.3
2012 Management of Information Incompleteness in Rough Non-deterministic Information Analysis
Hiroshi Sakai, Michinori Nakata, Dominik Slezak
IPMU (1)3
2012 Rough SQL - Semantics and Execution
Dominik Slezak, Piotr Synak, Graham Toppin, Jakub Wroblewski, Janusz Borkowski
IPMU (2)1
2010 Injecting domain knowledge into a granular database engine: a position paper
abstract
We discuss how to use techniques from such fields as text processing and knowledge management to better handle text attributes in the Infobright's RDBMS engine. Our approach leads to a rich interface for domain experts who wish to share their knowledge about data content and, on the other hand, it remains unnoticeable to data users. It enables to improve data storage, data access, and data compression, with no changes required at the database schema level.
Dominik Slezak, Graham Toppin
CIKM1
2010 Attribute selection with fuzzy decision reducts
Chris Cornelis, Richard Jensen, Germán Hurtado Martín, Dominik Slezak
Inf. Sci.4
2010 Introduction to the special issue on advanced information retrieval and databases
Aijun An, Dominik Slezak
J. Intell. Inf. Syst.2
2009 Data warehouse technology by infobright
abstract
We discuss Infobright technology with respect to its main features and architectural differentiators. We introduce the upcoming research and development projects that may be of special interest to the academic and industry communities.
Dominik Slezak, Victoria Eastwood
SIGMOD Conference1
2009 Degrees of conditional (in)dependence: A framework for approximate Bayesian networks and examples related to the rough set-based feature selection
Dominik Slezak
Inf. Sci.1
2008 Brighthouse: an analytic data warehouse for ad-hoc queries
abstract
Brighthouse is a column-oriented data warehouse with an automatically tuned, ultra small overhead metadata layer called Knowledge Grid, that is used as an alternative to classical indexes. The advantages of column-oriented data storage, as well as data compression have already been well-documented, especially in the context of analytic, decision support querying. This paper demonstrates additional benefits resulting from Knowledge Grid for compressed, column-oriented databases. In particular, we explain how it assists in query optimization and execution, by minimizing the need of data reads and data decompression.
Dominik Slezak, Jakub Wroblewski, Victoria Eastwood, Piotr Synak
Proc. VLDB Endow.1
2003 Application of Temporal Descriptors to Musical Instrument Sound Recognition
Alicja Wieczorkowska, Jakub Wroblewski, Piotr Synak, Dominik Slezak
J. Intell. Inf. Syst.4
1999 Classification Algorithms Based on Linear Combinations of Features
Dominik Slezak, Jakub Wroblewski
PKDD1
1997 Neural Networks Design: Rough Set Approach to Continuous Data
Hung Son Nguyen, Marcin S. Szczuka, Dominik Slezak
PKDD3