EDBT 2026 Demo / reviewers in the wild / expert
Dominik Slezak
dblp:53/3187
· DBLP profile ↗
39ranked-venue papers in the field
10as first author
11since 2021 · last 2025
0000-0003-2453-4974ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 17 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 12 (3 first)Other / Interdisciplinary · 5 (1 first)Database Systems & Data Management · 2 (2 first)Data Mining & Knowledge Discovery · 2 (1 first)Information Retrieval & Web Search · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Information Granulation for Hierarchical Feature Selection in Detection of Anomalies in IoT Devices
Lukasz Wawrowski, Konrad Chwelatiuk, Marcin Michalak 0001, Iwona Kostorz, Dominik Slezak, Piotr Biczyk, Blazej Adamczyk, Maksym Brzeczek |
IEEE Big Data | 5 |
| 2024 | Do Data Scientists Dream About Their Skills' Assessment? - Transforming a Competition Platform Into an Assessment PlatformabstractWe present a platform for automatic assessment of technical data science skills (hard skills) and competencies that help to apply those technical skills in practice (soft skills). The platform serves so-called assessment platform tasks that resemble data-focused tasks typical for online data science competitions. Actually, these tasks are designed based on international conference competitions that we have been organizing on our competition platform knowledgepit.ai. The main idea relies on our observation that the given person’s behavior during a competition (dynamics of submitting solutions, activity on competition forum, etc.) can be correlated with his/her soft competencies. Accordingly, our goal is to translate the specifics of international conference competitions into a framework that investigates the behaviors of data science job candidates, employees who want to improve their work in data science projects, students who wish to build their future professional careers on AI solutions, etc. We claim that such behaviors – if properly measured based on solving our assessment platform tasks – can effectively indicate both hard and soft types of competencies. Dominik Slezak, Andrzej Janusz, Maciej Swiechowski, Agnieszka Chadzynska-Krasowska, Jacek Kaminski |
IEEE Big Data | 1 |
| 2024 | IEEE Big Data Cup 2024 Report: Predicting Chess Puzzle Difficulty at KnowledgePit.aiabstractWe summarize the results of the IEEE BigData 2024 Cup: Predicting Chess Puzzle Difficulty – a data science competition organized at the knowledgepit.ai platform in association with the IEEE BigData 2024 conference. We describe the competition goal and tie it to existing research on human-computer interaction, focusing on task difficulty estimation and aligning human and AI behavior. We explain how we acquired and processed the data, separately for training and testing datasets. We review submitted solutions and evaluate their performance by comparing them to a simple benchmark. We explain how the achieved rating differences translate to user experience when solving chess puzzles. We further explore the concept of chess puzzle difficulty by replicating competition results with puzzle ratings obtained using chess bots only. We conclude with a summary of our findings and directions for future studies, as well as an invitation to the next edition of the competition in 2025. Jan Zysko, Maciej Swiechowski, Sebastian Stawicki, Katarzyna Jagiela, Andrzej Janusz, Dominik Slezak |
IEEE Big Data | 6 |
| 2023 | IEEE BigData Cup 2023 Report: Object Recognition with Muon Tomography Using Cosmic RaysabstractWe summarize the results of the IEEE BigData 2023 Cup: Object Recognition with Muon Tomography using Cosmic Rays - a data mining competition organized at the KnowledgePit.ai platform in association with the IEEE BigData 2023 conference. We describe the challenge at the heart of the competition task, as well as the data acquisition and preparation steps. We present the entire process of preparing experiments and subsequent data analysis for the purpose of recognizing X-rayed objects using muon tomography techniques. We conclude this analysis by presenting the baseline as well as the winning solution of the object segmentation algorithms for the research space reconstruction and object classification. Mateusz Wnuk, Jan Dziuba, Andrzej Janusz, Dominik Slezak |
IEEE Big Data | 4 |
| 2023 | A practical study of methods for deriving insightful attribute importance rankings using decision bireductsabstractSubject matter experts (SMEs) often rely on attribute importance rankings to verify machine learning models, acquire insights into their outcomes, and gain a deeper understanding of the investigated phenomena. To further increase their usefulness, we introduce a new approach to the evaluation of attribute rankings produced by any machine learning method . As a real-world case study , we investigate the attribute importance scores produced using XGBoost and decision bireducts on the data gathered by an HR company, where the goal is to predict the willingness of candidates to change their job. For this task, XGBoost delivers accurate models but fails to identify many attributes that are important to SMEs. In comparison, decision bireducts lead to models that are easier to interpret and explore the data with a higher focus on the diversity of attributes. The ensembles of decision bireducts deliver comparable accuracy and their associated attribute rankings are more insightful than those of XGBoost. Andrzej Janusz, Dominik Slezak, Sebastian Stawicki, Krzysztof Stencel |
Inf. Sci. | 2 |
| 2022 | Tensor-based Approach to Big Data Processing and Machine LearningabstractWe present an approach to tensor compression and decomposition, as well as to a design of data processing algorithms on the top of them. Our implementation uses the popular scalable data processing framework Apache Parquet to effectively store the data. This library does not directly store tensors as native data types, but we slightly changed its implementation for our purpose using its specific data storage format and extending it with additional compression. We summarize the performance of tensor storage, as well as the effectiveness of multiple machine learning methods and their hyperparameter tuning. Maciej Bartoszuk, Jaroslaw Litwin, Mateusz Wnuk, Dominik Slezak |
IEEE Big Data | 4 |
| 2022 | EVEAL - Expected Variance Estimation for Active LearningabstractRegression problems frequently occur in the surrounding world, therefore are unavoidable in real-world applications. However, to obtain a model with desired generalization performance, a vast amount of labels usually is required. In many scenarios, obtaining unlabelled data is relatively inexpensive, therefore active learning approaches may be used to reduce the needed annotation effort. Most uncertainty-based regression active learning algorithms use variance estimation of model predictions to choose informative samples. Those algorithms do not incorporate knowledge about the data distribution for the given task. In this paper, we propose a novel algorithm to incorporate information about data distribution and combine it with variance estimation as an informativeness function. Experiments conducted on four data sets show that the proposed approach outperforms standard variance-based sampling by a margin, and indicate its robustness. Daniel Kaluza, Andrzej Janusz, Dominik Slezak |
IEEE Big Data | 3 |
| 2022 | Approximation of the expectation-maximization algorithm for Gaussian mixture models on big dataabstractGaussian mixture models are a very useful tool for modeling data distribution. While estimating parameters using the expectation-maximization algorithm, this approach does not scale well with big datasets, especially if it is necessary to prepare many models for the proper selection of metaparameters. In this article we present an approximation of the expectation-maximization algorithm obtained by merging crucial subsets of the dataset, that differ slightly in their effect on the expectation-maximization loss function, into information granules. Furthermore, application examples comparing new method with the classical approach are shown. Mateusz Przyborowski, Dominik Slezak |
IEEE Big Data | 2 |
| 2022 | IEEE BigData Cup 2022 Report Privacy-preserving Matching of Encrypted ImagesabstractWe summarize the results of IEEE BigData 2022 Cup: Privacy-preserving Matching of Encrypted Images - a data mining challenge organized at the KnowledgePit platform in association with the IEEE BigData 2022 conference. We describe the challenge in the hearth of the competition task, as well as the data acquisition and preparation steps. We also provide a brief overview of the top-performing solutions submitted by participants. Finally, we present results of the post-competition data analysis, in which we consider the similarity of solutions submitted by various teams in terms of their errors on the test data. We conclude this analysis by discussion of the significance and impact of the competition results on the underlying problem of constructing effective and efficient anonymization algorithms for monitoring in the DOOH advertising industry. Marcin S. Szczuka, Andrzej Janusz, Boguslaw Cyganek, Jakub Grabek, Lukasz Przebinda, Andzelika Zalewska-Küpçü, Andrzej Bukala, Dominik Slezak |
IEEE Big Data | 8 |
| 2022 | Learning multimodal entity representations and their ensembles, with applications in a data-driven advisory framework for video game players
Andrzej Janusz, Daniel Kaluza, Maciej Matraszek, Lukasz Grad, Maciej Swiechowski, Dominik Slezak |
Inf. Sci. | 6 |
| 2021 | Predicting Victories in Video Games - IEEE BigData 2021 Cup ReportabstractWe summarize the results of IEEE BigData 2021 Cup: Predicting Victories in Video Games - a data mining challenge organized at the KnowledgePit platform in association with the IEEE BigData 2021 conference. We describe the competition task, as well as the data acquisition and preprocessing steps. We also provide a brief overview of the top-performing solutions submitted by participants. Finally, we present results of the post-competition data analysis, in which we consider the similarity of solutions submitted by various teams in terms of their errors on the test data. We conclude this analysis by demonstrating a method for constructing an ensemble of submitted solutions. Such an ensemble performs better than any of the individual solutions submitted during the competition. Maciej Matraszek, Andrzej Janusz, Maciej Swiechowski, Dominik Slezak |
IEEE BigData | 4 |
| 2020 | Predicting Escalations in Customer Support: Analysis of Data Mining Challenge ResultsabstractWe summarize IEEE Big Data Cup: Predicting Escalations in Customer Support - a data mining competition organized jointly by companies Information Builders and QED Software at the KnowledgePit platform, in the frame of the 2020 IEEE International Conference on Big Data. We discuss the motivation for organizing this event and highlight the factors that make it such a challenging topic. We describe the data provided to participants and formulate the competition task. We also provide an overview of competition results with a detailed analysis of a few selected solutions. Finally, we present a novel functionality of the KnowledgePit platform - an analytic module that allows organizers to investigate selected solutions using a convenient GUI and provides in-depth insights about their quality. Andrzej Janusz, Guohua Hao, Daniel Kaluza, Tony Li, Robert Wojciechowski, Dominik Slezak |
IEEE BigData | 6 |
| 2020 | Reinventing Infobright's Concept of Rough Calculations on Granulated Tables for the Purpose of Accelerating Modern Data Processing FrameworksabstractWe present an approach to data and information granulation known from the Infobright Community Edition (ICE) analytical database engine, now reimplemented within the two popular scalable data processing frameworks: Apache Parquet and ROOT. Both of these libraries, do not directly realize the idea of resolving queries based on rough-set-driven calculations on granulated data statistics, what was one of the biggest accelerators in ICE. We summarize the implementation of such level of operations and compare the performance of analytical SQL queries in ROOT, Parquet, and ICE. Mateusz Wnuk, Sebastian Stawicki, Dominik Slezak |
IEEE BigData | 3 |
| 2019 | IEEE BigData 2019 Cup: Suspicious Network Event Recognitionabstract“IEEE BigData 2019 Cup: Suspicious Network Event Recognition” was a data mining competition organized jointly by companies Security On-Demand and QED Software at the KnowledgePit online platform, in association with the IEEE BigData 2019 conference. The scope of this challenge referred to the notions of cybersecurity analytics and network alert evaluation. In this paper, we summarize the results of our competition. We explain how data sets had been prepared before it was possible to make them available to competition participants. We describe the baseline scoring models that we designed as a reference for participants, and we demonstrate how critical for their performance was the aspect of appropriate feature engineering. We also discuss the results of experiments conducted to verify the (un)suitability of deep recurrent neural networks in this particular case. In some sense, we show that there are no “perfect” machine learning approaches that could be applied equally successfully to every data science undertaking. Andrzej Janusz, Daniel Kaluza, Agnieszka Chadzynska-Krasowska, Bartek Konarski, Joel Holland, Dominik Slezak |
IEEE BigData | 6 |
| 2019 | On resilient feature selection: Computational foundations of r-C-reducts
Marek Grzegorowski, Dominik Slezak |
Inf. Sci. | 2 |
| 2018 | Toward Machine Learning on Granulated Data - a Case of Compact Autoencoder-based Representations of Satellite ImagesabstractWe consider a problem of learning from compact representations of images for a purpose of object recognition and content-based image retrieval. We discuss a motivation for using compressed images in those tasks and indicate exemplary applications related to analysis on the data from satellites. Finally, we show some preliminary results of experiments conducted to demonstrate the impact of the image data granulation on the quality of classification. We empirically compare the performance of prediction models trained on original images, images compressed using autoencoders, and on images whose quality was lowered in order to reduce their size. Mateusz Przyborowski, Tomasz Tajmajer, Lukasz Grad, Andrzej Janusz, Piotr Biczyk, Dominik Slezak |
IEEE BigData | 6 |
| 2018 | Similarity-based Detection of Fertile Days at OvuFriendabstractWe discuss recent AI-related developments at OvuFriend's online platform which is designed to assist families in overcoming infertility problems. One of functionalities of the platform is to detect fertile days basing on often incomplete and uncertain data provided by the users. Besides discussing the particular layers of the underlying OvuFriend's system architecture, we concentrate on one of the proposed fertile day detection models which is based on the idea of utilizing multivariate similarities between the current cycle and the past cycles available for the given user or for users who have a similar profile, with an additional self-checking procedure that enables the algorithm to neglect insufficiently reliable inputs. Lukasz Sosnowski, Wojciech Chaber, Lukasz Milobedzki, Tomasz Penza, Jadwiga Sosnowska, Karol Zaleski, Joanna Fedorowicz, Iwona Szymusik, Dominik Slezak |
IEEE BigData | 9 |
| 2018 | How to Match Jobs and Candidates - A Recruitment Support System Based on Feature Engineering and Advanced Analytics
Andrzej Janusz, Sebastian Stawicki, Michal Drewniak, Krzysztof Ciebiera, Dominik Slezak, Krzysztof Stencel |
IPMU (2) | 5 |
| 2018 | SENSEI: An Intelligent Advisory System for the eSport Community and Casual PlayersabstractIn this article, we describe the SENSEI system. It helps players to improve their skills in popular eSports games. We discuss the main goals of the system and explain the associated challenges. We also present its conceptual architecture which aims at enabling full automation of the data acquisition and analytic processes. The system is expected to provide in-depth analytics of players' performance and give practical advice regarding possible improvements. Thus its architecture allows players to provide feedback and manually label important concepts. Finally, we discuss our first case study - an advisory system for popular collectible card video games. Andrzej Janusz, Dominik Slezak, Sebastian Stawicki, Krzysztof Stencel |
WI | 2 |
| 2018 | Grail: A Framework for Adaptive and Believable AI in Video GamesabstractWe describe a framework - called Grail - which aims at providing developers with tools for implementing AI in games. There is a whole variety of games and the role of AI in them can vary from case to case. Thus, the main challenge is to create a system allowing for meeting various design goals with relatively easy to use interfaces. We present the conceptual architecture of Grail and algorithms chosen by us to cover a wide spectrum of use cases: Planning, Utility System, Simplified Games with Tree Search and scripting. We believe that together they fulfill the requirements of a flexible AI engine. Maciej Swiechowski, Dominik Slezak |
WI | 2 |
| 2018 | Bireducts with tolerance relations
María José Benítez-Caballero, Jesús Medina 0001, Eloísa Ramírez-Poussa, Dominik Slezak |
Inf. Sci. | 4 |
| 2018 | A framework for learning and embedding multi-sensor forecasting models into a decision support system: A case study of methane concentration in coal mines
Dominik Slezak, Marek Grzegorowski, Andrzej Janusz, Michal Kozielski, Sinh Hoa Nguyen, Marek Sikora, Sebastian Stawicki, Lukasz Wróbel |
Inf. Sci. | 1 |
| 2018 | A new approximate query engine based on intelligent capture and fast transformations of granulated data summariesabstractWe outline the processes of intelligent creation and utilization of granulated data summaries in the engine aimed at fast approximate execution of analytical SQL statements. We discuss how to use the introduced engine for the purposes of ad-hoc data exploration over large and quickly increasing data collected in a heterogeneous or distributed fashion. We focus on mechanisms that transform input data summaries into result sets representing query outcomes. We also illustrate how our computational principles can be put together with other paradigms of scaling and harnessing data analytics. Dominik Slezak, Rick Glick, Pawel Betlinski, Piotr Synak |
J. Intell. Inf. Syst. | 1 |
| 2017 | On the role of feature space granulation in feature selection processesabstractInformation granulation plays an important role in the process of scaling up modern machine learning and knowledge discovery algorithms. By employing compact descriptions of granules - whereby granules are defined as collections of original data elements gathered together by means of their similarity, proximity or functionality - one can drastically accelerate computations and, moreover, make the results of those computations more meaningful for domain experts. In this paper, we summarize some of the feature space granulation approaches introduced by now. We discuss the meaning of similarity, proximity and functionality while considering the granules of physically existing or potentially derivable attributes. We also show several examples of utilization of the granulation structures defined over the feature spaces in the feature selection algorithms. As a case study, we consider the algorithms developed within the theory of rough sets, aimed at finding irreducible subsets of attributes that are sufficient to distinguish between the cases belonging to different target decision classes. Marek Grzegorowski, Andrzej Janusz, Dominik Slezak, Marcin S. Szczuka |
IEEE BigData | 3 |
| 2017 | Scalable cyber-security analytics with a new summary-based approximate query engineabstractA growing need for scalable solutions for both machine learning and interactive analytics exists in the area of cyber-security. Machine learning aims at segmentation and classification of log events, which leads towards optimization of the threat monitoring processes. The tools for interactive analytics are required to resolve the uncertain cases, whereby machine learning algorithms are not able to provide a convincing outcome and human expertise is necessary. In this paper we focus on a case study of a security operations platform, whereby typical layers of information processing are integrated with a new database engine dedicated to approximate analytics. The engine makes it possible for the security experts to query massive log event data sets in a standard relational style. The query outputs are received orders of magnitude faster than any of the existing database solutions running with comparable resources and, in addition, they are sufficiently accurate to make the right decisions about suspicious corner cases. The engine internals are driven by the principles of information granulation and summary-based processing. They also refer to the ideas of data quantization, approximate computing, rough sets and probability propagation. In the paper we study how the engine's parameters can influence its performance within the considered environment. In addition to the results of experiments conducted on large data sets, we also discuss some of our high level design decisions including the choice of an approximate query result accuracy measure that should reflect the specifics of the considered threat monitoring operations. Dominik Slezak, Agnieszka Chadzynska-Krasowska, Joel Holland, Piotr Synak, Rick Glick, Marcin Perkowski |
IEEE BigData | 1 |
| 2015 | Granular modeling with fuzzy comparatorsabstractWe present an overview of an approach to solving various real-life tasks related to Computational Intelligence by means of modeling with information granules. The particular methods of building the model are based on networks of fuzzy comparators. We demonstrate that comparator networks are a powerful and versatile tool suitable for applications. The introduction of methodology is accompanied with brief presentation of its formal basis. We also list existing and prospective, practical applications of the described approach. Lukasz Sosnowski, Marcin S. Szczuka, Dominik Slezak |
IEEE BigData | 3 |
| 2014 | Implementing algorithms of rough set theory and fuzzy rough set theory in the R package "RoughSets"
Lala Septem Riza, Andrzej Janusz, Christoph Bergmeir, Chris Cornelis, Francisco Herrera, Dominik Slezak, José Manuel Benítez 0001 |
Inf. Sci. | 6 |
| 2014 | Processing and mining complex data streams
Jerzy Stefanowski, Alfredo Cuzzocrea, Dominik Slezak |
Inf. Sci. | 3 |
| 2012 | Management of Information Incompleteness in Rough Non-deterministic Information Analysis
Hiroshi Sakai, Michinori Nakata, Dominik Slezak |
IPMU (1) | 3 |
| 2012 | Rough SQL - Semantics and Execution
Dominik Slezak, Piotr Synak, Graham Toppin, Jakub Wroblewski, Janusz Borkowski |
IPMU (2) | 1 |
| 2010 | Injecting domain knowledge into a granular database engine: a position paperabstractWe discuss how to use techniques from such fields as text processing and knowledge management to better handle text attributes in the Infobright's RDBMS engine. Our approach leads to a rich interface for domain experts who wish to share their knowledge about data content and, on the other hand, it remains unnoticeable to data users. It enables to improve data storage, data access, and data compression, with no changes required at the database schema level. Dominik Slezak, Graham Toppin |
CIKM | 1 |
| 2010 | Attribute selection with fuzzy decision reducts
Chris Cornelis, Richard Jensen, Germán Hurtado Martín, Dominik Slezak |
Inf. Sci. | 4 |
| 2010 | Introduction to the special issue on advanced information retrieval and databases
Aijun An, Dominik Slezak |
J. Intell. Inf. Syst. | 2 |
| 2009 | Data warehouse technology by infobrightabstractWe discuss Infobright technology with respect to its main features and architectural differentiators. We introduce the upcoming research and development projects that may be of special interest to the academic and industry communities. Dominik Slezak, Victoria Eastwood |
SIGMOD Conference | 1 |
| 2009 | Degrees of conditional (in)dependence: A framework for approximate Bayesian networks and examples related to the rough set-based feature selection
Dominik Slezak |
Inf. Sci. | 1 |
| 2008 | Brighthouse: an analytic data warehouse for ad-hoc queriesabstractBrighthouse is a column-oriented data warehouse with an automatically tuned, ultra small overhead metadata layer called Knowledge Grid, that is used as an alternative to classical indexes. The advantages of column-oriented data storage, as well as data compression have already been well-documented, especially in the context of analytic, decision support querying. This paper demonstrates additional benefits resulting from Knowledge Grid for compressed, column-oriented databases. In particular, we explain how it assists in query optimization and execution, by minimizing the need of data reads and data decompression. Dominik Slezak, Jakub Wroblewski, Victoria Eastwood, Piotr Synak |
Proc. VLDB Endow. | 1 |
| 2003 | Application of Temporal Descriptors to Musical Instrument Sound Recognition
Alicja Wieczorkowska, Jakub Wroblewski, Piotr Synak, Dominik Slezak |
J. Intell. Inf. Syst. | 4 |
| 1999 | Classification Algorithms Based on Linear Combinations of Features
Dominik Slezak, Jakub Wroblewski |
PKDD | 1 |
| 1997 | Neural Networks Design: Rough Set Approach to Continuous Data
Hung Son Nguyen, Marcin S. Szczuka, Dominik Slezak |
PKDD | 3 |