Andrzej Janusz

dblp:02/2713 · DBLP profile ↗
← Back
21ranked-venue papers in the field
8as first author
11since 2021 · last 2026
0000-0002-9763-1399ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 10 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 4 (2 first)Other / Interdisciplinary · 4 (3 first)Business Process & Enterprise Data · 2 (1 first)Data Mining & Knowledge Discovery · 1
YearPublicationVenuePosition
2026 DupliMend: Online Detection and Refinement of Imprecise Activity Labels
Savandi Kalukapuge, Andrzej Janusz, Moe Thandar Wynn
CAiSE (2)2
2025 ThreatTrace: Cyber-Attack Detection Through Trace Abstraction and Soft Clustering
Andrzej Janusz, Savandi Kalukapuge, Moe Thandar Wynn
CAiSE (2)1
2024 Do Data Scientists Dream About Their Skills' Assessment? - Transforming a Competition Platform Into an Assessment Platform
abstract
We present a platform for automatic assessment of technical data science skills (hard skills) and competencies that help to apply those technical skills in practice (soft skills). The platform serves so-called assessment platform tasks that resemble data-focused tasks typical for online data science competitions. Actually, these tasks are designed based on international conference competitions that we have been organizing on our competition platform knowledgepit.ai. The main idea relies on our observation that the given person’s behavior during a competition (dynamics of submitting solutions, activity on competition forum, etc.) can be correlated with his/her soft competencies. Accordingly, our goal is to translate the specifics of international conference competitions into a framework that investigates the behaviors of data science job candidates, employees who want to improve their work in data science projects, students who wish to build their future professional careers on AI solutions, etc. We claim that such behaviors – if properly measured based on solving our assessment platform tasks – can effectively indicate both hard and soft types of competencies.
Dominik Slezak, Andrzej Janusz, Maciej Swiechowski, Agnieszka Chadzynska-Krasowska, Jacek Kaminski
IEEE Big Data2
2024 IEEE Big Data Cup 2024 Report: Predicting Chess Puzzle Difficulty at KnowledgePit.ai
abstract
We summarize the results of the IEEE BigData 2024 Cup: Predicting Chess Puzzle Difficulty – a data science competition organized at the knowledgepit.ai platform in association with the IEEE BigData 2024 conference. We describe the competition goal and tie it to existing research on human-computer interaction, focusing on task difficulty estimation and aligning human and AI behavior. We explain how we acquired and processed the data, separately for training and testing datasets. We review submitted solutions and evaluate their performance by comparing them to a simple benchmark. We explain how the achieved rating differences translate to user experience when solving chess puzzles. We further explore the concept of chess puzzle difficulty by replicating competition results with puzzle ratings obtained using chess bots only. We conclude with a summary of our findings and directions for future studies, as well as an invitation to the next edition of the competition in 2025.
Jan Zysko, Maciej Swiechowski, Sebastian Stawicki, Katarzyna Jagiela, Andrzej Janusz, Dominik Slezak
IEEE Big Data5
2023 IEEE BigData Cup 2023 Report: Object Recognition with Muon Tomography Using Cosmic Rays
abstract
We summarize the results of the IEEE BigData 2023 Cup: Object Recognition with Muon Tomography using Cosmic Rays - a data mining competition organized at the KnowledgePit.ai platform in association with the IEEE BigData 2023 conference. We describe the challenge at the heart of the competition task, as well as the data acquisition and preparation steps. We present the entire process of preparing experiments and subsequent data analysis for the purpose of recognizing X-rayed objects using muon tomography techniques. We conclude this analysis by presenting the baseline as well as the winning solution of the object segmentation algorithms for the research space reconstruction and object classification.
Mateusz Wnuk, Jan Dziuba, Andrzej Janusz, Dominik Slezak
IEEE Big Data3
2023 A practical study of methods for deriving insightful attribute importance rankings using decision bireducts
abstract
Subject matter experts (SMEs) often rely on attribute importance rankings to verify machine learning models, acquire insights into their outcomes, and gain a deeper understanding of the investigated phenomena. To further increase their usefulness, we introduce a new approach to the evaluation of attribute rankings produced by any machine learning method . As a real-world case study , we investigate the attribute importance scores produced using XGBoost and decision bireducts on the data gathered by an HR company, where the goal is to predict the willingness of candidates to change their job. For this task, XGBoost delivers accurate models but fails to identify many attributes that are important to SMEs. In comparison, decision bireducts lead to models that are easier to interpret and explore the data with a higher focus on the diversity of attributes. The ensembles of decision bireducts deliver comparable accuracy and their associated attribute rankings are more insightful than those of XGBoost.
Andrzej Janusz, Dominik Slezak, Sebastian Stawicki, Krzysztof Stencel
Inf. Sci.1
2022 EVEAL - Expected Variance Estimation for Active Learning
abstract
Regression problems frequently occur in the surrounding world, therefore are unavoidable in real-world applications. However, to obtain a model with desired generalization performance, a vast amount of labels usually is required. In many scenarios, obtaining unlabelled data is relatively inexpensive, therefore active learning approaches may be used to reduce the needed annotation effort. Most uncertainty-based regression active learning algorithms use variance estimation of model predictions to choose informative samples. Those algorithms do not incorporate knowledge about the data distribution for the given task. In this paper, we propose a novel algorithm to incorporate information about data distribution and combine it with variance estimation as an informativeness function. Experiments conducted on four data sets show that the proposed approach outperforms standard variance-based sampling by a margin, and indicate its robustness.
Daniel Kaluza, Andrzej Janusz, Dominik Slezak
IEEE Big Data2
2022 IEEE BigData Cup 2022 Report Privacy-preserving Matching of Encrypted Images
abstract
We summarize the results of IEEE BigData 2022 Cup: Privacy-preserving Matching of Encrypted Images - a data mining challenge organized at the KnowledgePit platform in association with the IEEE BigData 2022 conference. We describe the challenge in the hearth of the competition task, as well as the data acquisition and preparation steps. We also provide a brief overview of the top-performing solutions submitted by participants. Finally, we present results of the post-competition data analysis, in which we consider the similarity of solutions submitted by various teams in terms of their errors on the test data. We conclude this analysis by discussion of the significance and impact of the competition results on the underlying problem of constructing effective and efficient anonymization algorithms for monitoring in the DOOH advertising industry.
Marcin S. Szczuka, Andrzej Janusz, Boguslaw Cyganek, Jakub Grabek, Lukasz Przebinda, Andzelika Zalewska-Küpçü, Andrzej Bukala, Dominik Slezak
IEEE Big Data2
2022 Prescriptive Analytics for Optimization of FMCG Delivery Plans
Marek Grzegorowski, Andrzej Janusz, Stanislaw Lazewski, Maciej Swiechowski, Monika Jankowska
IPMU (2)2
2022 Learning multimodal entity representations and their ensembles, with applications in a data-driven advisory framework for video game players
Andrzej Janusz, Daniel Kaluza, Maciej Matraszek, Lukasz Grad, Maciej Swiechowski, Dominik Slezak
Inf. Sci.1
2021 Predicting Victories in Video Games - IEEE BigData 2021 Cup Report
abstract
We summarize the results of IEEE BigData 2021 Cup: Predicting Victories in Video Games - a data mining challenge organized at the KnowledgePit platform in association with the IEEE BigData 2021 conference. We describe the competition task, as well as the data acquisition and preprocessing steps. We also provide a brief overview of the top-performing solutions submitted by participants. Finally, we present results of the post-competition data analysis, in which we consider the similarity of solutions submitted by various teams in terms of their errors on the test data. We conclude this analysis by demonstrating a method for constructing an ensemble of submitted solutions. Such an ensemble performs better than any of the individual solutions submitted during the competition.
Maciej Matraszek, Andrzej Janusz, Maciej Swiechowski, Dominik Slezak
IEEE BigData2
2020 Predicting Escalations in Customer Support: Analysis of Data Mining Challenge Results
abstract
We summarize IEEE Big Data Cup: Predicting Escalations in Customer Support - a data mining competition organized jointly by companies Information Builders and QED Software at the KnowledgePit platform, in the frame of the 2020 IEEE International Conference on Big Data. We discuss the motivation for organizing this event and highlight the factors that make it such a challenging topic. We describe the data provided to participants and formulate the competition task. We also provide an overview of competition results with a detailed analysis of a few selected solutions. Finally, we present a novel functionality of the KnowledgePit platform - an analytic module that allows organizers to investigate selected solutions using a convenient GUI and provides in-depth insights about their quality.
Andrzej Janusz, Guohua Hao, Daniel Kaluza, Tony Li, Robert Wojciechowski, Dominik Slezak
IEEE BigData1
2019 IEEE BigData 2019 Cup: Suspicious Network Event Recognition
abstract
“IEEE BigData 2019 Cup: Suspicious Network Event Recognition” was a data mining competition organized jointly by companies Security On-Demand and QED Software at the KnowledgePit online platform, in association with the IEEE BigData 2019 conference. The scope of this challenge referred to the notions of cybersecurity analytics and network alert evaluation. In this paper, we summarize the results of our competition. We explain how data sets had been prepared before it was possible to make them available to competition participants. We describe the baseline scoring models that we designed as a reference for participants, and we demonstrate how critical for their performance was the aspect of appropriate feature engineering. We also discuss the results of experiments conducted to verify the (un)suitability of deep recurrent neural networks in this particular case. In some sense, we show that there are no “perfect” machine learning approaches that could be applied equally successfully to every data science undertaking.
Andrzej Janusz, Daniel Kaluza, Agnieszka Chadzynska-Krasowska, Bartek Konarski, Joel Holland, Dominik Slezak
IEEE BigData1
2018 Toward Machine Learning on Granulated Data - a Case of Compact Autoencoder-based Representations of Satellite Images
abstract
We consider a problem of learning from compact representations of images for a purpose of object recognition and content-based image retrieval. We discuss a motivation for using compressed images in those tasks and indicate exemplary applications related to analysis on the data from satellites. Finally, we show some preliminary results of experiments conducted to demonstrate the impact of the image data granulation on the quality of classification. We empirically compare the performance of prediction models trained on original images, images compressed using autoencoders, and on images whose quality was lowered in order to reduce their size.
Mateusz Przyborowski, Tomasz Tajmajer, Lukasz Grad, Andrzej Janusz, Piotr Biczyk, Dominik Slezak
IEEE BigData4
2018 How to Match Jobs and Candidates - A Recruitment Support System Based on Feature Engineering and Advanced Analytics
Andrzej Janusz, Sebastian Stawicki, Michal Drewniak, Krzysztof Ciebiera, Dominik Slezak, Krzysztof Stencel
IPMU (2)1
2018 SENSEI: An Intelligent Advisory System for the eSport Community and Casual Players
abstract
In this article, we describe the SENSEI system. It helps players to improve their skills in popular eSports games. We discuss the main goals of the system and explain the associated challenges. We also present its conceptual architecture which aims at enabling full automation of the data acquisition and analytic processes. The system is expected to provide in-depth analytics of players' performance and give practical advice regarding possible improvements. Thus its architecture allows players to provide feedback and manually label important concepts. Finally, we discuss our first case study - an advisory system for popular collectible card video games.
Andrzej Janusz, Dominik Slezak, Sebastian Stawicki, Krzysztof Stencel
WI1
2018 A framework for learning and embedding multi-sensor forecasting models into a decision support system: A case study of methane concentration in coal mines
Dominik Slezak, Marek Grzegorowski, Andrzej Janusz, Michal Kozielski, Sinh Hoa Nguyen, Marek Sikora, Sebastian Stawicki, Lukasz Wróbel
Inf. Sci.3
2017 On the role of feature space granulation in feature selection processes
abstract
Information granulation plays an important role in the process of scaling up modern machine learning and knowledge discovery algorithms. By employing compact descriptions of granules - whereby granules are defined as collections of original data elements gathered together by means of their similarity, proximity or functionality - one can drastically accelerate computations and, moreover, make the results of those computations more meaningful for domain experts. In this paper, we summarize some of the feature space granulation approaches introduced by now. We discuss the meaning of similarity, proximity and functionality while considering the granules of physically existing or potentially derivable attributes. We also show several examples of utilization of the granulation structures defined over the feature spaces in the feature selection algorithms. As a case study, we consider the algorithms developed within the theory of rough sets, aimed at finding irreducible subsets of attributes that are sufficient to distinguish between the cases belonging to different target decision classes.
Marek Grzegorowski, Andrzej Janusz, Dominik Slezak, Marcin S. Szczuka
IEEE BigData2
2014 Implementing algorithms of rough set theory and fuzzy rough set theory in the R package "RoughSets"
Lala Septem Riza, Andrzej Janusz, Christoph Bergmeir, Chris Cornelis, Francisco Herrera, Dominik Slezak, José Manuel Benítez 0001
Inf. Sci.2
2012 An Ensemble Approach to Multi-label Classification of Textual Data
Karol Kurach, Krzysztof Pawlowski, Lukasz Romaszko, Marcin Tatjewski, Andrzej Janusz, Hung Son Nguyen
ADMA5
2010 Discovering Rules-Based Similarity in Microarray Data
Andrzej Janusz
IPMU1