Radoslaw Cybulski

dblp:330/3235 · DBLP profile ↗
← Back
3ranked-venue papers in the field
3as first author
3since 2021 · last 2024
—ORCID · none

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 3 (3 first)
YearPublicationVenuePosition
2024 About Granular Rough Computing: Concept-Dependent Granulation Powered by Map Reduce
abstract
This work continues a cycle of research grounded in the foundational theories of granular computing introduced by Zadeh, rough set theory proposed by Pawlak, and particularly the concept-dependent granulation methodology developed within Polkowski’s theoretical framework. Building on our prior studies, we advance the application of granulation methods for knowledge extraction from decision-making systems. This approach forms knowledge granules—clusters of similar objects based on selected measures—allowing us to structure the universe of objects into prototypical representations that capture recurring data patterns. Such representations have demonstrated high efficacy in classification tasks, preserving data accuracy while achieving substantial reductions—up to 98%—in the size of training systems. In this paper, we further test the scalability of this method in large-scale data processing by employing MapReduce within the Apache Hadoop ecosystem, enhancing computational efficiency through distributed execution in a Java-based environment.
Radoslaw Cybulski
IEEE Big Data1
2023 Data Streaming in Concept-Dependent Granulation
abstract
In the evolving framework of granular computing, the continuous flow of data, often referred to as data streaming, presents both challenges and opportunities for the granulation process. Building upon the foundational works of Professor Zadeh and the granulation techniques based in rough set theory, this paper investigates the integration of data streaming with the concept-dependent granulation method. Recognizing the inherent challenges posed by the vastness and dynamism of streaming data, we explore the potential of adapting the concept-dependent granulation to accommodate and process these streams efficiently. Drawing inspiration from our previous works on random sampling and data decomposition, we introduce a novel approach that utilizes the real-time nature of data streams to enhance the granulation process. Our experimental studies, aim to evaluate the effectiveness of this method in terms of granulation quality and computational efficiency. Preliminary results suggest that integrating data streaming with concept-dependent granulation not only preserves the integrity of the granulated information but also offers significant advantages in processing large-scale dynamic data. In this work, we have verified the possibility of detecting the amount of data necessary to achieve the relevant classification efficiency, without having to process the entire data. This paper serves as a bridge, connecting the established methodologies of granulation with the emerging challenges of big data streaming, and sets the foundation for future research in this domain.
Radoslaw Cybulski, Piotr Artiemjew
IEEE Big Data1
2022 Accelerating concept-dependent granulation technique using data decomposition
abstract
The granular computing approach as a paradigm in approximate reasoning deals with the processing of knowledge into granules that consist of entities similar in information content. Within the framework of rough set theory, proposed 40 years ago by Zdzislaw Pawlak and developed since then by many authors, granulation is an important area of research. Granulation techniques have found application in many areas of data mining in classification, feature selection, clustering and approximation processes for decision-making systems, among others. The current work is located in the development of approximation techniques for decision systems using rough inclusions proposed by Polkowski. Polkowski proposed the hypothesis that granules induced in a data set of a universe of objects should lead to new object representing them, and such granulated counterparts should preserve the information content of the data. This hypothesis has been verified in a number of research papers by Polkowski and Artiemjew. It has been proven that one of the best granulation techniques from the proposed family of methods is concept-dependent granulation. In which, on selected data, the degree of approximation of decision-making systems reaches more than 90 percent reduction in the size of training systems while maintaining the classification efficiency of the original training data. The undoubted disadvantage of this technique is the quadratic computational complexity - which forces us to look for a way to apply it on big data sets.The current work is one in a series of papers on the application of techniques used to deal with big data sets in the context of accelerating the concept-dependent granulation method. In this paper, we test whether granulation of training data divided into subgroups with post-granulation object fusion is competitive - in terms of running time and classification efficiency - to using a full training system in the granulation and classification process. The posed experimental problem is verified on a selected decision-making systems from UCI repository and using simple kNN classifier.
Radoslaw Cybulski, Piotr Artiemjew
IEEE Big Data1