EDBT 2026 Demo / reviewers in the wild / expert
Urszula Stanczyk
dblp:92/3310
· DBLP profile ↗
27ranked-venue papers
20as first author
13since 2021 · last 2025
0000-0002-5071-7187ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 19 first-author · 13 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Constructing classifier committee exploiting transformations of attribute domainsabstractReaching a decision through deliberations of some component decision-making units is the notion applied in many domains. Collective judgement, combining information coming from many sources, with diverse properties or based on various data, can lead to enriched knowledge and result in more accurate predictions and enhanced understanding of the domain. The paper describes the research in which the constructed classifier committees used one type of learner and relied on the input data transformed by different discretisation processes. Based on partially distributed data and continuous and discrete representations of available attributes, several defined voting scenarios were investigated. They were employed for two well-known inducers applied to the task of authorship attribution in the domain of stylometric analysis of texts. The results from the performed experiments enable to observe conditions where decision aggregation based on characterising property of supervised discretisation algorithms can be advantageous, providing insights on roles played by the attributes in the decision-making process and the impact of data transformation. Urszula Stanczyk, Grzegorz Baron |
KES | 1 |
| 2024 | Domain-specific data characteristics: A study on meaning of stylometric sub-concepts and in-class imbalanceabstractIn the context of data imbalance probably the most investigated problem is imbalance of classes, as learning from the data with this characteristic makes detection of existing patterns for all classes more difficult. However, other problems related to imbalance also exists and the paper addresses such cases where classes are balanced, but there is in-class imbalance. Such imbalance can be caused by uneven representation of sub-concepts. When there is a noticeable difference between the numbers of samples belonging to sub-concepts, this can turn the under-represented sub-concepts into disjuncts. Data irregularities of this type can hinder recognition, therefore actions are typically taken to restore balance. In the investigations described, the issue was studied in the stylometric domain and various classifiers were applied to the data that was balanced, then imbalanced, and finally with restored balance. The experiments show that the specifics of the domain of application can put its own mark on the data which is difficult to overcome by standard processing such as under- or oversampling. Observed dependence on a learner and dataset makes the issue even more complex and layered, and shows the need for deeper studies. Urszula Stanczyk |
KES | 1 |
| 2024 | Observations of data characteristics and irregularities through domain-oriented transformations of attributesabstractThe paper presents research dedicated to observations of relations between attribute properties and discretisation. In the investigations described, the gradually increasing sets of features were discretised by selected approaches, and several variants of data were constructed. The continuous, partially discrete, and completely translated datasets were explored by the chosen classifiers and their performance studied in the context of a number of discretised attributes, discretisation procedures, and the way of processing of features and datasets. The stylometric problem of authorship attribution was the machine learning task under study. The experimental results enable to observe closer the specificity of style-markers employed as characteristic features, and indicate conditions for efficient recognition of authorship. They can be extended to other application domains with similar characteristics. Urszula Stanczyk, Grzegorz Baron |
KES | 1 |
| 2024 | Weighting Attributes Based on the Greedy Algorithm PropertiesabstractEstimation of importance for considered features is an important issue for any knowledge exploration process and it can be executed by a variety of approaches. In the research reported in this study, the primary aim was the development of a methodology for creating attribute rankings. Based on the properties of the greedy algorithm for inducing decision rules, a new application of this algorithm has been proposed. Instead of constructing a single ordering of features, attributes were weighted multiple times. The input datasets were discretised with several algorithms representing supervised and unsupervised discretisation approaches. Each resulting discrete data variant was exploited to construct a ranking of attributes. The effectiveness of the obtained rankings was confirmed through a rule filtering process governed by weighted attributes. The methodology was applied to the stylometric task of authorship attribution. The experimental outcomes demonstrate the value of the proposed research method, as it generally led to improved predictions while taking into account a noticeably decreased sets of attributes and decision rules. Beata Zielosko, Urszula Stanczyk, Kamil Jablonski |
KES | 2 |
| 2023 | Filtering Decision Rules Driven by Sequential Forward and Backward Selection of Attributes: An Illustrative Example in Stylometric DomainabstractThe paper presents investigations concerning the decision rule filtering process controlled by the estimated relevance of available attributes.In the conducted study, two search directions were used, sequential forward selection and sequential backward elimination.The steps of sequential search were governed by three rankings obtained for variables, all related to characteristics of data and rules that can be induced, as follows, (i) a ranking based on the weighting factor referring to the occurrence of attributes in generated decision reducts, (ii) the OneR ranking exploiting short rule properties, and (iii) the proposed ranking defined through the operation of greedy algorithm for rule induction.The three rankings were confronted and compared from the perspective of their usefulness for the selection of rules performed in the two directions and with two strategies for rule selection.The resulting sets of rules were analysed with respect to the properties of the constituent decision rules and from the point of performance for all constructed rulebased classifiers.Substantial experiments were carried out in the stylometric domain, treating the task of authorship attribution as classification.The results obtained indicate that for all three rankings and search paths it was possible to obtain a noticeable reduction of attributes while at least maintaining the power of inducers, at the same time improving characteristics of rule sets. Beata Zielosko, Urszula Stanczyk, Kamil Jablonski |
FedCSIS | 2 |
| 2023 | How transformations of representation for input data can affect the properties of induced decision reducts and rulesabstractDecision reducts and rules belong to forms used for the representation of knowledge learnt from input data while using a rough set approach in the exploration stage. As with any patterns that capture properties of data, the size is considered as the most important indicator of their quality, as a more general and concise description allows for improved understanding and simplified interpretation. The size of a reduct is given by its cardinality, and for a rule it is a number of included conditions, both referred to as lengths. These lengths undergo transformations along with data transformations, so such processing as discretisation leads to changed characteristics of existing patterns and through them also reducts and rules. The paper presents research with contributions in the observations focused on the influence of supervised and unsupervised discretisation on properties of attributes, induced decision reducts, and rules, offering a deeper insight into relations among them. Urszula Stanczyk |
KES | 1 |
| 2023 | Should supervised discretisation always be trusted unreservedly? On combining characteristics of supervised and unsupervised discretisation algorithms in two-step processingabstractThe paper presents a description of the research methodology dedicated to a two-step discretisation process applied to the input numeric data, with combining the characteristics of selected supervised and unsupervised algorithms, which leads to extended processing of some attributes in train and test sets. The methodology was illustrated with the investigations carried out in the domain of stylometric analysis of texts, for two datasets prepared for the task of binary authorship attribution. The several variants of transformed input data obtained were subjected to exploration using two selected machine learning methods capable of inducing knowledge from both continuous and categorical forms, namely the PART and J48 classifiers. The results from the experiments indicate that, as can be expected, supervised transformations of data work well enough, however, they do not always return the best outcome. The two-step processing of some attributes shows sufficient promise to warrant a closer study, as opposed to always unconditionally relying only on supervised algorithms as outperforming all other approaches. Urszula Stanczyk, Grzegorz Baron |
KES | 1 |
| 2022 | Evaluation of importance for condition attributes based on quality of decision reductsabstractRelative or decision reducts belong with mechanisms dedicated to feature selection, and they are embedded in rough set approach to data processing. Algorithms for reduct construction typically aim at dimensionality reduction aspect, searching for smallest reducts, which are considered as the most advantageous from the point of view of knowledge representation. However, classifiers build on reduced data models, based on reducts, can significantly vary in performance. Therefore, to ensure quality of predictions, other characteristics of reducts, apart from their cardinalities, need to be taken into account. The paper presents research in which estimation of reduct quality through their characteristics was reflected in calculation of the proposed weighting factors leading to attribute rankings. These rankings were next employed in the process of filtering decision rules, inferred by classic rough set approach. Constructed rule-based classifiers were applied in the stylometric domain to solve a task of authorship attribution. Urszula Stanczyk |
KES | 1 |
| 2022 | On heterogeneity or sub-classes aspect in construction of stylometric input datasetsabstractStylometric analysis of texts relies on learning characteristic traits of writing styles for authors. Once these patterns are discovered, they can be compared to the ones present in other text samples, to recognise their authorship. This recognition can be compromised if input datasets are prepared without taking into consideration possible stratification of the input space, leading to specific grouping of datapoints, or sub-classes within distinguished classes. The paper shows research dedicated to construction of various structures of input datasets, and combinations of such structures between train and test sets. In the research the influence of different stratification forms on the performance of selected popular classification systems was observed. To minimise the number of influencing factors, a task of authorship attribution was performed as binary classification with balanced classes. Stylometric descriptors exploited belonged to lexical and syntactic group, giving frequencies of occurrence for chosen style-markers. It resulted in real-valued attributes and these values were explored without applying discretisation, in order to avoid the possible bias of this procedure on observations. Urszula Stanczyk, Grzegorz Baron |
KES | 1 |
| 2022 | Ranking of attributes - comparative study based on data from stylometric domainabstractThe area of feature selection methods constantly expands along with the development of artificial intelligence domain, and has great impact on almost every field, whenever data is processed and explored. The paper presents research where a ranking method was proposed, inspired by an approach which comes from an algorithm for induction of decision rules. The ranking procedure was based on calculation of standard deviation for attributes, taking into account assigned class labels. This method was compared with another ranking mechanism, a modified version of popular Relief algorithm, with incorporating characteristics of variables by supervised discretisation. Comparison of obtained results included the aspect of knowledge representation as well as the perspective of the accuracy for constructed rule-based classifiers. The experiments were performed on datasets from stylometry domain, where authorship attribution was considered as a classification task, and stylometric descriptors as characteristic features defining writing styles of authors. Beata Zielosko, Urszula Stanczyk, Krzysztof Zabinski |
KES | 2 |
| 2021 | Standard vs. non-standard cross-validation: evaluation of performance in a space with structured distribution of datapointsabstractCross-validation is a popularly used approach to evaluation of performance for classifiers. It relies on random selection of independent samples for training and testing, and assumes that if any similarities among samples exist, they do not lead to known grouping of datapoints in the input space. If these conditions are violated, as it may happen for datasets with some structure of samples included, standard cross-validation can return biased results even for many folds. In the paper the research on cross-validation was reported for application to stylometric datasets, describing a task of authorship attribution. The comparison of standard and non-standard processing was presented. In the latter case, selected subsets of examples were swapped over between training and test sets several times. The experiments with three popular classifiers showed that standard cross-validation tended to give over-optimistic results, whereas non-standard processing was more guarded, and by that more reliable. To avoid high computational costs involved, evaluation based on averaged predictions for limited numbers of test sets can be considered as a reasonable compromise. Grzegorz Baron, Urszula Stanczyk |
KES | 2 |
| 2021 | Weighting factor for attributes based on reduct cardinalityabstractEstimation of attribute importance can be obtained by a mechanism that allows to assign some weight to variables. Weighting attributes can lead to their ordering, which, in turn, can be exploited for feature selection and reduction. Decision reducts constitute an example of a mechanism aiming at dimensionality reduction, embedded in rough set approach to data mining. The paper presents research works, where the process of weighting was driven by the proposed factor based on reducts, with varying their sets, and the results were analysed through the perspective of reduct cardinality, since it is typically considered as the most significant indicator of reduct quality. Constructed rankings of variables were used for inferring sets of decision rules from gradually decreasing numbers of features, and then the performance of the rule classifiers was tested. The experiments show that for the weighting factor to be useful for feature reduction, not only reduct cardinalities, but also the numbers of reducts found need to be taken into account. Urszula Stanczyk |
KES | 1 |
| 2021 | Condition attributes, properties of decision rules, and discretisation: Analysis of relations and dependenciesabstractWhen mining of input data is focused on rule induction, knowledge, discovered in exploration of existing patterns, is stored in combinations of certain conditions on attributes included in rule premises, leading to specific decisions. Through their properties, such as lengths, supports, cardinalities of rule sets, inferred rules characterise relations detected among variables. The paper presents research dedicated to analysis of these dependencies, considered in the context of various discretisation methods applied to the input data from stylometric domain. For induction of decision rules from data, Classical Rough Set Approach was employed. Next, based on rule properties, several factors were proposed and evaluated, reflecting characteristics of available condition attributes. They allowed to observe how variables and rule sets changed depending on applied discretisation algorithms. Beata Zielosko, Urszula Stanczyk |
KES | 2 |
| 2020 | Performance evaluation for ranking-based discretisationabstractDiscretisation often constitutes a part of initial data preparation stage. It translates continuous domain of features into granular, by assigning a number of intervals to represent attributes’ values by nominal categories. Typically all real-valued features are subjected to transformations, regardless of their characteristics. The paper presents research on discretisation executed with a discerning approach. To all available attributes, feature selection mechanisms were employed, in the form of rankings that order variables based on their importance. Exploiting this discovered knowledge on attributes, discretisation was then driven by a ranking, and either highest or lowest ranking features were selected for transformation. The influence of selective discretisation on the performance of classification systems was studied for three popular inducers. The procedure was employed in the field of stylometry, and a task of authorship recognition, considered as a binary classification with balanced classes. The experiments show that discretisation based on importance of features can lead to better performance than in the case of transformations applied to all attributes. Grzegorz Baron, Urszula Stanczyk |
KES | 2 |
| 2020 | Assessing quality of decision reductsabstractThe paper presents research focused on decision reducts, a feature reduction mechanism inherent to rough sets theory. As a reduct enables to protect the discriminative properties of attributes with respect to described concepts, from the point of data representation, a reduct length is considered to be the most important measure of its quality. However, such approach is insufficient while taking into account the performance of a reduct-based rule classifier applied to test samples. When many reducts of the same length are available, they can lead to vastly different predictions. The paper provides a description for the proposed procedure for iterative reduct generation, which results in decrease of diversity in the observed levels of accuracy, supporting reduct selection. The procedure was applied for binary classification with balanced classes, for the stylometric task of authorship attribution. Urszula Stanczyk, Beata Zielosko |
KES | 1 |
| 2020 | Reduct-based ranking of attributesabstractThe paper is dedicated to the area of feature selection, in particular a notion of attribute rankings that allow to estimate importance of variables. In the research presented for ranking construction a new weighting factor was defined, based on relative reducts. A reduct constitutes an embedded mechanism of feature selection, specific to rough set theory. The proposed factor takes into account the number of reducts in which a given attribute exists, as well as lengths of reducts. Two approaches for reduct generation were employed and compared, with search executed by a genetic algorithm. To validate the usefulness of the reduct-based rankings in the process of feature reduction, for gradually decreasing subsets of attributes, selected through rankings, sets of decision rules were induced in classical rough set approach. The performance of all rule classifiers was evaluated, and experimental results showed that the proposed rankings led to at least the same, or even increased classification accuracy for reduced sets of features than in the case of operating on the entire set of condition attributes. The experiments were performed on datasets from stylometry domain, with treating authorship attribution as a classification task, and stylometric descriptors as characteristic features defining writing styles. Beata Zielosko, Urszula Stanczyk |
KES | 2 |
| 2020 | Heuristic-based feature selection for rough set approachabstractThe paper presents the proposed research methodology, dedicated to the application of greedy heuristics as a way of gathering information about available features. Discovered knowledge, represented in the form of generated decision rules, was employed to support feature selection and reduction process for induction of decision rules with classical rough set approach. Observations were executed over input data sets discretised by several methods. Experimental results show that elimination of less relevant attributes through the proposed methodology led to inferring rule sets with reduced cardinalities, while maintaining rule quality necessary for satisfactory classification. Urszula Stanczyk, Beata Zielosko |
Int. J. Approx. Reason. | 1 |
| 2019 | On Approaches to Discretisation of Stylometric Data and Conflict Resolution in Decision MakingabstractThe paper presents research on unsupervised and supervised discretisation of input data used in execution of stylometric tasks of authorship attribution. Basing on numeric characterisation of writing styles, recognition of authorship is performed by decision rules, as their transparent structure enhances understanding of discovered knowledge. The performance of rule classifiers, constructed in rough set approach, is studied in the context of a strategy employed for resolving conflicts. It is also contrasted with that of other selected inducers. Urszula Stanczyk, Beata Zielosko |
KES | 1 |
| 2017 | Filtering Decision Rules with Continuous Attributes Governed by Discretisation
Urszula Stanczyk |
ISMIS | 1 |
| 2017 | Evaluating Importance for Numbers of Bins in Discretised Learning and Test Sets
Urszula Stanczyk |
KES-IDT (1) | 1 |
| 2016 | Measuring Quality of Decision Rules Through Ranking of Conditional Attributes
Urszula Stanczyk |
KES-IDT (1) | 1 |
| 2015 | Ranking of characteristic features in combined wrapper approaches to selectionabstractThe performance of a classification system of any type can suffer from irrelevant or redundant data, contained in characteristic features that describe objects of the universe. To estimate relevance of attributes and select their subset for a constructed classifier typically either a filter, wrapper, or an embedded approach, is implemented. The paper presents a combined wrapper framework, where in a pre-processing step, a ranking of variables is established by a simple wrapper model employing sequential backward search procedure. Next, another predictor exploits this resulting ordering of features in their reduction. The proposed methodology is illustrated firstly for a binary classification task of authorship attribution from stylometric domain, and then for additional verification for a waveform dataset from UCI machine learning repository. Urszula Stanczyk |
Neural Comput. Appl. | 1 |
| 2014 | RELIEF-based Selection of Decision RulesabstractWhen constructing rule classifiers for pattern recognition and classification tasks we can induce only some small set of decision rules, such as a minimal cover that is sufficient for recognition of learning samples, a subset of rules satisfying requirements for example with respect to rule support or strength, or a complete set of rules. Once some set is inferred, another approach becomes available, that of filtering out a group of rules meeting some given criteria. The paper presents the latter methodology, where all decision rules on examples are generated within Dominance-Based Rough Set Approach and the process of filtering exploits a ranking of conditional attributes obtained through Relief algorithm. The procedures are applied in the domain of computational stylistics or stylometry, dedicated to analysis of linguistic styles observable in samples of writing. Urszula Stanczyk |
KES | 1 |
| 2013 | Relative Reduct-Based Estimation of Relevance for Stylometric Features
Urszula Stanczyk |
ADBIS | 1 |
| 2013 | Establishing Relevance of Characteristic Features for Authorship Attribution with ANN
Urszula Stanczyk |
DEXA (2) | 1 |
| 2013 | On Preference Order of DRSA Conditional Attributes for Computational Stylistics
Urszula Stanczyk |
DEXA (2) | 1 |
| 2011 | Application of DRSA-ANN Classifier in Computational Stylistics
Urszula Stanczyk |
ISMIS | 1 |