Agnieszka Nowak-Brzezinska

dblp:62/3698 · also Agnieszka Nowak · DBLP profile ↗
← Back
28ranked-venue papers
17as first author
11since 2021 · last 2025
0000-0001-7238-1170ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 14 first-author · 11 since 2021Theory of computation · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorSystems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Categorical Pseudo-Labeling with Iterative Cluster Expansion
abstract
Labeling categorical data is a critical challenge in many real-world applications where expert annotation is costly, time-consuming, and often impractical. Moreover, we very often face the problem of training sets that are too small or unrepresentative to apply supervised learning techniques. Traditional pseudo-labeling techniques, primarily designed for numerical data, struggle to handle the unique characteristics of categorical datasets. To address this, we introduce Categorical Pseudo-Labeling with Iterative Cluster Expansion (CPLICE), a novel clustering-based approach that leverages qualitative similarity measures to iteratively expand labeled datasets. Unlike probability-based pseudo-labeling, CPLICE selects data points based on their distance from established clusters, progressively refining labels while handling ambiguous cases through dynamic reassignments. Our results demonstrate that CPLICE outperforms traditional tree-based models in low-label environments, providing a stable, scalable solution for effective data labeling. This method is particularly valuable for real-world scenarios such as medical diagnostics, customer segmentation, and fraud detection, where robust handling of categorical data is essential for accurate, interpretable models.
Weronika Lazarz, Agnieszka Nowak-Brzezinska
KES2
2025 Evaluating Clustering Quality in Categorical Data: A Comparative Analysis and Novel Metrics
abstract
Abstract
Weronika Lazarz, Agnieszka Nowak-Brzezinska
KES2
2025 A hybrid knowledge-based machine learning approach for ticket classification in the AWS cloud
abstract
Automated classification of customer support tickets is vital for modern service desks to reduce workload and improve response times. Machine learning (ML) methods provide flexibility and accuracy but often lack explainability and incur high computational costs. Knowledge-based systems (KBS) offer transparency and determinism but struggle with ambiguous or unseen inputs. This paper presents a hybrid approach combining a lightweight rule-based classifier with a fallback ML model, deployed in a scalable serverless AWS environment. We evaluated logistic regression with TF-IDF, sentence embedding classifiers, and the hybrid model on a curated multilingual support dataset using accuracy, F1 score, precision, and recall. Logistic regression outperformed embedding-based classifiers, achieving the highest macro-averaged F1 scores. The hybrid system leverages domain-specific rules to handle high-confidence cases and defers to ML when rules do not apply, consistently outperforming standalone ML models. Additional experiments tested robustness across multiple runs, train-test splits, and cross-lingual generalization with German data. Deployed using AWS Lambda, S3, and DynamoDB, the solution balances low latency and scalability. The results demonstrate the hybrid approach’s practical advantages for reliable, efficient support ticket classification.
Agnieszka Nowak-Brzezinska, Dawid Drabek
KES1
2024 Edge-based categorical clustering for data with related features
abstract
Supervised methods have been the leading methods for inferring from data, but unsupervised models, particularly with data clustering, are becoming more common in business processes. Cluster analysis methods are extremely effective when we do not have a clearly defined decision-making class, or the data contain anomalies that are difficult to identify. A particularly big problem arises with the efficient processing of categorical data. Such data type is, for example, amino acid sequences, where their co-occurrence is extremely important. Therefore, we propose a new approach to clustering categorical sequential data based on graph edges -GECC. This algorithm not only identifies clusters well but also, due to its implementation based on a graph structure, it is fast and does not burden memory resources. This article will show that our method perfectly reflects the natural data clusters.
Weronika Lazarz, Agnieszka Nowak-Brzezinska
KES2
2024 Analysis of selected algorithms for detecting outliers in data
abstract
This study involves analysis of various machine learning models for anomaly detection, focusing on diverse aspects such as sampling types, model tuning, and the size of training sets. By exploring a broad spectrum of supervised, unsupervised, and semi-supervised algorithms, the research aims to furnish a nuanced understanding of model performance dynamics within the anomaly detection domain. A significant finding from this analysis is the superior performance of tuned models over their untuned counterparts, thereby underscoring the pivotal role of hyperparameter tuning in augmenting algorithmic efficiency. The examination of sampling methodologies reveals that non-augmented sampling techniques strike an optimal balance between accuracy and training duration. Although hybrid and oversampling methods necessitate extended training periods, they nonetheless yield competitive outcomes. Conversely, undersampling approaches facilitate expedited training processes although at the expense of reduced average AUCPR values.
Agnieszka Nowak-Brzezinska, Dawid Jasiak
KES1
2023 Detecting outliers in rule-based knowledge bases using Self-Organizing Map and Local Outlier Factor algorithms
abstract
Our research deals with intelligent decision support systems based on rule-based knowledge bases. Decision support systems use rules ”If a condition, then a decision” as a form of knowledge representation. In the process of inference, which mirrors the process of human reasoning, we look for rules that confirm the facts and thus generate new knowledge. Such rule-based knowledge bases can (and often do) contain outlier rules. Our goal is to find such unusual rules. Thanks to this, we can influence the completeness of the knowledge base by finding unusual rules and asking domain experts to supplement knowledge in a rare area. To enhance the effectiveness of decision support systems, we conducted separate investigations into two distinct methods. The first method involved the utilisation of the Local Outlier Factor (LOF) algorithm in detecting rule outliers, while the second method employed the Self-Organizing Maps (SOM) algorithm for the same purpose. Our experiments not only confirmed the effectiveness of both the LOF and SOM algorithms but also involved comparing the results obtained from both methods. The discovery of outlier rules can aid knowledge engineers and domain experts in knowledge exploration and enhance the completeness of the knowledge base, which is crucial for decision support systems.
Czeslaw Horyn, Agnieszka Nowak-Brzezinska
KES2
2023 Comparative analysis of selected algorithms for qualitative data clustering
abstract
Data clustering can apply to both numeric and categorical attributes. The numerical characteristics of data are all such values that can be written as digits. This work is not directly devoted to this category of attributes. The research's primary and most crucial goal is to compare the clustering processes, the quality of the obtained clusters, and the clustering results between the selected algorithms for clustering numerical sets (using coding techniques) and qualitative data sets. Design experiments were performed on three datasets. The algorithms used to group the observations into classes were the k-means algorithm and the kmodes algorithm. Three different distance determination measures were used in the case of the k-means algorithm - the Euclidean measure, the urban measure, and the Chebyshev measure. It can be concluded that for the structures of the three selected data sets, the k-modes algorithm presents better results of the clustering quality and thus more effectively distributes observations to classes. It also does it in a much shorter time than the k-means algorithm. The keyword used here, however, is dataset structure.
Agnieszka Nowak-Brzezinska, Jakub Królik
KES1
2023 The similarity of the similarity measures in the context of clustering algorithms for categorical data
abstract
This research aims to assess the similarity between determining the value of the similarity of objects by using various measures. Particularly for categorical data, the used similarity measures may behave differently depending on whether the analyzed data are binary, non-binary, or mixed. In the case of categorical data, it begins to matter how many attributes we have and how many different values they can take. Cluster analysis is an essential technique in unsupervised machine learning. Object-clustering cannot be determined without determining similarities between objects in the set and searching for clusters with the most significant internal consistency and external separatory. Therefore, we need knowledge about which measure behaves similarly and which differs depending on the nature of the analyzed data. According to our knowledge, there is no such research in the literature. In research, we analyzed six different measures for eight different datasets (differing in size and data type). The results say, among others, which measures take the longest time for calculations and which are the most correlated with others, so they can be used interchangeably without fear of losing valuable information about the mutual similarity of objects.
Agnieszka Nowak-Brzezinska, Wojciech Ogieglo
KES1
2022 The Quality of Clustering Data Containing Outliers
Agnieszka Nowak-Brzezinska, Igor Gaibei
ACIIDS (2)1
2022 Self-Organizing Map algorithm as a tool for outlier detection
abstract
The research addresses the problem of outlier detection using the (unsupervised) self-organizing map (S OM) algorithm introduced by Kohonen [1]. Since the outliers can indicate something scientifically interesting, we need to analyze these issue deeply. In this research, we do not assume that after outliers are detected, they are removed. We hand them over to domain experts asking for further exploration. We have experience applying the LOF (local outlier factor) algorithm to detect outliers in real data. Now we wanted to check how a new algorithm, an S OM algorithm, will behave. We compare the results of outlier detection process using the LOF algorithm with the results of using the S OM algorithm. We were interested whether a data type significantly influences the efficiency of the process of outlier detection. It is obvious that qualitative data is much more difficult to explore than quantitative data. We chose two various dataset (quantitive and qualitative). We kept changing the learning parameters: a learning rate, a neighborhood radius, an activation function, a distance measure and others and analyzed the accuracy of the outlier detection process. The accuracy of the S OM algorithm for outlier detection is satisfying. This algorithm is faster than the LOF algoritm. The results of the experiments show that there is a strong negative correlation between the learning parameters and the training time of an SOM map and the QE quantization error computing.
Agnieszka Nowak-Brzezinska, Czeslaw Horyn
KES1
2021 Outliers in Covid 19 data based on Rule representation - the analysis of LOF algorithm
abstract
The article concerns the detection of outliers in rule-based knowledge bases containing data on Covid 19 cases. The authors move from the automatic generation of a rule-based knowledge base from source data by clustering rules in the knowledge base to optimize inference processes and to detecting unusual rules allowing for the optimal structure of rule groups. The paper presents a two-phase procedure, wherein in the first phase, we look for the optimal structure of rule clusters when there are outlier rules in the knowledge base. In the second phase, we detect outliers in the rules using the LOF (Local Outlier Factor) algorithm. Then we eliminate the unusual rules from the database and check whether the selected cluster quality measures are responded positively to the elimination of outliers, which would indicate that the rules were rightly considered outliers. The performed experiments confirmed the effectiveness of the LOF algorithm and selected cluster quality measures in the context of detecting atypical rules. The detection of such rules can support knowledge engineers or domain experts in knowledge mining to improve the completeness of the knowledge base, which is usually the basis of the decision support system.
Agnieszka Nowak-Brzezinska, Czeslaw Horyn
KES1
2020 Outliers in rules - the comparision of LOF, COF and KMEANS algorithms
abstract
The aim of the article is the analysis of using LOF, COF and Kmeans algorithms for outlier detection in rule based knowledge bases. The subject of outlier mining is very important nowadays. Outliers in rules mean unusual rules which are rare in comparison to others and should be explored further by the domain expert. In the research the authors use the outlier detection methods to find a given (1%, 5%, 10%) number of outliers in rules. Then, they analyze which of seven various quality indices, that they used for all rules and after removing selected outliers, improve the quality of rule clusters. In the experimental stage the authors used six different knowledge bases. The results show that the optimal results were achieved for COF outlier detection algorithm as the one for which, among all analyzed quality indices, the cluster quality improved most frequently.
Agnieszka Nowak-Brzezinska, Czeslaw Horyn
KES1
2019 Exploration of rule-based knowledge bases: A knowledge engineer's support
Agnieszka Nowak-Brzezinska, Alicja Wakulicz-Deja
Inf. Sci.1
2018 Methods of Rule Clusters' Representation in Domain Knowledge Bases
Agnieszka Nowak-Brzezinska
ICCCI (2)1
2018 Experimental Implementation of Web-Based Knowledge Base Verification Module
Roman Siminski, Agnieszka Nowak-Brzezinska, Michal Siminski
ICCCI (2)2
2018 Different Methods for Cluster's Representation and Their Impact on the Effectiveness of Searching Through Such a Structure
Tomasz Xieski, Agnieszka Nowak-Brzezinska
ICCCI (2)2
2017 Knowledge Exploration in Medical Rule-Based Knowledge Bases
Agnieszka Nowak-Brzezinska, Tomasz Rybotycki, Roman Siminski, Malgorzata Przybyla-Kasperek
ICCCI (2)1
2017 Decision Fusion Methods in a Dispersed Decision System - A Comparison on Medical Data
Malgorzata Przybyla-Kasperek, Agnieszka Nowak-Brzezinska, Roman Siminski
ICCCI (2)2
2017 Comparison of similarity measures in context of rules clustering
abstract
This paper introduces five similarity measures, very well known in literature, but not because of using them to compare rules between themselves and choose the most similar one. Rules in knowledge bases are a very specific type of data representation and it is necessary to compare them carefully. The goal of the paper is to analyze the influence of using different similarity measures on the number of clusters, or the size of the representatives of the created clusters of rules. The results of the experiments are presented in Section III in order to discuss the significance of the analyzed measures and methods of rules creating.
Agnieszka Nowak-Brzezinska, Tomasz Rybotycki
INISTA1
2017 Outlier mining in rule-based knowledge bases
abstract
This paper introduces an approach to outlier mining in the context of rule-based knowledge bases. Rules in knowledge bases are a very specific type of data representation and it is necessary to analyze them carefully, especially when they differ from each other. The goal of the paper is to analyze the influence of using different similarity measures and clustering methods on the number of outliers discovered during the mining process. The results of the experiments are presented in Section V in order to discuss the significance of the analyzed parameters.
Agnieszka Nowak-Brzezinska
INISTA1
2016 Mining Medical Knowledge Bases
Agnieszka Nowak-Brzezinska, Tomasz Rybotycki, Roman Siminski, Malgorzata Przybyla-Kasperek
ICCCI (2)1
2016 Intersection Method, Union Method, Product Rule and Weighted Average Method in a Dispersed Decision-Making System - a Comparative Study on Medical Data
Malgorzata Przybyla-Kasperek, Agnieszka Nowak-Brzezinska
ICCCI (2)2
2016 KBExplorator and KBExpertLib as the Tools for Building Medical Decision Support Systems
Roman Siminski, Agnieszka Nowak-Brzezinska
ICCCI (2)2
2016 Mining Rule-based Knowledge Bases Inspired by Rough Set Theory
abstract
Rule-based knowledge bases are constantly increasing in volume, thus the knowledge stored as a set of rules is getting progressively more complex and when rules are not organized into any structure, the system is inefficient. The aim of this paper is to improve the performance of mining knowledge b ases by modification of both their structure and inference algorithms, which in author’s opinion, lead to improve the efficiency of the inference process. The good performance of this approach is shown through an extensive experimental study carried out on a collection of real knowledge bases. Experiments prove that rules partition enables reducing significantly the percentage of the knowledge base analysed during the inference process. It was also proved that the form of the group’s representative plays an important role in the efficiency of the inference process.
Agnieszka Nowak-Brzezinska
Fundam. Informaticae1
2014 Exploratory Clustering and Visualization
abstract
In this work the topic of applying clustering as a knowledge extraction method from real-world data is discussed. Authors propose a two-phase cluster creation and visualization technique, which combines hierarchical and density-based algorithms1. What is more, authors analyze the impact of data sampling on the result of searching through such a structure. Particular attention was also given to the problem of cluster visualization. Authors review selected, two-dimensional approaches, stating their advantages and drawbacks in the context of representing complex cluster structures.
Agnieszka Nowak-Brzezinska, Tomasz Xieski
KES1
2013 Complex Decision Systems and Conflicts Analysis Problem
abstract
This paper discusses the issues related to the conflict analysis method and the rough set theory, process of global decision-making on the basis of knowledge which is stored in several local knowledge bases. The value of the rough set theory and conflict analysis applied in practical decision support systems with complex domain knowledge are expressed. The furthermore examples of decision support systems with complex domain knowledge are presented in this article. The paper proposes a new approach to the organizational structure of a multi-agent decision-making system, which operates on the basis of dispersed knowledge. In the presented system, the local knowledge bases will be combined into groups in a dynamic way. We will seek to designate groups of local bases on which the test object is classified to the decision classes in a similar manner. Then, a process of knowledge inconsistencies elimination will be implemented for created groups. Global decisions will be made using one of the methods for analysis of conflicts.
Alicja Wakulicz-Deja, Agnieszka Nowak-Brzezinska, Malgorzata Przybyla-Kasperek
Fundam. Informaticae2
2012 Visual modeling of condition monitoring systems
abstract
In this paper a graphical approach for Condition Monitoring Systems (CM) based on Model Driven Architecture is presented; in particular, the software application “Smart Monitoring Agent” or SMA. This graphical approach starts from the idea of modeling, applying and visualizing non-hierarchical relations. The presented method makes use of well-defined data models that take the solution to a superior level of configurability, using only graphical tools. Finally it is explained how this approach enables the flexible knowledge exchange between condition monitoring experts and software engineers.
Maciej Zygmunt, Marek Budyn, Michal Orkisz, James R. Ottewill, Victor H. Jaramillo, Agnieszka Nowak-Brzezinska
ETFA6
2006 Optimization of Speech Recognition by Clustering of Phones
Agnieszka Nowak-Brzezinska, Alicja Wakulicz-Deja, Sebastian Bachlinski
Fundam. Informaticae1