VLDB 2026 Research / reviewers in the wild / expert
Agnieszka Nowak-Brzezinska
dblp:62/3698 · also Agnieszka Nowak
· DBLP profile ↗
28ranked-venue papers
17as first author
11since 2021 · last 2025
0000-0001-7238-1170ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 14 first-author · 11 since 2021Theory of computation · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorSystems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Categorical Pseudo-Labeling with Iterative Cluster ExpansionabstractLabeling categorical data is a critical challenge in many real-world applications where expert annotation is costly, time-consuming, and often impractical. Moreover, we very often face the problem of training sets that are too small or unrepresentative to apply supervised learning techniques. Traditional pseudo-labeling techniques, primarily designed for numerical data, struggle to handle the unique characteristics of categorical datasets. To address this, we introduce Categorical Pseudo-Labeling with Iterative Cluster Expansion (CPLICE), a novel clustering-based approach that leverages qualitative similarity measures to iteratively expand labeled datasets. Unlike probability-based pseudo-labeling, CPLICE selects data points based on their distance from established clusters, progressively refining labels while handling ambiguous cases through dynamic reassignments. Our results demonstrate that CPLICE outperforms traditional tree-based models in low-label environments, providing a stable, scalable solution for effective data labeling. This method is particularly valuable for real-world scenarios such as medical diagnostics, customer segmentation, and fraud detection, where robust handling of categorical data is essential for accurate, interpretable models. Weronika Lazarz, Agnieszka Nowak-Brzezinska |
KES | 2 |
| 2025 | Evaluating Clustering Quality in Categorical Data: A Comparative Analysis and Novel MetricsabstractAbstract Weronika Lazarz, Agnieszka Nowak-Brzezinska |
KES | 2 |
| 2025 | A hybrid knowledge-based machine learning approach for ticket classification in the AWS cloudabstractAutomated classification of customer support tickets is vital for modern service desks to reduce workload and improve response times. Machine learning (ML) methods provide flexibility and accuracy but often lack explainability and incur high computational costs. Knowledge-based systems (KBS) offer transparency and determinism but struggle with ambiguous or unseen inputs. This paper presents a hybrid approach combining a lightweight rule-based classifier with a fallback ML model, deployed in a scalable serverless AWS environment. We evaluated logistic regression with TF-IDF, sentence embedding classifiers, and the hybrid model on a curated multilingual support dataset using accuracy, F1 score, precision, and recall. Logistic regression outperformed embedding-based classifiers, achieving the highest macro-averaged F1 scores. The hybrid system leverages domain-specific rules to handle high-confidence cases and defers to ML when rules do not apply, consistently outperforming standalone ML models. Additional experiments tested robustness across multiple runs, train-test splits, and cross-lingual generalization with German data. Deployed using AWS Lambda, S3, and DynamoDB, the solution balances low latency and scalability. The results demonstrate the hybrid approach’s practical advantages for reliable, efficient support ticket classification. Agnieszka Nowak-Brzezinska, Dawid Drabek |
KES | 1 |
| 2024 | Edge-based categorical clustering for data with related featuresabstractSupervised methods have been the leading methods for inferring from data, but unsupervised models, particularly with data clustering, are becoming more common in business processes. Cluster analysis methods are extremely effective when we do not have a clearly defined decision-making class, or the data contain anomalies that are difficult to identify. A particularly big problem arises with the efficient processing of categorical data. Such data type is, for example, amino acid sequences, where their co-occurrence is extremely important. Therefore, we propose a new approach to clustering categorical sequential data based on graph edges -GECC. This algorithm not only identifies clusters well but also, due to its implementation based on a graph structure, it is fast and does not burden memory resources. This article will show that our method perfectly reflects the natural data clusters. Weronika Lazarz, Agnieszka Nowak-Brzezinska |
KES | 2 |
| 2024 | Analysis of selected algorithms for detecting outliers in dataabstractThis study involves analysis of various machine learning models for anomaly detection, focusing on diverse aspects such as sampling types, model tuning, and the size of training sets. By exploring a broad spectrum of supervised, unsupervised, and semi-supervised algorithms, the research aims to furnish a nuanced understanding of model performance dynamics within the anomaly detection domain. A significant finding from this analysis is the superior performance of tuned models over their untuned counterparts, thereby underscoring the pivotal role of hyperparameter tuning in augmenting algorithmic efficiency. The examination of sampling methodologies reveals that non-augmented sampling techniques strike an optimal balance between accuracy and training duration. Although hybrid and oversampling methods necessitate extended training periods, they nonetheless yield competitive outcomes. Conversely, undersampling approaches facilitate expedited training processes although at the expense of reduced average AUCPR values. Agnieszka Nowak-Brzezinska, Dawid Jasiak |
KES | 1 |
| 2023 | Detecting outliers in rule-based knowledge bases using Self-Organizing Map and Local Outlier Factor algorithmsabstractOur research deals with intelligent decision support systems based on rule-based knowledge bases. Decision support systems use rules ”If a condition, then a decision” as a form of knowledge representation. In the process of inference, which mirrors the process of human reasoning, we look for rules that confirm the facts and thus generate new knowledge. Such rule-based knowledge bases can (and often do) contain outlier rules. Our goal is to find such unusual rules. Thanks to this, we can influence the completeness of the knowledge base by finding unusual rules and asking domain experts to supplement knowledge in a rare area. To enhance the effectiveness of decision support systems, we conducted separate investigations into two distinct methods. The first method involved the utilisation of the Local Outlier Factor (LOF) algorithm in detecting rule outliers, while the second method employed the Self-Organizing Maps (SOM) algorithm for the same purpose. Our experiments not only confirmed the effectiveness of both the LOF and SOM algorithms but also involved comparing the results obtained from both methods. The discovery of outlier rules can aid knowledge engineers and domain experts in knowledge exploration and enhance the completeness of the knowledge base, which is crucial for decision support systems. Czeslaw Horyn, Agnieszka Nowak-Brzezinska |
KES | 2 |
| 2023 | Comparative analysis of selected algorithms for qualitative data clusteringabstractData clustering can apply to both numeric and categorical attributes. The numerical characteristics of data are all such values that can be written as digits. This work is not directly devoted to this category of attributes. The research's primary and most crucial goal is to compare the clustering processes, the quality of the obtained clusters, and the clustering results between the selected algorithms for clustering numerical sets (using coding techniques) and qualitative data sets. Design experiments were performed on three datasets. The algorithms used to group the observations into classes were the k-means algorithm and the kmodes algorithm. Three different distance determination measures were used in the case of the k-means algorithm - the Euclidean measure, the urban measure, and the Chebyshev measure. It can be concluded that for the structures of the three selected data sets, the k-modes algorithm presents better results of the clustering quality and thus more effectively distributes observations to classes. It also does it in a much shorter time than the k-means algorithm. The keyword used here, however, is dataset structure. Agnieszka Nowak-Brzezinska, Jakub Królik |
KES | 1 |
| 2023 | The similarity of the similarity measures in the context of clustering algorithms for categorical dataabstractThis research aims to assess the similarity between determining the value of the similarity of objects by using various measures. Particularly for categorical data, the used similarity measures may behave differently depending on whether the analyzed data are binary, non-binary, or mixed. In the case of categorical data, it begins to matter how many attributes we have and how many different values they can take. Cluster analysis is an essential technique in unsupervised machine learning. Object-clustering cannot be determined without determining similarities between objects in the set and searching for clusters with the most significant internal consistency and external separatory. Therefore, we need knowledge about which measure behaves similarly and which differs depending on the nature of the analyzed data. According to our knowledge, there is no such research in the literature. In research, we analyzed six different measures for eight different datasets (differing in size and data type). The results say, among others, which measures take the longest time for calculations and which are the most correlated with others, so they can be used interchangeably without fear of losing valuable information about the mutual similarity of objects. Agnieszka Nowak-Brzezinska, Wojciech Ogieglo |
KES | 1 |
| 2022 | The Quality of Clustering Data Containing Outliers
Agnieszka Nowak-Brzezinska, Igor Gaibei |
ACIIDS (2) | 1 |
| 2022 | Self-Organizing Map algorithm as a tool for outlier detectionabstractThe research addresses the problem of outlier detection using the (unsupervised) self-organizing map (S OM) algorithm introduced by Kohonen [1]. Since the outliers can indicate something scientifically interesting, we need to analyze these issue deeply. In this research, we do not assume that after outliers are detected, they are removed. We hand them over to domain experts asking for further exploration. We have experience applying the LOF (local outlier factor) algorithm to detect outliers in real data. Now we wanted to check how a new algorithm, an S OM algorithm, will behave. We compare the results of outlier detection process using the LOF algorithm with the results of using the S OM algorithm. We were interested whether a data type significantly influences the efficiency of the process of outlier detection. It is obvious that qualitative data is much more difficult to explore than quantitative data. We chose two various dataset (quantitive and qualitative). We kept changing the learning parameters: a learning rate, a neighborhood radius, an activation function, a distance measure and others and analyzed the accuracy of the outlier detection process. The accuracy of the S OM algorithm for outlier detection is satisfying. This algorithm is faster than the LOF algoritm. The results of the experiments show that there is a strong negative correlation between the learning parameters and the training time of an SOM map and the QE quantization error computing. Agnieszka Nowak-Brzezinska, Czeslaw Horyn |
KES | 1 |
| 2021 | Outliers in Covid 19 data based on Rule representation - the analysis of LOF algorithmabstractThe article concerns the detection of outliers in rule-based knowledge bases containing data on Covid 19 cases. The authors move from the automatic generation of a rule-based knowledge base from source data by clustering rules in the knowledge base to optimize inference processes and to detecting unusual rules allowing for the optimal structure of rule groups. The paper presents a two-phase procedure, wherein in the first phase, we look for the optimal structure of rule clusters when there are outlier rules in the knowledge base. In the second phase, we detect outliers in the rules using the LOF (Local Outlier Factor) algorithm. Then we eliminate the unusual rules from the database and check whether the selected cluster quality measures are responded positively to the elimination of outliers, which would indicate that the rules were rightly considered outliers. The performed experiments confirmed the effectiveness of the LOF algorithm and selected cluster quality measures in the context of detecting atypical rules. The detection of such rules can support knowledge engineers or domain experts in knowledge mining to improve the completeness of the knowledge base, which is usually the basis of the decision support system. Agnieszka Nowak-Brzezinska, Czeslaw Horyn |
KES | 1 |
| 2020 | Outliers in rules - the comparision of LOF, COF and KMEANS algorithmsabstractThe aim of the article is the analysis of using LOF, COF and Kmeans algorithms for outlier detection in rule based knowledge bases. The subject of outlier mining is very important nowadays. Outliers in rules mean unusual rules which are rare in comparison to others and should be explored further by the domain expert. In the research the authors use the outlier detection methods to find a given (1%, 5%, 10%) number of outliers in rules. Then, they analyze which of seven various quality indices, that they used for all rules and after removing selected outliers, improve the quality of rule clusters. In the experimental stage the authors used six different knowledge bases. The results show that the optimal results were achieved for COF outlier detection algorithm as the one for which, among all analyzed quality indices, the cluster quality improved most frequently. Agnieszka Nowak-Brzezinska, Czeslaw Horyn |
KES | 1 |
| 2019 | Exploration of rule-based knowledge bases: A knowledge engineer's support
Agnieszka Nowak-Brzezinska, Alicja Wakulicz-Deja |
Inf. Sci. | 1 |
| 2018 | Methods of Rule Clusters' Representation in Domain Knowledge Bases
Agnieszka Nowak-Brzezinska |
ICCCI (2) | 1 |
| 2018 | Experimental Implementation of Web-Based Knowledge Base Verification Module
Roman Siminski, Agnieszka Nowak-Brzezinska, Michal Siminski |
ICCCI (2) | 2 |
| 2018 | Different Methods for Cluster's Representation and Their Impact on the Effectiveness of Searching Through Such a Structure
Tomasz Xieski, Agnieszka Nowak-Brzezinska |
ICCCI (2) | 2 |
| 2017 | Knowledge Exploration in Medical Rule-Based Knowledge Bases
Agnieszka Nowak-Brzezinska, Tomasz Rybotycki, Roman Siminski, Malgorzata Przybyla-Kasperek |
ICCCI (2) | 1 |
| 2017 | Decision Fusion Methods in a Dispersed Decision System - A Comparison on Medical Data
Malgorzata Przybyla-Kasperek, Agnieszka Nowak-Brzezinska, Roman Siminski |
ICCCI (2) | 2 |
| 2017 | Comparison of similarity measures in context of rules clusteringabstractThis paper introduces five similarity measures, very well known in literature, but not because of using them to compare rules between themselves and choose the most similar one. Rules in knowledge bases are a very specific type of data representation and it is necessary to compare them carefully. The goal of the paper is to analyze the influence of using different similarity measures on the number of clusters, or the size of the representatives of the created clusters of rules. The results of the experiments are presented in Section III in order to discuss the significance of the analyzed measures and methods of rules creating. Agnieszka Nowak-Brzezinska, Tomasz Rybotycki |
INISTA | 1 |
| 2017 | Outlier mining in rule-based knowledge basesabstractThis paper introduces an approach to outlier mining in the context of rule-based knowledge bases. Rules in knowledge bases are a very specific type of data representation and it is necessary to analyze them carefully, especially when they differ from each other. The goal of the paper is to analyze the influence of using different similarity measures and clustering methods on the number of outliers discovered during the mining process. The results of the experiments are presented in Section V in order to discuss the significance of the analyzed parameters. Agnieszka Nowak-Brzezinska |
INISTA | 1 |
| 2016 | Mining Medical Knowledge Bases
Agnieszka Nowak-Brzezinska, Tomasz Rybotycki, Roman Siminski, Malgorzata Przybyla-Kasperek |
ICCCI (2) | 1 |
| 2016 | Intersection Method, Union Method, Product Rule and Weighted Average Method in a Dispersed Decision-Making System - a Comparative Study on Medical Data
Malgorzata Przybyla-Kasperek, Agnieszka Nowak-Brzezinska |
ICCCI (2) | 2 |
| 2016 | KBExplorator and KBExpertLib as the Tools for Building Medical Decision Support Systems
Roman Siminski, Agnieszka Nowak-Brzezinska |
ICCCI (2) | 2 |
| 2016 | Mining Rule-based Knowledge Bases Inspired by Rough Set TheoryabstractRule-based knowledge bases are constantly increasing in volume, thus the knowledge stored as a set of rules is getting progressively more complex and when rules are not organized into any structure, the system is inefficient. The aim of this paper is to improve the performance of mining knowledge b ases by modification of both their structure and inference algorithms, which in author’s opinion, lead to improve the efficiency of the inference process. The good performance of this approach is shown through an extensive experimental study carried out on a collection of real knowledge bases. Experiments prove that rules partition enables reducing significantly the percentage of the knowledge base analysed during the inference process. It was also proved that the form of the group’s representative plays an important role in the efficiency of the inference process. Agnieszka Nowak-Brzezinska |
Fundam. Informaticae | 1 |
| 2014 | Exploratory Clustering and VisualizationabstractIn this work the topic of applying clustering as a knowledge extraction method from real-world data is discussed. Authors propose a two-phase cluster creation and visualization technique, which combines hierarchical and density-based algorithms1. What is more, authors analyze the impact of data sampling on the result of searching through such a structure. Particular attention was also given to the problem of cluster visualization. Authors review selected, two-dimensional approaches, stating their advantages and drawbacks in the context of representing complex cluster structures. Agnieszka Nowak-Brzezinska, Tomasz Xieski |
KES | 1 |
| 2013 | Complex Decision Systems and Conflicts Analysis ProblemabstractThis paper discusses the issues related to the conflict analysis method and the rough set theory, process of global decision-making on the basis of knowledge which is stored in several local knowledge bases. The value of the rough set theory and conflict analysis applied in practical decision support systems with complex domain knowledge are expressed. The furthermore examples of decision support systems with complex domain knowledge are presented in this article. The paper proposes a new approach to the organizational structure of a multi-agent decision-making system, which operates on the basis of dispersed knowledge. In the presented system, the local knowledge bases will be combined into groups in a dynamic way. We will seek to designate groups of local bases on which the test object is classified to the decision classes in a similar manner. Then, a process of knowledge inconsistencies elimination will be implemented for created groups. Global decisions will be made using one of the methods for analysis of conflicts. Alicja Wakulicz-Deja, Agnieszka Nowak-Brzezinska, Malgorzata Przybyla-Kasperek |
Fundam. Informaticae | 2 |
| 2012 | Visual modeling of condition monitoring systemsabstractIn this paper a graphical approach for Condition Monitoring Systems (CM) based on Model Driven Architecture is presented; in particular, the software application “Smart Monitoring Agent” or SMA. This graphical approach starts from the idea of modeling, applying and visualizing non-hierarchical relations. The presented method makes use of well-defined data models that take the solution to a superior level of configurability, using only graphical tools. Finally it is explained how this approach enables the flexible knowledge exchange between condition monitoring experts and software engineers. Maciej Zygmunt, Marek Budyn, Michal Orkisz, James R. Ottewill, Victor H. Jaramillo, Agnieszka Nowak-Brzezinska |
ETFA | 6 |
| 2006 | Optimization of Speech Recognition by Clustering of Phones
Agnieszka Nowak-Brzezinska, Alicja Wakulicz-Deja, Sebastian Bachlinski |
Fundam. Informaticae | 1 |