Katarzyna Kaczmarek-Majer

dblp:138/3661 · also Katarzyna Kaczmarek · DBLP profile ↗
← Back
20ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0003-0422-9366ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Fuzzy linguistic summaries and the double negation property
abstract
In this contribution, we formalize the general case of linguistic summaries with multiple summarizers and qualifiers. We extend the definitions present in the literature to handle cases in the following form: Q R1 ⋆. .. ⋆ R k y's are P1 ⋄. .. ⋄ P l where ⋆, ⋄ ∈ {AND, OR}. Moreover, we study the consistency of fuzzy summaries, focusing mainly on the double negation property (DN). We consider different negations for linguistic quantifiers and show which can be used for preserving the DN property.
Katarzyna Mis, Katarzyna Kaczmarek-Majer, Michal Baczynski 0001
Fuzzy Sets Syst.2
2025 Incremental learning and granular computing from evolving data streams: An application to speech-based bipolar disorder diagnosis
abstract
We apply an evolving granular-computing modeling approach, called evolving Optimal Granular System (eOGS), to bipolar mood disorder (BD) diagnosis based on speech data streams. The eOGS online learning algorithm reveals information granules in the flow and design the structure and parameters of a granular rule-based model with a certain degree of interpretability based on acoustic attributes obtained from phone calls made over 7 months to the Psychiatry department of a hospital. A multi-objective programming problem that trades-off information specificity, model compactness, and numerical and granular error indices is presented. Spectral and prosodic attributes are ranked and selected based on a hybrid Pearson-Spearman correlation coefficient . Low attribute-class correlation, ranging from 0.03 to 0.07, is observed, as well as high class overlap, which is typical in the psychiatric field. eOGS models for BD recognition overcome alternative computational-intelligence models, namely, Dynamic Evolving Neural-Fuzzy Inference System (DENFIS) and Fuzzy-set-Based evolving Modeling (FBeM-Gauss), by a small margin in both best and average cases; followed by eXtended Takagi-Sugeno (xTS) and evolving Takagi Sugeno (eTS) types of models. The proposed eOGS model using only 8 of the original acoustic attributes, and about 15 ‘If-Then’ inference rules, has exhibited the best root mean square error , 0.1361, and 91.8% accuracy in sharp BD class estimates. Granules associated to linguistic labels and a granular input-output map offer human understandability with relation to the inherent process of generating class estimates. Linguistically readable eOGS rules may assist physicians in explaining symptoms and making a diagnosis.
Daniel F. Leite, Gabriella Casalino, Katarzyna Kaczmarek-Majer, Giovanna Castellano
Fuzzy Sets Syst.3
2024 Fuzzy Linguistic Summaries for Hidden Markov Models
Katarzyna Kaczmarek-Majer, Michal Baczynski 0001, Olgierd Hryniewicz, Katarzyna Mis, Weronika Mucha, Filip Wichrowski
IPMU (3)1
2024 Explainable Impact of Partial Supervision in Semi-Supervised Fuzzy Clustering
abstract
Controlling the impact of partial supervision on the outcomes of modeling is of uttermost importance in semisupervised fuzzy clustering. Semi-Supervised Fuzzy C-Means (SSFCMeans), a specific model we consider, uses a single hyperparameter called a scaling factor α to weigh the impact of partially labeled data. This concept became widespread and was reused directly in many works building on SSFCMeans, or even applied to other fuzzy clustering algorithms such as Possibilistic C-Means. However, none of the works challenged the original interpretation of α which suggests that the impact of partial supervision is directly proportional to the scaling factor. We fill the above research gap and thoroughly analyze this relationship. We provide novel explanations of the scaling factor α in terms of the key element of fuzzy clustering - the membership values. We prove that the impact of partial supervision is a non-linear function of α. Our approach is rooted in the explainability framework, which distinguishes interpretation from an explanation and treats the latter as superior. Explaining the scaling factor leads to an explainable impact of partial supervision and enables greater control of it. Finally, built on the novel explanations, we propose a unified, analytically justified framework for selecting the value of the hyperparameter α that is based on the crossvalidation approach. We illustrate that the proposed framework enables an extensive analysis of the impact of partial supervision in SSFCMeans with a simulation experiment.
Kamil Kmita, Katarzyna Kaczmarek-Majer, Olgierd Hryniewicz
IEEE Trans. Fuzzy Syst.2
2024 Combined Defuzzification Under Shared Constraint
abstract
Defuzzification of fuzzy sets is an important aspect of fuzzy processing, as it determines how the fuzziness is dealt with when performing the final step to obtain crisp solutions. In this contribution, we consider the possibility of defuzzifying multiple general type-1 fuzzy sets, given that their defuzzified values are bound by a single, known constraint. This stems from an application in spatial data processing, where we obtained a number of fuzzy sets whose defuzzified values should sum up to a crisp and known value. It is possible to defuzzify each fuzzy set individually and rescale the outcomes to meet the constraint. However, considering that the fuzzy set contains information on what constitutes good values, accounting for the constraint in the defuzzification process has the potential to yield better results. We considered this as an optimization problem and investigated appropriate goal functions. The approach and goal functions are discussed with respect to typical properties of defuzzifiers.
Jörg Verstraete, Weronika Radziszewska, Katarzyna Kaczmarek-Majer, Sebastian Bykuc
IEEE Trans. Fuzzy Syst.3
2022 Impact of clustering unlabeled data on classification: case study in bipolar disorder
abstract
Currently, it is possible to collect a large amount of data from sensors.At the same time, data are often only partially labeled.For example, in the context of smartphonebased monitoring of mental state, there are much more data collected from smartphones than those collected from psychiatrists about the mental state.The approach presented in this paper is designed to examine if unlabeled data can improve the accuracy of classification tasks in the considered case study of classifying a patient's state.First, unlabeled data are represented by clusters membership through Fuzzy C-means algorithm which corresponds to the uncertainty of the patient's condition in this disease.Secondly, the classification is performed using two well-known algorithms, Random Forest and SVM.The obtained results indicate a minimal improvement in the quality of classification thanks to the use of membership in clusters.These results are promising due to both, the accuracy and interpretability.
Olga Kaminska, Katarzyna Kaczmarek-Majer, Olgierd Hryniewicz
FedCSIS2
2022 Experimental evaluation of the accuracy of an ensemble of fuzzy methods for classification of episodes in bipolar disorder
abstract
Clinical practice confirms that speech can support the diagnosis of several mental disorders. For example, reduced speech activity, changes in specific voice features, and pause-related measures were found to be sensitive markers of depressive symptoms. Considering the possibility of continuous speech data collection via a smartphone app, voice analysis has great potential for monitoring mental states. Nevertheless, there is still a need to select the most effective validation approaches for solving the task of predicting the mental state. Those validation approaches shall consider that the data collected from sensors and the response variables considered in this BD application problem are subject to various sources of uncertainty. The aim of the study is to perform an experimental evaluation of the accuracy of top-performing crisp and fuzzy methods, such as Naive Bayes Network, SOTA algorithm, Fuzzy Rule, Probabilistic Neural Network, Decision Tree, Gradient Boosted Tree, Random Forest, Tree Ensemble, and an ensemble approach that combines them. Various training and testing scenarios are considered for each of these methods, consisting of a given percentage of all observations. Additionally, the results from multiple methods are aggregated using the dominant function. Thus, the most frequent rating is taken and a metric based on fuzzy numbers is also considered for comparative purposes. The preliminary results of numerical experiments are promising. The sensitive point is the vicinity of the threshold of transition to a disease state. It should be noted that due to minor differences inherent in such cases, it seems intuitive to use fuzzy numbers to determine the patient’s assessment. Experiments confirmed also that the ranking of methods depends on the choice of the training set and evaluation metric.
Katarzyna Kaczmarek-Majer, Adam Kiersztyn
FUZZ-IEEE1
2022 Confidence path regularization for handling label uncertainty in semi-supervised learning: use case in bipolar disorder monitoring
abstract
Semi-supervised learning has gained great interest because of its ability to combine unlabeled data with – potentially few – labeled observations in a training process. However, in some application contexts, one can question whether all available labels are equally valid. For example, in the context of bipolar disorder (BD) remote monitoring, a common practice is to extrapolate the psychiatrist’s assessment onto some fixed time window surrounding the visit, the so-called ground truth period. In consequence, all data from this period are labeled with the same category. Such an approach may potentially result in misguided supervision affecting the model’s performance. In this paper, we consider the problem of label uncertainty, assuming that the labels are crisp, but they may be assigned to particular observations with varying confidence. We propose a novel method called Confidence Path Regularization (CPR) that incorporates this uncertainty into the fuzzy c-means semi-supervised learning. The proposed CPR approach is a novel method for automatic, data-driven handling of label uncertainty. We achieve it by estimating the confidence factor for each labeled observation. In addition, CPR allows for the exploration of potential class-specific patterns in the adjusted confidence. The proposed method is illustrated with experiments on partially labeled data about speech characteristics collected from smartphone application for BD monitoring. In this particular applied scenario, we also use additional contextual data to improve the construction of confidence paths. It is shown that the proposed CPR approach enables to reflect the varying confidence in labels as compared with the nominal approach which assigns the majority of observations to the same class associated with relevant ground truth period
Kamil Kmita, Gabriella Casalino, Giovanna Castellano, Olgierd Hryniewicz, Katarzyna Kaczmarek-Majer
FUZZ-IEEE5
2022 Expert-in-the-loop Stepwise Regression and its Application in Air Pollution Modeling
abstract
In this work, we provide a statistical procedure to integrate expert preferences towards explanatory variables in stepwise forward regression. The proposed method builds on the traditional stepwise linear regression and goal programming. The procedure is validated experimentally for real-life data from various sources aiming at predicting air pollution. The practical goal is to predict the annual concentrations of two health-related air pollutants, namely PM10 (Particulate Matter that is 10 micrometers or less in diameter) and NO2 (Nitrogen Dioxide). The main finding from this work is that inclusion of expert knowledge leads to more robust and accurate predictive models. Considering the limited size of data from air pollution monitoring stations, additional expert knowledge enabled to select most meaningful explanatory variables, and as the consequence the statistical inference lead to the improved predictions. The main contribution of this work is the proposed simple but solid expert-in-the-loop stepwise forward linear regression method allowing to include expert preferences. Experiments confirm that the proposed procedure is not only more interpretable but also delivers more accurate predictions for the considered air pollutants concentrations.
Milosz Fraszczyk, Katarzyna Kaczmarek-Majer, Olgierd Hryniewicz, Krzysztof Skotak, Anna Degórska
IS2
2022 Fuzzy Linguistic Summaries for Explaining Online Semi-Supervised Learning
abstract
Intelligent systems for the medical domain often require processing data streams that evolve over time and are only partially labeled. At the same time, the need for explanations is of utmost importance not only due to various regulations, but also to increase trust among systems’ users. In this work, an online data-driven learning method with focus on the explainability of evolving models equipped with incremental semi-supervised learning algorithms is considered. The proposed method combines: (i) the Dynamic Incremental Semi-Supervised Fuzzy C-Means (DISSFCM) algorithm to incrementally classify subsets of data; with (ii) Linguistic Summarization, which provides explanations of the classification results in terms of short sentences in a natural language. The approach has been illustrated for streaming data collected from voice calls of patients affected by Bipolar Disorder. The results show the effectiveness of the proposed method in classifying instances belonging to healthy and affective states, and explaining the approximate reasoning behind the classification of new acoustic data related to patients.
Katarzyna Kaczmarek-Majer, Gabriella Casalino, Giovanna Castellano, Daniel F. Leite, Olgierd Hryniewicz
IS1
2022 Explaining smartphone-based acoustic data in bipolar disorder: Semi-supervised fuzzy clustering and relative linguistic summaries
abstract
Smartphones enable to collect large data streams about phone calls that, once combined with Computational Intelligence techniques, bring great potential for improving the monitoring of patients with mental illnesses. However, the acoustic data streams recorded in uncontrolled environments are dynamically changing due to various sources of uncertainty. In addition, such acoustic data are usually difficult to interpret by psychiatrists. Within this study, we propose an approach based on Linguistic Summaries with Fuzzy Clustering (LS-FC) aiming at the development of human-consistent and easily interpretable summaries about relations between acoustic data and mental state of a patient affected by Bipolar Disorder, e.g., Most calls in the state of hypomania have low loudness compared to the state of euthymia [T = 1]. To capture the dynamics of acoustic data streams, we apply a dynamic incremental semi-supervised fuzzy clustering that synthesizes data into clusters. These clusters are represented by prototypes which are used for the construction of the membership functions describing linguistic terms e.g., low loudness, and then, linguistic summaries. The main contribution of this paper is the incorporation of information about clusters’ prototypes in the generation of linguistic summaries. The primary goal of this research is explainability. The semi-supervised learning algorithm is used mainly for deriving clusters and building improved linguistic summaries. Numerical results indicate that linguistic summaries provide intuitive and clear information about voice features in a patient’s affective state and they are consistent with clinical observation. In particular, during most calls in hypomania/mania both the quality of the patient’s voice and the dynamics of change in the spectrum signal reflected in spectral flux are low compared to euthymia. The proposed approach enables to summarize large data streams into meaningful descriptions that, although relatively simple, offer information granules that are very intuitive for clinicians and are promising to support the smartphone-based monitoring of bipolar disorder patients to inform about the potential change of mental state.
Katarzyna Kaczmarek-Majer, Gabriella Casalino, Giovanna Castellano, Olgierd Hryniewicz, Monika Dominiak
Inf. Sci.1
2022 PLENARY: Explaining black-box models in natural language through fuzzy linguistic summaries
abstract
We introduce an approach called PLENARY (exPlaining bLack-box modEls in Natural lAnguage thRough fuzzY linguistic summaries), which is an explainable classifier based on a data-driven predictive model. Neural learning is exploited to derive a predictive model based on two levels of labels associated with the data. Then, model explanations are derived through the popular SHapley Additive exPlanations (SHAP) tool and conveyed in a linguistic form via fuzzy linguistic summaries. The linguistic summarization allows translating the explanations of the model outputs provided by SHAP into statements expressed in natural language. PLENARY accounts for the imprecision related to model outputs by summarizing them into simple linguistic statements and for the imprecision related to the data labeling process by including additional domain knowledge in the form of middle-layer labels. PLENARY is validated on preprocessed speech signals collected from smartphones from patients with bipolar disorder and on publicly available mental health survey data. The experiments confirm that fuzzy linguistic summarization is an effective technique to support meta-analyses of the outputs of AI models. Also, PLENARY improves explainability by aggregating low-level attributes into high-level information granules, and by incorporating vague domain knowledge into a multi-task sequential and compositional multilayer perceptron. SHAP explanations translated into fuzzy linguistic summaries significantly improve understanding of the predictive modelling process and its outputs.
Katarzyna Kaczmarek-Majer, Gabriella Casalino, Giovanna Castellano, Monika Dominiak, Olgierd Hryniewicz, Olga Kaminska, Gennaro Vessio, Natalia Díaz Rodríguez
Inf. Sci.1
2021 Intelligent analysis of data streams about phone calls for bipolar disorder monitoring
abstract
Voice features from everyday phone conversations are regarded as a sensitive digital marker of mood phases in bipolar disorder. At the same time, although acoustic data collected from smartphones are relatively large, their psychiatric labelling is usually very limited, and there is still a need for intelligent and interpretable approaches to process such multiple data streams with a low percentage of labelling. Furthermore, both acoustic data and psychiatric labels are subject to several sources of uncertainty (e.g., irregular phone usage, background noises, subjectivity in psychiatric evaluation). To cope with these characteristics of an acoustic data stream, this paper introduces an intelligent qualitative and quantitative analysis based on the Dynamic Incremental Semi-Supervised Fuzzy C-Means algorithm (DISSFCM) for supporting bipolar disorder monitoring. The proposed approach is illustrated with real-life data collected from smartphones and psychiatric assessments of a bipolar disorder patient. Analysis of the dynamics of data streams basing on the cluster prototypes from fuzzy semi-supervised learning is a highly novel approach. It is also showed that the DISSFCM algorithm obtains relatively high classification performance (accuracy ranging from 0.66 to 0.76) already with 25% labelling percentage, thanks to the splitting mechanism that is adapting the number of clusters to the structure of data.
Gabriella Casalino, Giovanna Castellano, Katarzyna Kaczmarek-Majer, Olgierd Hryniewicz
FUZZ-IEEE3
2021 Possibilistic aggregation of inhomogeneous streams of data
abstract
Streams of data collected from sensors are usually large and inhomogeneous in time. In this paper, we consider the case when data consist of subsegments of different lengths governed by possibly different probability distributions. The data describing consecutive subsegments are presented in the form of histograms. Next, these subsegments are grouped in larger segments whose characteristics, such as measures of location or variability, are used in further analysis. We present a possibilistic method for the aggregation of subsegment data represented by histograms into segment data represented by possibilistic distributions. The performance of the proposed method is illustrated in monitoring of bipolar disorder patients using their voice data collected from smartphones.
Olgierd Hryniewicz, Katarzyna Kaczmarek-Majer
FUZZ-IEEE2
2020 Dynamic Incremental Semi-supervised Fuzzy Clustering for Bipolar Disorder Episode Prediction
Gabriella Casalino, Giovanna Castellano, Francesco Galetta, Katarzyna Kaczmarek-Majer
DS4
2020 Acoustic Feature Selection with Fuzzy Clustering, Self Organizing Maps and Psychiatric Assessments
Olga Kaminska, Katarzyna Kaczmarek-Majer, Olgierd Hryniewicz
IPMU (1)2
2019 Control charts based on fuzzy costs for monitoring short autocorrelated time series
Olgierd Hryniewicz, Katarzyna Kaczmarek-Majer, Karol R. Opara
Int. J. Approx. Reason.2
2019 Application of linguistic summarization methods in time series forecasting
Katarzyna Kaczmarek-Majer, Olgierd Hryniewicz
Inf. Sci.1
2018 Model Averaging Approach to Forecasting the General Level of Mortality
Marcin Bartkowiak, Katarzyna Kaczmarek-Majer, Aleksandra Rutkowska, Olgierd Hryniewicz
IPMU (1)2
2013 Linguistic knowledge about temporal data in Bayesian linear regression model to support forecasting of time series
Katarzyna Kaczmarek-Majer, Olgierd Hryniewicz
FedCSIS1