Agnieszka Wosiak

dblp:152/5664 · DBLP profile ↗
← Back
16ranked-venue papers
11as first author
9since 2021 · last 2024
0000-0001-6124-1236ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 11 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 first-authorSoftware engineering, systems software and programming languages · 5 · 5 first-author
YearPublicationVenuePosition
2024 New Presence-Dependent Binary Similarity Measures for Pairwise Label Comparisons in Multi-label Classification
Agnieszka Wosiak, Rafal Wozniak
ICCCI (2)1
2024 Efficacy of feature selection and Classification algorithms in cancer remission using medical imaging
abstract
This study evaluates the efficacy of feature selection and Classification algorithms in predicting cancer remission using medical imaging. The goal of our research was to build a solution to estimate potential cancer remission based on just one single scan. We applied feature selection technique using a filtering approach and F-value to extract the most informative attributes from PET and CT images. We tested four Classification algorithms: Decision Tree, K-Nearest Neighbors, Random Forest, and Support Vector Classifier. Our findings indicate that the Random Forest, especially when applied to PET images, was the most effective for cancer remission, achieving an accuracy of 90%. This study demonstrates the potential of tailored feature selection and machine learning models to enhance the prediction of cancer remission outcomes in clinical settings.
Malgorzata Krzywicka, Agnieszka Wosiak
KES2
2024 New binary similarity measures for enhanced disease correlation analysis and comorbidity detection
abstract
This study evaluates the effectiveness of new binary similarity measures, SAMP and SPAN, for identifying comorbidity patterns in patient data. These measures enhance the recognition and Classification of comorbidities in binary medical data by excluding negative matches and including absence mismatches. Unlike traditional methods such as Jaccard, Dice, and Chi-square, which are often limited in detecting complex disease correlations, SAMP and SPAN demonstrate improved efficiency and clustering quality in handling comorbidities. The experiments revealed that these measures significantly enhance the detection of disease associations and patient clustering. The findings suggest that SAMP and SPAN can offer valuable insights into the interactions between different health conditions, potentially leading to more effective diagnostic and treatment strategies.
Agnieszka Wosiak, Klaudia Gabryelczak, Katarzyna Zykwinska
KES1
2023 Sensitivity of Standard Evaluation Metrics for Disease Classification and Progression Assessment Based on Whole-Body Imaging
abstract
The paper investigates the limitations of existing metrics in evaluating the classification of medical images, particularly whole-body imaging used in nuclear medicine for assessing disease progression. We demonstrated that while current metrics such as Recall (TPR), NPV, Miss rate (FNR), FOR, F1 Score, and AUC effectively assess the performance of classifiers, they fall short in capturing subtle changes in medical images that are crucial for clinical decision-making. We performed an in-depth evaluation of standard metrics and discuss potential improvements for developing a more comprehensive methodology to assess classification and disease progression based on medical images.
Malgorzata Krzywicka, Agnieszka Wosiak
KES2
2023 Improving Automatic Recognition of Emotional States Using EEG Data Augmentation Techniques
abstract
Emotion recognition is crucial for improving communication between humans and computers. Electroencephalography (EEG) signals can be used for this purpose, but the limited amount of available EEG data poses challenges in creating accurate classification models, particularly when using deep machine learning methods. To address this issue, we investigated the impact of data augmentation on the quality of emotion prediction for inter-subject classification based on public emotion EEG datasets, the MAHNOB. The effectiveness of different data augmentation techniques, including sliding windows of varying lengths, overlapping windows, and Gaussian noise, were tested. We verified the augmentation techniques by automated emotion classification using a shallow convolutional neural network and the Valence/Arousal model. The results demonstrate that data augmentation can significantly improve the accuracy of emotion prediction for individual subjects, with the noise method being the most effective. Augmentation using Gaussian noise achieves up to a 30% improvement for a single subject in both Arousal and Valence emotion model dimensions, compared to the baseline. Our findings highlight the potential of data augmentation as a promising approach for improving the accuracy of emotion recognition using EEG signals.
Patrycja Szczakowska, Agnieszka Wosiak, Katarzyna Zykwinska
KES2
2022 Unsupervised emotional state recognition based on clustering of EEG features
abstract
Efficient information retrieval from the EEG sensors is a complex and challenging task, particularly in the context of psychology, including emotional states. Therefore, different machine learning strategies are considered to improve the processes based on EEG signal analysis. Most of them use supervised approaches since EEG datasets usually include metadata and descriptions that can be used for learning. However, these descriptions are mainly based on self-reports of emotional states, which means that they may not be reliable or objective. The paper proposes an approach that incorporates unsupervised learning techniques as a solution supporting classification where classification labels may be uncertain. The research proved that our approach improves the recognition of emotions and gives results with an average accuracy greater by fve percentage points.
Aleksandra Dura, Agnieszka Wosiak
KES2
2021 EEG channel selection strategy for deep learning in emotion recognition
abstract
Emotions play an important role in everyday life and contribute to physical and mental health. Emotional states can be detected by electroencephalography (EEG signals). Efficient information retrieval from the EEG sensors is a complex and challenging task. Therefore, deep learning methods for EEG signal analysis attract more and more attention. Many researchers emphasize automated feature learning as the motivation for using deep learning approaches. We propose using a limited number of EEG channels as an input for a deep neural network. In the research, we confirm that our electrode selection enhances the learning process of the convolutional neural network. The classification accuracy for the reduced subset of electrodes yields results comparable to the full dataset in a significantly shorter time—the average learning time 58% faster using our proposed strategy.
Aleksandra Dura, Agnieszka Wosiak
KES2
2021 Using semantic enrichment methods in expert search system for recruitment process in IT corporation
abstract
The problem of intelligent information retrieval and semantic enrichment becomes more and more popular due to the difficulty of searching and analyzing large text datasets. The common approach assumes user manual queries in natural language. Various semantic enrichment methods and intelligent text searching allow obtaining more accurate results leading to broader knowledge and user satisfaction. This research presents state-of-the-art methods of searching with enrichment and building rankings of results for the expert recruitment process in IT industry. The proposed model implements full-text search, semantic enrichment, and machine learning to match experts with job offers. Different data sources on expert competencies were used, including curricula vitae, historical data, and Internet resources. The testing results confirm an improvement in the search quality compared to the existing systems in the recruitment company.
Agnieszka Wosiak
KES1
2021 Automated extraction of information from Polish resume documents in the IT recruitment process
abstract
The problem of automated information retrieval and semantic analysis becomes more and more popular due to the difficulty of searching and analyzing large text datasets. Also, the recruitment process implies searching vast amounts of usually unstructured text files and makes human analysis challenging or even impossible to perform. This research presents a hybrid approach to automated information retrieval for the recruitment process in the IT industry. The proposed model implements a multi-module system in terms of low resource language dictionaries and complex linguistic dependencies in Polish. The experimental results confirmed our proposed solution. The keyword recognition for different resume sections was 60%-160% higher than for separate common tools.
Agnieszka Wosiak
KES1
2020 Automated feature selection for obstructive sleep apnea syndrome diagnosis
abstract
The paper presents a methodology of computer data analysis supporting medical diagnosis of obstructive sleep apnea (OSA) based on the results of polysomnography. Based on a database of 5114 patients, methods of detecting OSA with high accuracy have been developed. It has been also confirmed that obesity is an important risk factor. The methods of computer diagnostics have been compared with commonly used STOP-BANG questionnaire. The key stage of methodology referred to distinguished features that are most related to moderate or severe OSA presence, and are easy to gather at the same time. As a result of our studies we can conclude that it is possible to use smartwatch devices in order to develop a system of preliminary diagnostics of obstructive sleep apnea, which allows in the future for increased availability of apnea tests, reduced costs and earlier diagnosis.
Agnieszka Wosiak, Rafal Kowalski
KES1
2018 Imputing Missing Values for Improved Statistical Inference Applied to Intrauterine Growth Restriction Problem
abstract
The paper describes the study on the problem of missing values in medical data collected to discover new dependencies between parameters in children born with intrauterine growth restriction disorder.The aim of the research is to propose a procedure that may be taken to improve the medical inference in the presence of missing data.The approach with use of unconditional mean and k-nearest neighbor imputation has been applied.The experiments proved that application of missing data imputation in original dataset yields more valuable dependencies when compared to original data, maintaining the confidence interval for goodness of fit with the original distribution above 90%.The discovered dependencies in data may establish the basis for new treatment procedures of children with intrauterine growth restriction disorder. Index Terms-missing values
Agnieszka Wosiak, Kinga Glinka, Agata Zamecznik, Katarzyna Niewiadomska-Jarosik
FedCSIS1
2017 Preprocessing compensation techniques for improved classification of imbalanced medical datasets
abstract
The paper describes the study on the problem of applying classification techniques in medical datasets with a class imbalance.The aim of the research is to identify factors that negatively affect classification results and propose actions that may be taken to improve the performance.To alleviate the impact of uneven and complex class distribution, methods of balancing the datasets are proposed and compared.The experiments were conducted on five datasets -three binary and two multiclass.They comprise several data preprocessing methods applied on data and the classification with different techniques.The study shows that for some datasets there exists a combination of a certain preprocessing method and a classification technique which outperforms other approaches.For datasets with complex distribution or too many features the ratio of correctly predicted labels may be low regardless what resampling method and classification technique has been applied.
Agnieszka Wosiak, Sylwia Karbowiak
FedCSIS1
2017 Unsupervised feature selection using reversed correlation for improved medical diagnosis
abstract
Statistical inference has been usually used for medical data analysis, however in many cases it appears not to be efficient enough. Cluster analysis enables finding out groups of similar instances, for which statistical models can be built more effectively. In the paper a feature selection method for finding clustering attributes, which are supposed to improve performance of statistical analysis, is proposed. The method consists in selecting reversed correlated features as attributes of cluster analysis. The proposed technique has been evaluated by experiments done on real data sets of cardiovascular cases. Experiment results showed that the presented approach stimulates efficacy of statistical inference applied to medical diagnosis.
Agnieszka Wosiak, Danuta Zakrzewska
INISTA1
2016 Supervised and Unsupervised Machine Learning for Improved Identification of Intrauterine Growth Restriction Types
abstract
This paper concerns automated identification of intrauterine growth restriction (IUGR) types by use of machine learning methods.The research presents a comparison of supervised and unsupervised learning covering single and hybrid classification, as well as clustering.Supervised learning techniques included bagging with Naïve Bayes, k-nearest neighbours (kNN), C4.5 and SMO as base classifiers, random forest as a variant of bagging with a decision tree as a base classifier, boosting with Naïve Bayes, SMO, kNN and C4.5 as base classifiers, and voting by all single classifiers using majority as a combination rule, as well as five single classification strategies: kNN, C4.5, Naïve Bayes, random tree and sequential minimal optimization algorithm for training support vector machines.Unsupervised learning encompassed k-means and expectation-maximization algorithms.The major conclusion drawn from the study was that hybrid classifiers have demonstrated their potential ability to identify more accurately symmetrical and asymmetrical types of IUGR, whereas the unsupervised learning techniques produced the worst results.
Agnieszka Wosiak, Agata Zamecznik, Katarzyna Niewiadomska-Jarosik
FedCSIS1
2015 On integrating clustering and statistical analysis for supporting cardiovascular disease diagnosis
abstract
Statistical analysis of medical data plays significant role in medical diagnostics development. However in many cases the statistics is not effective enough. In the paper we consider combining statistical inference with clustering in the preprocessing phase of data analysis. The proposed methodology is checked on cardiovascular data and used for developing methods of early diagnosis of hypertension in children. Experiments, conducted on the real data, have demonstrated that the proposed hybrid approach allowed to discover relationships which have not been identified by using only the statistical methods. We have observed approximately 30% growth in the number of correlations between diagnosed attributes. Moreover all the obtained statistically significant dependencies were stronger in clusters rather than in the whole datasets.
Agnieszka Wosiak, Danuta Zakrzewska
FedCSIS1
2014 Feature Selection for Classification Incorporating Less Meaningful Attributes in Medical Diagnostics
abstract
In medical diagnostics there is a constant need of searching for new methods of attribute acquiring, but it is difficult to asses if these new features can support the existing ones and can be useful in medical inference.In the paper the methodology of discovering features which are less informative while considering independently, however meaningful for diagnosis making, is investigated.The proposed methodology can contribute to better use of attributes, which have not been considered in the diagnostics process so far.The experimental study, which concerns arterial hypertension as one of the civilization diseases demanding early detection and improved treatment is presented.The experiments confirmed that additional attributes enable obtaining the diagnostic results comparable to the ones received by using the most obvious features.
Agnieszka Wosiak, Danuta Zakrzewska
FedCSIS1