VLDB 2026 Research / reviewers in the wild / expert
Barbara Probierz
dblp:150/2767
· DBLP profile ↗
19ranked-venue papers
11as first author
13since 2021 · last 2026
0000-0002-5122-2645ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 11 first-author · 13 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fine-Tuning of Multilingual Language Models for Low-Resource Smishing Detection Using LoRA
Natalia Krawczyk, Barbara Probierz, Jan Kozak |
ACIIDS (1) | 2 |
| 2025 | Emotion and Phrase-Based Patterns in Smishing: A Feature-Driven Detection FrameworkabstractSmishing, or SMS-based phishing, remains a significant threat to mobile users due to its use of concise and emotionally manipulative language. These messages often rely on psychological cues and high-risk phrases that are not fully captured by traditional feature engineering. This study proposes an improved smishing detection framework that adds emotion-based labels and phrase-level risk indicators to embedding-based models. We used GPT-4.5 to annotate emotional dimensions. We also extracted targeted phrasal indicators based on known smishing patterns. Experiments on a unified dataset of over 22000 SMS messages show that combining emotion-based and phrase-level features consistently improves classification accuracy. For example, in Logistic Regression with TF-IDF, accuracy increased from 0.9186 to 0.9221 after adding both types of features. Gradient Boosting with TF-IDF showed the largest performance gain, while Random Forest achieved the highest absolute accuracy in this setting (0.9432). For Word2Vec embeddings, similar patterns were observed: Logistic Regression improved the most (from 0.8721 to 0.8817), while XGBoost reached the highest overall accuracy (0.9447) with sentiment-enhanced input. However, in some cases, risk-based features led to a slight drop in performance, suggesting potential feature noise in highly optimized or shallow models. These findings indicate that combining emotion-driven and pattern-based features with classical embeddings enhances phishing detection performance, particularly in short messages where traditional methods may lack context awareness. Natalia Krawczyk, Barbara Probierz, Jan Kozak |
KES | 2 |
| 2024 | Identification of Users in a Gambling Problem with the Use of Machine Learning
Tomasz Jach, Barbara Probierz, Jan Kozak, Piotr Stefanski, Grzegorz Dziczkowski, Anita Hrabia, Przemyslaw Juszczuk, Szymon Glowania, Gabriel Wolek, Wojciech Sznapka, Lukasz Swierk, Natalia Joniec |
ACIIDS (2) | 2 |
| 2024 | Automation of the Analysis of Medical Interviews to Improve Diagnoses Using NLP for Medicine
Barbara Probierz, Aleksandra Stras |
ACIIDS (1) | 1 |
| 2024 | Detection of Candidate Skills from Job Offers and Comparison with ESCO Database
Grzegorz Dziczkowski, Barbara Probierz, Grzegorz Madyda |
ICCCI (1) | 2 |
| 2023 | Emotion Detection from Text in Social Networks
Barbara Probierz, Jan Kozak, Przemyslaw Juszczuk |
ACIIDS (1) | 1 |
| 2023 | Sign language interpreting - relationships between research in different areas - overviewabstractTranslation from the national language into sign language is an extremely important area of research and practice, which aims to ensure communication between deaf or hard of hearing people and the hearing community.The article provides an overview of the most important research on sign language interpretation conducted in various research areas.The latest scientific and theoretical achievements were presented, which contribute to a better understanding of the subject of sign language translation and the improvement of the quality of translation services.Our main goal is to identify outstanding areas of interdisciplinary research related to sign language translation and to identify links between these studies conducted in different areas.The conclusions of the article aim to broaden the knowledge and awareness of sign language translation and to identify areas that require further research and development.The work is linked to a project related to the application of machine learning in increasing accessibility for deaf people. Barbara Probierz, Jan Kozak, Adam Piasecki, Angelika Podlaszewska |
FedCSIS | 1 |
| 2023 | Knowledge graphs to an analysis and visualization of texts from scientific articlesabstractReviewing the literature is one of the key elements of scientific research that allows you to identify existing solutions and research niches. However, it can be difficult for researchers to find relevant scientific articles related to the research topic. A limited number of available sources of information, diverse ways of describing them, and a multitude of scientific publications mean that scientists often have to spend a lot of time and effort to find the information they need. For this reason, we propose a solution for text analysis and its presentation in the form of a graph visualization. In our research, we used Natural Language Processing (NLP) methods and word weighting measures such as TF and TF-IDF. In addition, knowledge graphs were used to present the results visually. The conducted research was based on the analysis of the content of scientific articles, which allowed to draw important conclusions related to the presentation of texts in graphic form. Experimental results also identified potential methods and suggestions for literature reviews in specific fields of science. The analysis of the methods used and the results obtained allowed for a better understanding of the potential of natural language processing and graph text representation in the analysis of scientific articles. Barbara Probierz, Jan Kozak |
KES | 1 |
| 2022 | A Comparative Study of Classification and Clustering Methods from Text of Books
Barbara Probierz, Jan Kozak, Anita Hrabia |
ACIIDS (2) | 1 |
| 2022 | Modelling an IT solution to anonymise selected data processed in digital documentsabstractAllowing access to real legal documents is an important element both for the development of science and the judiciary.On the other hand, protecting information about citizens or organizations, that appear in these documents, is crucial and required by law.Therefore, before the documents are distributed, the data anonymisation process should be carried out.Unfortunately, there is no perfect tool that can automatically anonymise documents in such a way, that the main concept of the document is preserved; especially in the case of documents written in inflectional language.The aim of this article is to show how important (and at the same time how difficult) is the task to identify personal or corporate data of a client, as well as other related personal data in documents that are subject to legal protection.We conducted research aimed at assessing the usefulness of IT techniques as well as decision rules and patterns in the anonymisation of legal documents.A set of real legal documents written in Polish was used for the research in which we identified selected types of data that need to be anonymised.Eventually, the obtained results were assessed by field experts.Additionally, in order to verify the effectiveness of the proposed solution, we conducted research on a set of 50,000 false identities with names, company names, addresses and other confidential information.The collection was created using Fake Name Generator 1 .The obtained results from both experiments confirmed that the solutions we proposed is accurate even in the case of real legal documents. Barbara Probierz, Tomasz Jach, Jan Kozak, Radoslaw Pacud, Tomasz Turek |
FedCSIS | 1 |
| 2022 | Clustering of scientific articles using natural language processingabstractWith the development of the Internet, the number of scientific articles published online has also increased significantly. In recent years there has been a steady increase in the number of articles available online. Despite providing keywords and preparing abstracts, it is often a problem for articles to be properly matched during editorial process - to match the article to the competence of a particular editor. It is also a hassle for researchers who want to be kept informed about the most relevant articles. In paper, we propose the use of natural language processing (NLP) in the process of clustering scientific articles. Our study looks at the impact of clustering on the actual keywords provided by the authors – we analyse both the application of NLP to the abstract itself and to the introduction, followed by clustering using the K–means algorithm. As a result of experiments conducted on more than 1500 scientific articles, we have shown that our proposed approach allows articles to be approximated by their subject matter as a result of the clustering performed. The obtained results show that the best distribution of scientific articles can be obtained using the TF-IDF measure, and the worst - using TF measure. Barbara Probierz, Jan Kozak, Anita Hrabia |
KES | 1 |
| 2021 | Adaptive Goal Function of Ant Colony Optimization in Fake News Detection
Barbara Probierz, Jan Kozak, Piotr Stefanski, Przemyslaw Juszczuk |
ICCCI | 1 |
| 2021 | Rapid detection of fake news based on machine learning methodsabstractNowadays, it is very important to quickly recognize the false information referred to as fake news. This is especially important in the case of news appearing on the Internet because of its wide and rapid spreading. It is equally important to be able to initially classify news as fake or true based on the title itself. In this paper, we propose an approach to classifying news based on the title without analyzing the other aspects. The obtained results will be compared with classification based on the whole news text. The goal of this work is to propose a method that balances between data analysis time and quality of classification in fake news prediction. We use natural language processing (NLP) to describe the title and text of the news. This is a complex process, requiring good analysis to be applied to classification. Therefore, the use of complex classifiers – in this case, classical ensemble methods – has been proposed in order to achieve a high quality of classification (measured by popular measure). In this paper, we present analyses of a real data set and results of news classification using the proposed model – including an ensemble of classifiers and single classifiers. Barbara Probierz, Piotr Stefanski, Jan Kozak |
KES | 1 |
| 2020 | The hybrid ant colony optimization and ensemble method for solving the data stream e-mail foldering problem
Jan Kozak, Przemyslaw Juszczuk, Barbara Probierz |
Neural Comput. Appl. | 3 |
| 2019 | Discovery of Leaders and Cliques in the Organization Based on Social Network Analysis
Barbara Probierz |
ICCCI (2) | 1 |
| 2018 | The Mechanism to Predict Folders in Automatic Classification Email Messages to Folders in the Mailboxes
Barbara Probierz |
ICCCI (2) | 1 |
| 2015 | Adaptive Ant Colony Decision Forest in Automatic Categorization of Emails
Urszula Boryczka, Barbara Probierz, Jan Kozak |
ACIIDS (1) | 2 |
| 2015 | A New Algorithm to Categorize E-mail Messages to Folders with Social Networks Analysis
Urszula Boryczka, Barbara Probierz, Jan Kozak |
ICCCI (2) | 2 |
| 2014 | An Ant Colony Optimization Algorithm for an Automatic Categorization of Emails
Urszula Boryczka, Barbara Probierz, Jan Kozak |
ICCCI | 2 |