VLDB 2026 Research / reviewers in the wild / expert
Marek Kozlowski
dblp:122/1782
· DBLP profile ↗
14ranked-venue papers
4as first author
4since 2021 · last 2024
0000-0002-6313-8387ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Detection of AI-Generated Emails - A Case StudyabstractThis work-in-progress paper investigates the problem of assessing and detecting if a text was written by a human or if it was generated by a language model. In our case study, we focused on email messages. For the purpose of experiments, we used a combination of publicly available email datasets with our in-house data, containing in total over 10k emails. Then, we generated their “copies” using large language models (LLMs) with specific prompts. We experimented with various classifiers and feature spaces. We achieved encouraging results, with the F1-scores of almost 0.99 for email messages in English and over 0.92 for the ones in Polish, using Random Forest as a classifier. We found that the detection model relied strongly on typographic and orthographic (spelling) imperfections of the analyzed emails and on statistics of sentence lengths. We also observed the inferior results obtained for Polish, highlighting a need for research in the direction of languages underrepresented in training models. Pawel Gryka, Kacper T. Gradon, Marek Kozlowski, Milosz Kutyla, Artur Janicki |
ARES | 3 |
| 2024 | Impact of Spelling and Editing Correctness on Detection of LLM-Generated EmailsabstractIn this paper, we investigated the impact of spelling and editing correctness on the accuracy of detection if an email was written by a human or if it was generated by a language model.As a dataset, we used a combination of publicly available email datasets with our in-house data, with over 10k emails in total.Then, we generated their "copies" using large language models (LLMs) with specific prompts.As a classifier, we used random forest, which yielded the best results in previous experiments.For English emails, we found a slight decrease in evaluation metrics if error-related features were excluded.However, for the Polish emails, the differences were more significant, indicating a decline in prediction quality by around 2% relative.The results suggest that the proposed detection method can be equally effective for English even if spelling-and grammar-checking tools are used.As for Polish, to compensate for error-related features, additional measures have to be undertaken. Pawel Gryka, Kacper T. Gradon, Marek Kozlowski, Milosz Kutyla, Artur Janicki |
FedCSIS | 3 |
| 2023 | Hybrid retrievers with generative re-rankersabstractThe passage retrieval task was announced during PolEval 2022 (SemEval-inspired evaluation campaign for natural language processing tools for Polish).Passage retrieval is a crucial part of modern open-domain question answering systems that rely on precise and efficient retrieval components to identify passages that contain correct answers.Our solution to this task is a multi-stage neural information retrieval system.The first stage consists of a candidate passage retrieval step in which passages are retrieved using federated search over sparse (BM25) and dense indexes (two FAISS indexes built using bi-encoder type retrievers based on Polish RoBERTa models).The second stage consists of a re-ranking step of the previously selected passages with a neural model, mt5-13b-mmarco.The model scores each passage by its relevance to a given query.The highest-scoring passages are then retained as the final result.Our system achieved second place in the competition. Marek Kozlowski |
FedCSIS | 1 |
| 2022 | Predicting the Costs of Forwarding Contracts Using XGBoost and a Deep Neural NetworkabstractThis article presents an application of an XGBoost and deep neural network ensemble as a solution for a task assigned at the FedCSIS 2022 Challenge: Predicting the Costs of Forwarding Contracts.We demonstrate that prediction quality can be improved by combining the two approaches.We present a neural network architecture based on three independent flows.We then discuss the influence of long short-term memory units on the risk of overfitting.Finally, we show that the static XGBoost model can complement a neural network that processes dynamic data. Lukasz Podlodowski, Marek Kozlowski |
FedCSIS | 2 |
| 2019 | Application of XGBoost to the cyber-security problem of detecting suspicious network traffic eventsabstractThis paper presents an application of XGBoost as a solution for a task associated with the IEEE BigData2019 Cup: Suspicious Network Event Recognition. As has been shown in the paper, the high-quality classification model can be based on independent predictions of each component in the sequence of network traffic events, then analyzed with statistical aggregation functions to generate the final prediction. We also propose the approach to this problem including handling high dimensionality space of IP addresses through encoding octets separately. Lukasz Podlodowski, Marek Kozlowski |
IEEE BigData | 2 |
| 2019 | Clustering of semantically enriched short textsabstractThe paper is devoted to the issue of clustering small sets of very short texts. Such texts are often incomplete and highly inconclusive, so establishing a notion of proximity between them is a challenging task. In order to cope with polysemy we adapt the SenseSearcher algorithm (SnS), by Kozlowski and Rybinski in Computational Intelligence 33(3): 335–367, 2017b . In addition, we test the possibilities of improving the quality of clustering ultra-short texts by means of enriching them semantically. We present two approaches, one based on neural-based distributional models, and the other based on external knowledge resources. The approaches are tested on SnSRC and other knowledge-poor algorithms. Marek Kozlowski, Henryk Rybinski |
J. Intell. Inf. Syst. | 1 |
| 2018 | Deep Learning Approaches towards Book Covers Classification
Przemyslaw Buczkowski, Antoni Sobkowicz, Marek Kozlowski |
ICPRAM | 3 |
| 2017 | Semantic Enriched Short Text Clustering
Marek Kozlowski, Henryk Rybinski |
ISMIS | 1 |
| 2017 | Word Sense Induction with Closed Frequent TermsetsabstractThe article is devoted to the problem of word sense induction. We propose a method for inducing senses from a raw text corpus. The proposed sense induction algorithm (called SenseSearcher, or SnS) is based on closed frequent sets, and as a result, it provides a multilevel sense representation. To a large extent, it is a knowledge‐poor approach, as it does not need any kind of structured knowledge base about senses and there is no deep language knowledge embedded. By discovering a hierarchy of senses, the algorithm enables identifying subsenses (fine‐grained senses). SnS discovers not only frequent (dominating) senses but also infrequent ones (dominated). The method was evaluated in two main areas: lexicography and information retrieval. With the use of the SnS algorithm, we provide a tool able to induce from a textual corpus a structure of senses, with a varying number of granularity levels. In the area of information retrieval, SnS can be used for clustering search result, according to the discovered senses. The experiments have shown that SnS performs better than the methods participating in the SemEval2013 WSI Task 11 competition, and most of the known search result clustering methods. Marek Kozlowski, Henryk Rybinski |
Comput. Intell. | 1 |
| 2017 | Intelligent information processing for building university knowledge baseabstractThere are many ready-to-use software solutions for building institutional scientific information platforms, most of which have functionality well suited to repository needs. However, there have already been discussions about various problems with institutional digital libraries. As a remedy, an approach that is researcher-centric (rather than document-centric) has been proposed recently in some systems. This paper is devoted to research aimed at tools for building knowledge bases for university research. We focus on the AI methods that have been elaborated and applied practically within our platform for building such knowledge bases. In particular we present a novel approach to data acquisition and the semantic enrichment of the acquired data. In addition, we present the algorithms applied in the real life system for experts profiling and retrieval. Jakub Koperwas, Lukasz Skonieczny, Marek Kozlowski, Piotr Andruszkiewicz, Henryk Rybinski, Waclaw Struk |
J. Intell. Inf. Syst. | 3 |
| 2016 | A novel method for dictionary translationabstractThe paper addresses the problem of automatic dictionary translation.The proposed method translates a dictionary by means of mining repositories in the source and target languages, without any directly given relationships connecting the two languages. It consists of two stages: (1) translation by lexical similarity, where words are compared graphically, and (2) translation by semantic similarity, where contexts are compared. In the experiments Polish and English version of Wikipedia were used as text corpora. The method and its phases are thoroughly analyzed. The results allow implementing this method in human-in-the-middle systems. Robert Krajewski, Henryk Rybinski, Marek Kozlowski |
J. Intell. Inf. Syst. | 3 |
| 2016 | A recommender system of reviewers and experts in reviewing problems
Jaroslaw Protasiewicz, Witold Pedrycz, Marek Kozlowski, Slawomir Dadas, Tomasz Stanislawek, Agata Kopacz, Malgorzata Galezewska |
Knowl. Based Syst. | 3 |
| 2014 | AI Platform for Building University Research Knowledge Base
Jakub Koperwas, Lukasz Skonieczny, Marek Kozlowski, Piotr Andruszkiewicz, Henryk Rybinski, Waclaw Struk |
ISMIS | 3 |
| 2014 | A Seed Based Method for Dictionary Translation
Robert Krajewski, Henryk Rybinski, Marek Kozlowski |
ISMIS | 3 |