VLDB 2026 Research / reviewers in the wild / expert
Aleksander Smywinski-Pohl
dblp:157/7699 · also Aleksander Pohl
· DBLP profile ↗
16ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0001-6684-0748ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 5 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Targum - a Multilingual New Testament Translation Corpus
Maciej Rapacz, Aleksander Smywinski-Pohl |
LREC | 2 |
| 2025 | Lemmatization of Polish Multi-word ExpressionsabstractThis paper explores the lemmatization of multi-word expressions (MWEs) and proper names in Polish -tasks complicated by linguistic irregularities and historical factors.Instead of using rule-based methods, we apply a machine learning approach with fine-tuned plT5 and mT5 models.We trained and validated the models on enhanced gold-standard data from the 2019 Pol-Eval task and evaluated the impact of additional fine-tuning on a silver-standard dataset derived from Wikipedia.Two setups were tested: one without context, and one using leftside context of the target MWE.Our best model achieved 86.23% AccCS (Accuracy Case-Sensitive), 89.43% AccCI (Accuracy Case-Insensitive), and a combined score of 88.79%, setting a new state-of-the-art for Polish MWE and named entity lemmatization, as confirmed by the PolEval maintainers.We also evaluated optimization and quantization techniques to reduce model size and inference time with modest quality loss. Magdalena Król, Aleksander Smywinski-Pohl, Zbigniew Kaleta, Pawel Lewkowicz |
EMNLP | 2 |
| 2025 | Are manual annotations necessary for statutory interpretations retrieval?abstractLegal research often involves finding cases where judges interpret legal concepts, helping professionals cite precedents and aiding public understanding. Tomer Libal, Aleksander Smywinski-Pohl, Adam Kaczmarczyk, Magdalena Król |
ICAIL | 2 |
| 2024 | A Legal Assistant for Accountable Decision-MakingabstractWe describe a chatbot system capable of providing legal assistance and advice. It combines several approaches to processing legals problems, including formal logic-driven and the use of language models. This combination allows building a chatbot that can realize two major goals: perform legal reasoning with 100% accuracy and provide the user with explanations allowing them to understand the legal problems connected with their case. Adam Kaczmarczyk, Tomer Libal, Aleksander Smywinski-Pohl |
JURIX | 3 |
| 2023 | PolEval 2022/23 Challenge Tasks and ResultsabstractThis paper summarizes the 2022/2023 edition of PolEval -an evaluation campaign for natural language processing tools for Polish.We describe the tasks organized in this edition, which are: Punctuation prediction from conversational language, Abbreviation disambiguation and Passage Retrieval.We also discuss the datasets prepared for each of the tasks, evaluation metrics chosen to rank the submissions and also sum up the approaches chosen by the participants to tackle the tasks. Lukasz Kobylinski, Maciej Ogrodniczuk, Piotr Rybak, Piotr Przybyla, Piotr Pezik, Agnieszka Mikolajczyk, Wojciech Janowski, Michal Marcinczuk, Aleksander Smywinski-Pohl |
FedCSIS | 9 |
| 2023 | Giving Examples Instead of Answering Questions: Introducing Legal Concept-Example SystemsabstractQuestion-Answering Systems (QASs) have seen a big development in recent years and various attempts have been made to extend them to the legal domain. Nevertheless, the needs and methodology of legal research relies often less on getting answers and more on finding positive and negative examples for certain legal concepts. In this paper, we introduce a sub-category within QASs that focuses on such legal tasks and design a methodology for the automated production of such systems. Tomer Libal, Aleksander Smywinski-Pohl |
JURIX | 2 |
| 2021 | Automatic extraction of amendments from polish statutory lawabstractThe article discusses the problem of automatic detection of amendments found in the Polish statutory law. We treat the problem as a token-classification task and we introduce a scheme constructed by analysis of more than 200 amending bills. We apply recent neural architectures such as BERT and BiRNN to the task of token classification. The achieved results of all models are very high as micro average F1 score ranges from 96.3% to 98.2% for BiRNN. The presented solution is a first step towards fully automatic structuring and application of amendments in the Polish statutory law. Aleksander Smywinski-Pohl, Mateusz Piech, Zbigniew Kaleta, Krzysztof Wrobel 0002 |
ICAIL | 1 |
| 2021 | Improving classifier training efficiency for automatic cyberbullying detection with Feature Density
Juuso Kalevi Kristian Eronen, Michal Ptaszynski, Fumito Masui, Aleksander Smywinski-Pohl, Gniewosz Leliwa, Michal Wroczynski |
Inf. Process. Manag. | 4 |
| 2021 | Meta-User2Vec model for addressing the user and item cold-start problem in recommender systemsabstractAbstract The cold-start scenario is a critical problem for recommendation systems, especially in dynamically changing domains such as online news services. In this research, we aim at addressing the cold-start situation by adapting an unsupervised neural User2Vec method to represent new users and articles in a multidimensional space. Toward this goal, we propose an extension of the Doc2Vec model that is capable of representing users with unknown history by building embeddings of their metadata labels along with item representations. We evaluate our proposed approach with respect to different parameter configurations on three real-world recommendation datasets with different characteristics. Our results show that this approach may be applied as an efficient alternative to the factorization machine-based method when the user and item metadata are used and hence can be applied in the cold-start scenario for both new users and new items. Additionally, as our solution represents the user and item labels in the same vector space, we can analyze the spatial relations among these labels to reveal latent interest features of the audience groups as well as possible data biases and disparities. Joanna Misztal-Radecka, Bipin Indurkhya, Aleksander Smywinski-Pohl |
User Model. User Adapt. Interact. | 3 |
| 2020 | Data Augmentation for Sentiment Analysis in English - The Online Approach
Michal Jungiewicz, Aleksander Smywinski-Pohl |
ICANN (2) | 2 |
| 2019 | Automatic Construction of a Polish Legal Dictionary with Mappings to Extra-Legal Terms Established via Word EmbeddingsabstractThe primary objective of this research is finding correspondence between legal and extra-legal terms in Polish by employing unsupervised methods, such as statistics and word embeddings. We investigate the possibility to construct a legal dictionary automatically by employing statistical methods for identifying the legal terms (including multi-word entities) and then finding correspondence between these terms and extralegal terminology used by laymen, by employing word embeddings inducing algorithms. We compare two popular libraries word2vec and GloVe in a synthetic experiment showing the superiority of word2vec CBOW negative sampling variant in the described problem. Aleksander Smywinski-Pohl, Karol Lasocki, Krzysztof Wrobel 0002, Marek Strzalta |
ICAIL | 1 |
| 2019 | Application of Character-Level Language Models in the Domain of Polish Statutory Law
Aleksander Smywinski-Pohl, Krzysztof Wrobel 0002, Karol Lasocki, Michal Jungiewicz |
JURIX | 1 |
| 2018 | Improving Text Classification with Vectors of Reduced PrecisionabstractThis paper presents the analysis of the impact of a floating-point number precision reduction on the quality of text classification. The precision reduction of the vectors representing the data (e.g. TF-IDF representation in our case) allows for a decrease of computing time and memory footprint on dedicated hardware platforms. The impact of precision reduction on the classification quality was performed on 5 corpora, using 4 different classifiers. Also, dimensionality reduction was taken into account. Results indicate that the precision reduction improves classification accuracy for most cases (up to 25% of error reduction). In general, the reduction from 64 to 4 bits gives the best scores and ensures that the results will not be worse than with the full floating-point representation. Krzysztof Wrobel 0002, Maciej Wielgosz, Marcin Pietron, Michal Karwatowski, Jerzy Duda, Aleksander Smywinski-Pohl |
ICAART (2) | 6 |
| 2015 | Comparison of language models trained on written texts and speech transcripts in the context of automatic speech recognitionabstractWe investigate whether language models used in automatic speech recognition (ASR) should be trained on speech transcripts rather than on written texts.By calculating log-likelihood statistic for part-of-speech (POS) n-grams, we show that there are significant differences between written texts and speech transcripts.We also test the performance of language models trained on speech transcripts and written texts in ASR and show that using the former results in greater word error reduction rates (WERR), even if the model is trained on much smaller corpora.For our experiments we used the manually labeled one million subcorpus of the National Corpus of Polish and an HTK acoustic model. Sebastian Dziadzio, Aleksandra Nabozny, Aleksander Smywinski-Pohl, Bartosz Ziólko |
FedCSIS | 3 |
| 2013 | Knowledge-based Named Entity Recognition in Polish
Aleksander Smywinski-Pohl |
FedCSIS | 1 |
| 2012 | Improving Wikipedia Miner Word Sense Disambiguation Algorithm
Aleksander Smywinski-Pohl |
FedCSIS | 1 |