VLDB 2026 Research / reviewers in the wild / expert
Filip Gralinski
dblp:07/3244
· DBLP profile ↗
14ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0001-8066-4533ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Oddballness: universal anomaly detection with language modelsabstractWe present a new method to detect anomalies in texts (in general: in sequences of any data), using language models, in a totally unsupervised manner. The method considers probabilities (likelihoods) generated by a language model, but instead of focusing on low-likelihood tokens, it considers a new metric defined in this paper: oddballness. Oddballness measures how “strange” a given token is according to the language model. We demonstrate in grammatical error detection tasks (a specific case of text anomaly detection) that oddballness is better than just considering low-likelihood events, if a totally unsupervised setup is assumed. Filip Gralinski, Ryszard Staruch, Krzysztof Jurkiewicz |
COLING | 1 |
| 2025 | reVISION: A Polish Benchmark for Evaluating Vision-Language Models on Multimodal National Exam DataabstractVision-Language (VL) models have gained significant popularity in recent years for tasks involving data extraction and image recognition.In this paper, we introduce reVISION, a largescale Polish benchmark for evaluating such models, comprising over 39k questions extracted from Polish national exams.We assess both the models' general knowledge and their ability to handle tasks that go beyond plain text, including tables and the recognition of specialized visual objects.Our study also examines the effects of image resizing and the inclusion of additional context via OCR-extracted text.To explore the correlation between human and model performance, we compare their results across multiple exam years. Michal Ciesiólka, Filip Gralinski |
FedCSIS | 2 |
| 2023 | Modeling Spaced Repetition with LSTMs
Jakub Pokrywka, Marcin Biedalak, Filip Gralinski, Krzysztof Biedalak |
CSEDU (2) | 3 |
| 2023 | CCpdf: Building a High Quality Corpus for Visually Rich Documents from Web Crawl Data
Michal Turski, Tomasz Stanislawek, Karol Kaczmarek, Pawel Dyda, Filip Gralinski |
ICDAR (3) | 5 |
| 2022 | Using Transformer models for gender attribution in PolishabstractGender identification is the task of predicting the gender of an author of a given text.Some languages, including Polish, exhibit gender-revealing syntactic expression.In this paper, we investigate machine learning methods for gender identification in Polish.For the evaluation, we use large (780M words) corpus "He Said She Said", created by grepping (for author's gender identification) gender-revealing syntactic expressions and normalizing all these expressions to masculine form (for preventing classifiers from using syntactic features).In this work, we evaluate TF-IDF based, fastText, LSTM and RoBERTa models, differentiating self-contained and non-selfcontained approaches.We also provide a human baseline.We report large improvements using pre-trained RoBERTa models and discuss the possible contamination of test data for the best pre-trained model. Karol Kaczmarek, Jakub Pokrywka, Filip Gralinski |
FedCSIS | 3 |
| 2022 | Temporal Language Modeling for Short Text Document Classification with TransformersabstractLanguage models are typically trained on solely text data, not utilizing document timestamps, which are available in most internet corpora.In this paper, we examine the impact of incorporating timestamp into transformer language model in terms of downstream classification task and masked language modeling on 2 short texts corpora.We examine different timestamp components: day of the month, month, year, weekday.We test different methods of incorporating date into the model: prefixing date components into text input and adding trained date embeddings.Our study shows, that such a temporal language model performs better than a regular language model for both documents from training data time span and unseen time span.That holds true for classification and language modeling.Prefixing date components into text performs no worse than training special date components embeddings. Jakub Pokrywka, Filip Gralinski |
FedCSIS | 2 |
| 2021 | Successive Halving Top-k OperatorabstractWe propose a differentiable successive halving method of relaxing the top-k operator, rendering gradient-based optimization possible. The need to perform softmax iteratively on the entire vector of scores is avoided using a tournament-style selection. As a result, a much better approximation of top-k and lower computational cost is achieved compared to the previous approach. Michal Pietruszka, Lukasz Borchmann, Filip Gralinski |
AAAI | 3 |
| 2021 | LAMBERT: Layout-Aware Language Modeling for Information ExtractionabstractWe introduce a simple new approach to the problem of understanding documents where non-trivial layout influences the local semantics. To this end, we modify the Transformer encoder architecture in a way that allows it to use layout features obtained from an OCR system, without the need to re-learn language semantics from scratch. We only augment the input of the model with the coordinates of token bounding boxes, avoiding, in this way, the use of raw images. This leads to a layout-aware language model which can then be fine-tuned on downstream tasks. The model is evaluated on an end-to-end information extraction task using four publicly available datasets: Kleister NDA, Kleister Charity, SROIE and CORD. We show that our model achieves superior performance on datasets consisting of visually rich documents, while also outperforming the baseline RoBERTa on documents with flat layout (NDA \(F_{1}\) increase from 78.50 to 80.42). Our solution ranked first on the public leaderboard for the Key Information Extraction from the SROIE dataset, improving the SOTA \(F_{1}\)-score from 97.81 to 98.17. Lukasz Garncarek, Rafal Powalski, Tomasz Stanislawek, Bartosz Topolski, Piotr Halama, Michal Turski, Filip Gralinski |
ICDAR (1) | 7 |
| 2021 | Kleister: Key Information Extraction Datasets Involving Long Documents with Complex LayoutsabstractThe relevance of the Key Information Extraction (KIE) task is increasingly important in natural language processing problems. But there are still only a few well-defined problems that serve as benchmarks for solutions in this area. To bridge this gap, we introduce two new datasets (Kleister NDA and Kleister Charity). They involve a mix of scanned and born-digital long formal English-language documents. In these datasets, an NLP system is expected to find or infer various types of entities by employing both textual and structural layout features. The Kleister Charity dataset consists of 2,788 annual financial reports of charity organizations, with 61,643 unique pages and 21,612 entities to extract. The Kleister NDA dataset has 540 Non-disclosure Agreements, with 3,229 unique pages and 2,160 entities to extract. We provide several state-of-the-art baseline systems from the KIE domain (Flair, BERT, RoBERTa, LayoutLM, LAMBERT), which show that our datasets pose a strong challenge to existing models. The best model achieved an 81.77% and an 83.57% F1-score on respectively the Kleister NDA and the Kleister Charity datasets. We share the datasets to encourage progress on more in-depth and complex information extraction tasks. Tomasz Stanislawek, Filip Gralinski, Anna Wróblewska, Dawid Lipinski, Agnieszka Kaliska, Paulina Rosalska, Bartosz Topolski, Przemyslaw Biecek |
ICDAR (1) | 2 |
| 2021 | Dynamic Boundary Time Warping for sub-sequence matching with few examples
Lukasz Borchmann, Dawid Jurkiewicz, Filip Gralinski, Tomasz Górecki |
Expert Syst. Appl. | 3 |
| 2020 | From Dataset Recycling to Multi-Property Extraction and BeyondabstractThis paper investigates various Transformer architectures on the WikiReading Information Extraction and Machine Reading Comprehension dataset. The proposed dual-source model outperforms the current state-of-the-art by a large margin. Next, we introduce WikiReading Recycled - a newly developed public dataset, and the task of multiple-property extraction. It uses the same data as WikiReading but does not inherit its predecessor’s identified disadvantages. In addition, we provide a human-annotated test set with diagnostic subsets for a detailed analysis of model performance. Tomasz Dwojak, Michal Pietruszka, Lukasz Borchmann, Jakub Chledowski, Filip Gralinski |
CoNLL | 5 |
| 2016 | "He Said She Said" ― a Male/Female Corpus of Polish
Filip Gralinski, Lukasz Borchmann, Piotr Wierzchon |
LREC | 1 |
| 2010 | Mining Parenthetical Translations for Polish-English Lexica
Filip Gralinski |
CICLing | 1 |
| 2009 | An Environment for Named Entity Recognition and Translation
Filip Gralinski, Krzysztof Jassem, Michal Marcinczuk |
EAMT | 1 |