EDBT 2026 Demo / reviewers in the wild / expert
Tomasz Stanislawek
dblp:174/3652
· DBLP profile ↗
4ranked-venue papers in the field
1as first author
4since 2021 · last 2023
0000-0003-1046-7563ORCID · reported
Domains — venue-derived; a paper can count in several
Other / Interdisciplinary · 4 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | ICDAR 2023 Competition on Document UnderstanDing of Everything (DUDE)
Jordy Van Landeghem, Rubèn Tito, Lukasz Borchmann, Michal Pietruszka, Dawid Jurkiewicz, Rafal Powalski, Pawel Józiak, Sanket Biswas, Mickaël Coustaty, Tomasz Stanislawek |
ICDAR (2) | 10 |
| 2023 | CCpdf: Building a High Quality Corpus for Visually Rich Documents from Web Crawl Data
Michal Turski, Tomasz Stanislawek, Karol Kaczmarek, Pawel Dyda, Filip Gralinski |
ICDAR (3) | 2 |
| 2021 | LAMBERT: Layout-Aware Language Modeling for Information ExtractionabstractWe introduce a simple new approach to the problem of understanding documents where non-trivial layout influences the local semantics. To this end, we modify the Transformer encoder architecture in a way that allows it to use layout features obtained from an OCR system, without the need to re-learn language semantics from scratch. We only augment the input of the model with the coordinates of token bounding boxes, avoiding, in this way, the use of raw images. This leads to a layout-aware language model which can then be fine-tuned on downstream tasks. The model is evaluated on an end-to-end information extraction task using four publicly available datasets: Kleister NDA, Kleister Charity, SROIE and CORD. We show that our model achieves superior performance on datasets consisting of visually rich documents, while also outperforming the baseline RoBERTa on documents with flat layout (NDA \(F_{1}\) increase from 78.50 to 80.42). Our solution ranked first on the public leaderboard for the Key Information Extraction from the SROIE dataset, improving the SOTA \(F_{1}\)-score from 97.81 to 98.17. Lukasz Garncarek, Rafal Powalski, Tomasz Stanislawek, Bartosz Topolski, Piotr Halama, Michal Turski, Filip Gralinski |
ICDAR (1) | 3 |
| 2021 | Kleister: Key Information Extraction Datasets Involving Long Documents with Complex LayoutsabstractThe relevance of the Key Information Extraction (KIE) task is increasingly important in natural language processing problems. But there are still only a few well-defined problems that serve as benchmarks for solutions in this area. To bridge this gap, we introduce two new datasets (Kleister NDA and Kleister Charity). They involve a mix of scanned and born-digital long formal English-language documents. In these datasets, an NLP system is expected to find or infer various types of entities by employing both textual and structural layout features. The Kleister Charity dataset consists of 2,788 annual financial reports of charity organizations, with 61,643 unique pages and 21,612 entities to extract. The Kleister NDA dataset has 540 Non-disclosure Agreements, with 3,229 unique pages and 2,160 entities to extract. We provide several state-of-the-art baseline systems from the KIE domain (Flair, BERT, RoBERTa, LayoutLM, LAMBERT), which show that our datasets pose a strong challenge to existing models. The best model achieved an 81.77% and an 83.57% F1-score on respectively the Kleister NDA and the Kleister Charity datasets. We share the datasets to encourage progress on more in-depth and complex information extraction tasks. Tomasz Stanislawek, Filip Gralinski, Anna Wróblewska, Dawid Lipinski, Agnieszka Kaliska, Paulina Rosalska, Bartosz Topolski, Przemyslaw Biecek |
ICDAR (1) | 1 |