Lukasz Borchmann

dblp:184/8672 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Unchecked and Overlooked: Addressing the Checkbox Blind Spot in Large Language Models with CheckboxQA
Michal Turski, Mateusz Chilinski, Lukasz Borchmann
ICDAR (4)3
2024 STable: Table Generation Framework for Encoder-Decoder Models
abstract
Michał Pietruszka, Michał Turski, Łukasz Borchmann, Tomasz Dwojak, Gabriela Nowakowska, Karolina Szyndler, Dawid Jurkiewicz, Łukasz Garncarek. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Michal Pietruszka, Michal Turski, Lukasz Borchmann, Tomasz Dwojak, Gabriela Nowakowska, Karolina Szyndler, Dawid Jurkiewicz, Lukasz Garncarek
EACL (1)3
2023 Document Understanding Dataset and Evaluation (DUDE)
abstract
We call on the Document AI (DocAI) community to reevaluate current methodologies and embrace the challenge of creating more practically-oriented benchmarks. Document Understanding Dataset and Evaluation (DUDE) seeks to remediate the halted research progress in understanding visually-rich documents (VRDs). We present a new dataset1with novelties related to types of questions, answers, and document layouts based on multi-industry, multi-domain, and multi-page VRDs of various origins, and dates. Moreover, we are pushing the boundaries of current methods by creating multi-task and multi-domain evaluation setups that more accurately simulate real-world situations where powerful generalization and adaptation under low-resource settings are desired. DUDE aims to set a new standard as a more practical, long-standing benchmark for the community, and we hope that it will lead to future extensions and contributions that address real-world challenges. Finally, our work illustrates the importance of finding more efficient ways to model language, images, and layout in DocAI.
Jordy Van Landeghem, Rafal Powalski, Rubèn Tito, Dawid Jurkiewicz, Matthew B. Blaschko, Lukasz Borchmann, Mickaël Coustaty, Marie-Francine Moens, Michal Pietruszka, Bertrand Anckaert, Tomasz Stanislawek, Pawel Józiak, Ernest Valveny
ICCV6
2023 ICDAR 2023 Competition on Document UnderstanDing of Everything (DUDE)
Jordy Van Landeghem, Rubèn Tito, Lukasz Borchmann, Michal Pietruszka, Dawid Jurkiewicz, Rafal Powalski, Pawel Józiak, Sanket Biswas, Mickaël Coustaty, Tomasz Stanislawek
ICDAR (2)3
2022 Sparsifying Transformer Models with Trainable Representation Pooling
abstract
We propose a novel method to sparsify attention in the Transformer model by learning to select the most-informative token representations during the training process, thus focusing on the task-specific parts of an input.A reduction of quadratic time and memory complexity to sublinear was achieved due to a robust trainable top-k operator.Our experiments on a challenging long document summarization task show that even our simple baseline performs comparably to the current SOTA, and with trainable pooling we can retain its top quality, while being 1.8× faster during training, 4.5× faster during inference and up to 13× more computationally efficient in the decoder. 1
Michal Pietruszka, Lukasz Borchmann, Lukasz Garncarek
ACL (1)2
2021 Successive Halving Top-k Operator
abstract
We propose a differentiable successive halving method of relaxing the top-k operator, rendering gradient-based optimization possible. The need to perform softmax iteratively on the entire vector of scores is avoided using a tournament-style selection. As a result, a much better approximation of top-k and lower computational cost is achieved compared to the previous approach.
Michal Pietruszka, Lukasz Borchmann, Filip Gralinski
AAAI2
2021 Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer
Rafal Powalski, Lukasz Borchmann, Dawid Jurkiewicz, Tomasz Dwojak, Michal Pietruszka, Gabriela Palka
ICDAR (2)2
2021 Dynamic Boundary Time Warping for sub-sequence matching with few examples
Lukasz Borchmann, Dawid Jurkiewicz, Filip Gralinski, Tomasz Górecki
Expert Syst. Appl.1
2020 From Dataset Recycling to Multi-Property Extraction and Beyond
abstract
This paper investigates various Transformer architectures on the WikiReading Information Extraction and Machine Reading Comprehension dataset. The proposed dual-source model outperforms the current state-of-the-art by a large margin. Next, we introduce WikiReading Recycled - a newly developed public dataset, and the task of multiple-property extraction. It uses the same data as WikiReading but does not inherit its predecessor’s identified disadvantages. In addition, we provide a human-annotated test set with diagnostic subsets for a detailed analysis of model performance.
Tomasz Dwojak, Michal Pietruszka, Lukasz Borchmann, Jakub Chledowski, Filip Gralinski
CoNLL3
2016 "He Said She Said" ― a Male/Female Corpus of Polish
Filip Gralinski, Lukasz Borchmann, Piotr Wierzchon
LREC2