EDBT 2026 Demo / reviewers in the wild / expert
Dawid Wisniewski
dblp:217/9960
· DBLP profile ↗
15ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0003-1194-7921ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 11 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Feature Engineering Pipeline for Non-stationary Financial Time Series: A ZigZag-Based Approach to Forex Decision Support
Pawel Kierkosz, Pawel Kolec, Bartlomiej Rudowicz, Adam Nowacki, Dawid Wisniewski |
DEXA (2) | 5 |
| 2026 | ForMaT: Dataset for Visually-Grounded Multilingual PDF TranslationabstractWe present ForMaT (Format-Preserving Multilingual Translation), a parallel corpus of 3,956 PDFs across 15 language pairs that preserves original layout metadata proposed for multimodal machine translation. To ensure structural diversity in the dataset, we employ K-Medoids sampling over 45 geometric features, capturing complex elements like nested tables and formulas to focus only on visually diverse PDF documents. Our evaluation reveals that current MT systems struggle with spatial grounding and geometric synchronization, often losing the link between text and its visual context. ForMaT provides a benchmark for developing layout-aware translation models that integrate visual and textual context for high-fidelity document reconstruction. Michal Ciesiólka, Dawid Wisniewski, Adrian Charkiewicz, Kamil Guttmann |
EAMT (1) | 2 |
| 2026 | Beyond Semantics: Measuring Fine-Grained Emotion Preservation in Small Language Model-Based Machine TranslationabstractPreserving affective nuance remains a challenge in Machine Translation (MT), where semantic equivalence often takes precedence over emotional fidelity. This paper evaluates the performance of three state-of-the-art Small Language Models (SLMs) – EuroLLM, Aya Expanse, and Gemma – in maintaining fine-grained emotions during backtranslation. Using the GoEmotions dataset, which comprises Reddit comments across 28 distinct categories, we assess emotional preservation across five European languages: German, French, Spanish, Italian, and Polish. Specifically, we investigate (i) the inherent capability of these SLMs to retain emotional sentiment, (ii) the efficacy of emotion-aware prompting in improving preservation, and (iii) the performance of ModernBERT as a contemporary alternative to BERT for emotion classification in MT evaluation. Dawid Wisniewski, Igor Czudy |
EAMT (1) | 1 |
| 2025 | Do Not Change Me: On Transferring Entities Without Modification in Neural Machine Translation - a Multilingual PerspectiveabstractCurrent machine translation models provide us with high-quality outputs in most scenarios. However, they still face some specific problems, such as detecting which entities should not be changed during translation. In this paper, we explore the abilities of popular NMT models, including models from the OPUS project, Google Translate, MADLAD, and EuroLLM, to preserve entities such as URL addresses, IBAN numbers, or emails when producing translations between four languages: English, German, Polish, and Ukrainian. We investigate the quality of popular NMT models in terms of accuracy, discuss errors made by the models, and examine the reasons for errors. Our analysis highlights specific categories, such as emojis, that pose significant challenges for many models considered. In addition to the analysis, we propose a new multilingual synthetic dataset of 36,000 sentences that can help assess the quality of entity transfer across nine categories and four aforementioned languages. Dawid Wisniewski, Miko Pokrywka, Zofia Rostek |
MTSummit (1) | 1 |
| 2025 | Exploring the Feasibility of Multilingual Grammatical Error Correction with a Single LLM up to 9B parameters: A Comparative Study of 17 ModelsabstractRecent language models can successfully solve various language-related tasks, and many understand inputs stated in different languages. In this paper, we explore the performance of 17 popular models used to correct grammatical issues in texts stated in English, German, Italian, and Swedish when using a single model to correct texts in all those languages. We analyze the outputs generated by these models, focusing on decreasing the number of grammatical errors while keeping the changes small. The conclusions drawn help us understand what problems occur among those models and which models can be recommended for multilingual grammatical error correction tasks. We list six models that improve grammatical correctness in all four languages and show that Gemma 9B is currently the best performing one for the languages considered. Dawid Wisniewski, Antoni Solarski, Artur Nowakowski |
MTSummit (1) | 1 |
| 2025 | Boosting Dual Quality detection with AI-based social media analysis
Maksim Brzezinski, Maciej Niemir, Krzysztof Muszynski, Mateusz Lango, Dawid Wisniewski |
Inf. Process. Manag. | 5 |
| 2024 | FAME-MT Dataset: Formality Awareness Made Easy for Machine Translation PurposesabstractPeople use language for various purposes. Apart from sharing information, individuals may use it to express emotions or to show respect for another person. In this paper, we focus on the formality level of machine-generated translations and present \textbf{FAME-MT} – a dataset consisting of 11.2 million translations between 15 European source languages and 8 European target languages classified to formal and informal classes according to target sentence formality. This dataset can be used to fine-tune machine translation models to ensure a given formality level for 8 European target languages considered. We describe the dataset creation procedure, the analysis of the dataset’s quality showing that \textbf{FAME-MT} is a reliable source of language register information, and we construct a publicly available proof-of-concept machine translation model that uses the dataset to steer the formality level of the translation. Currently, it is the largest dataset of formality annotations, with examples expressed in 112 European language pairs. The dataset is made available online. Dawid Wisniewski, Zofia Rostek, Artur Nowakowski |
EAMT (1) | 1 |
| 2023 | Fine-Grained and Complex Food Entity Recognition Benchmark for Ingredient SubstitutionabstractFood computing is currently fast-growing into an innovative area of knowledge extraction. However, benchmarks for information extraction from semi-structured data, especially when dealing with more complex relations, are scarce in this domain. In this paper, we introduce a benchmark aimed at information extraction of complex entities to support ingredient substitution tasks. Firstly, we present a new dataset – called TASTEset – for fine-grained recognition of food entities in culinary recipes. Secondly, we provide complex entity annotations for substitution on top of the fine-grained entity mentions, which we carefully prepared. We share the dataset and the tasks to encourage progress on more in-depth and complex information extraction from recipes. Agnieszka Lawrynowicz, Anna Wróblewska, Agnieszka Kaliska, Maciej Pawlowski, Dawid Wisniewski, Witold Sosnowski, Jakub Dutkiewicz |
K-CAP | 5 |
| 2022 | BigCQ: Generating a Synthetic Set of Competency Questions Formalized into SPARQL-OWL (Student Abstract)abstractWe present a method for constructing synthetic datasets of Competency Questions translated into SPARQL-OWL queries. This method is used to generate BigCQ, the largest set of CQ patterns and SPARQL-OWL templates that can provide translation examples to automate assessing the completeness and correctness of ontologies. Dawid Wisniewski, Jedrzej Potoniec, Agnieszka Lawrynowicz |
AAAI | 1 |
| 2022 | Sarcastic RoBERTa: A RoBERTa-Based Deep Neural Network Detecting Sarcasm on Twitter
Maciej Hercog, Piotr Jaronski, Jan Kolanowski, Pawel Mieczynski, Dawid Wisniewski, Jedrzej Potoniec |
DaWaK | 5 |
| 2022 | Should We Afford Affordances? Injecting ConceptNet Knowledge into BERT-Based Models to Improve Commonsense Reasoning AbilityabstractAbstract Recent years have shown that deep learning models pre-trained on large text corpora using the language model objective can help solve various tasks requiring natural language understanding. However, many commonsense concepts are underrepresented in online resources because they are too obvious for most humans. To solve this problem, we propose the use of affordances – common-sense knowledge that can be injected into models to increase their ability to understand our world. We show that injecting ConceptNet knowledge into BERT-based models leads to an increase in evaluation scores measured on the PIQA dataset. Andrzej Gretkowski, Dawid Wisniewski, Agnieszka Lawrynowicz |
EKAW | 2 |
| 2021 | SeeQuery: An Automatic Method for Recommending Translations of Ontology Competency Questions into SPARQL-OWLabstractOntology authoring is a complicated and error-prone process since the knowledge being modeled is expressed using logic-based formalisms, in which logical consequences of the knowledge have to be foreseen. To make that process easier, competency questions (CQs), being questions expressed in natural language are often stated to trace both the correctness and completeness of the ontology at a given time. However, CQs have to be translated into a formal language, like ontology query language (SPARQL-OWL), to query the ontology. Since the translation step is time-consuming and requires familiarity with the query language used, in this paper, we propose an automatic method named SeeQuery, which recommends SPARQL-OWL queries being translations of CQs stated against a given ontology. It consists of a pipeline of transformations based on template matching and filling, being motivated by the biggest to date publicly available CQ to SPARQL-OWL datasets. We provide a detailed description of SeeQuery and evaluate the method on a separate set of 2 ontologies with their CQs. It is, to date, the only automatic method available for recommending SPARQL-OWL queries out of CQs. The source code of SeeQuery is available at: https://github.com/dwisniewski/SeeQuery. Dawid Wisniewski, Jedrzej Potoniec, Agnieszka Lawrynowicz |
CIKM | 1 |
| 2021 | Incorporating Presuppositions of Competency Questions into Test-Driven Development of Ontologies (S)abstractOntology authoring is a complicated and error-prone process since the knowledge being modelled is expressed using logic-based formalisms, in which logical consequences of the knowledge have to be foreseen.Many approaches intended to make this task easier, use competency questions (CQs), being questions expressed in natural language to trace both the correctness and completeness of the ontology at a given time.However, CQs hold so-called presuppositions that have to be satisfied by the ontology to obtain meaningful answers from CQs.Moreover, CQs have to be expressed using a formal language, like ontology query language (SPARQL-OWL), to query the ontology.In this paper, we propose an extension of test-driven ontology development approach by formalization of presupposition satisfaction tests in terms of SPARQL-OWL queries, as well as providing translations of CQs into SPARQL-OWL queries if presupposition tests are passed.We provide a detailed description of the proposed framework and how to incorporate such tests in the workflow of test-driven development of ontologies.It is the first framework available for formalization of SPARQL-OWL queries out of CQs with their presupposition tests. Jedrzej Potoniec, Dawid Wisniewski, Agnieszka Lawrynowicz |
SEKE | 2 |
| 2020 | RecipeNLG: A Cooking Recipes Dataset for Semi-Structured Text GenerationabstractSemi-structured text generation is a non-trivial problem.Although last years have brought lots of improvements in natural language generation, thanks to the development of neural models trained on large scale datasets, these approaches still struggle with producing structured, context-and commonsense-aware texts.Moreover, it is not clear how to evaluate the quality of generated texts.To address these problems, we introduce RecipeNLG -a novel dataset of cooking recipes.We discuss the data collection process and the relation between the semi-structured texts and cooking recipes.We use the dataset to approach the problem of generating recipes.Finally, we make use of multiple metrics to evaluate the generated recipes. Michal Bien, Michal Gilski, Martyna Maciejewska, Wojciech Taisner, Dawid Wisniewski, Agnieszka Lawrynowicz |
INLG | 5 |
| 2019 | Analysis of Ontology Competency Questions and their formalizations in SPARQL-OWL
Dawid Wisniewski, Jedrzej Potoniec, Agnieszka Lawrynowicz, C. Maria Keet |
J. Web Semant. | 1 |