EDBT 2026 Demo / reviewers in the wild / expert
Mateusz Lango
dblp:180/0059
· DBLP profile ↗
21ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0003-2881-5642ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 5 first-author · 14 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reasoning Gets Harder for LLMs Inside A DialogueabstractLarge Language Models (LLMs) achieve strong performance on many reasoning benchmarks, yet these evaluations typically focus on isolated tasks that differ from real-world usage in task-oriented dialogue (TOD).In this setting, LLMs must perform reasoning inherently while generating text and adhering to instructions on role, format, and style.This mismatch raises concerns about whether benchmark performance accurately reflects models' reasoning robustness in TOD setting.We investigate how framing reasoning tasks within TOD affects LLM performance by introducing BOUL-DER, a new dynamic benchmark covering eight travel-related tasks that require arithmetic, spatial, and temporal reasoning with both commonsense and formal aspects.Each problem is presented in both isolated and dialogue-based variants, enabling controlled comparison while mitigating data contamination.Experiments on eight LLMs reveal a substantial and consistent performance gap between isolated and dialogue settings.Through ablations and qualitative analysis, we show that this gap is largely driven by the multi-turn nature of dialogue, with additional effects from role conditioning and tool-use requirements.Our results highlight the need to evaluate LLM reasoning in realistic interactive scenarios.1 Ivan Kartác, Mateusz Lango, Ondrej Dusek |
ACL (1) | 2 |
| 2025 | OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMsabstractLarge Language Models (LLMs) have demonstrated great potential as evaluators of NLG systems, allowing for high-quality, reference-free, and multi-aspect assessments. However, existing LLM-based metrics suffer from two major drawbacks: reliance on proprietary models to generate training data or perform evaluations, and a lack of fine-grained, explanatory feedback. We introduce OpeNLGauge, a fully open-source, reference-free NLG evaluation metric that provides accurate explanations based on individual error spans. OpeNLGauge is available as a two-stage ensemble of larger open-weight LLMs, or as a small fine-tuned evaluation model, with confirmed generalizability to unseen tasks, domains and aspects. Our extensive meta-evaluation shows that OpeNLGauge achieves competitive correlation with human judgments, outperforming state-of-the-art models on certain tasks while maintaining full reproducibility and providing explanations more than twice as accurate. Ivan Kartác, Mateusz Lango, Ondrej Dusek |
INLG | 2 |
| 2025 | Do My Eyes Deceive Me? A Survey of Human Evaluations of Hallucinations in NLGabstractHallucinations are one of the most pressing challenges for large language models (LLMs). While numerous methods have been proposed to detect and mitigate them automatically, human evaluation continues to serve as the gold standard. However, these human evaluations of hallucinations show substantial variation in definitions, terminology, and evaluation practices. In this paper, we survey 64 studies involving human evaluation of hallucination published between 2019 and 2024, to investigate how hallucinations are currently defined and assessed. Our analysis reveals a lack of consistency in definitions and exposes several concerning methodological shortcomings. Crucial details, such as evaluation guidelines, user interface design, inter-annotator agreement metrics, and annotator demographics, are frequently under-reported or omitted altogether. Patrícia Schmidtová, Eduardo Calò, Simone Balloccu, Dimitra Gkatzia, Rudali Huidrom, Mateusz Lango, Fahime Same, Vilém Zouhar, Saad Mahamood, Ondrej Dusek |
INLG | 6 |
| 2025 | How (un)faithful are explainable LLM-based NLG metrics?abstractExplainable NLG metrics are becoming a popular research topic; however, the faithfulness of the explanations they provide is typically not evaluated. In this work, we propose a testbed for assessing the faithfulness of span-based metrics by performing controlled perturbations of their explanations and observing changes in the final score. We show that several popular LLM evaluators do not consistently produce faithful explanations. Alex Terentowicz, Mateusz Lango, Ondrej Dusek |
INLG | 2 |
| 2025 | Counterfactual Explanations with Probabilistic Guarantees on their Robustness to Model Change
Ignacy Stepka, Jerzy Stefanowski, Mateusz Lango |
KDD (1) | 3 |
| 2025 | Boosting Dual Quality detection with AI-based social media analysis
Maksim Brzezinski, Maciej Niemir, Krzysztof Muszynski, Mateusz Lango, Dawid Wisniewski |
Inf. Process. Manag. | 4 |
| 2024 | Polish-ASTE: Aspect-Sentiment Triplet Extraction Datasets for PolishabstractAspect-Sentiment Triplet Extraction (ASTE) is one of the most challenging and complex tasks in sentiment analysis. It concerns the construction of triplets that contain an aspect, its associated sentiment polarity, and an opinion phrase that serves as a rationale for the assigned polarity. Despite the growing popularity of the task and the many machine learning methods being proposed to address it, the number of datasets for ASTE is very limited. In particular, no dataset is available for any of the Slavic languages. In this paper, we present two new datasets for ASTE containing customer opinions about hotels and purchased products expressed in Polish. We also perform experiments with two ASTE techniques combined with two large language models for Polish to investigate their performance and the difficulty of the assembled datasets. The new datasets are available under a permissive licence and have the same file format as the English datasets, facilitating their use in future research. Marta Lango, Borys Naglik, Mateusz Lango, Iwo Naglik |
LREC/COLING | 3 |
| 2024 | Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMsabstractSimone Balloccu, Patrícia Schmidtová, Mateusz Lango, Ondrej Dusek. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Simone Balloccu, Patrícia Schmidtová, Mateusz Lango, Ondrej Dusek |
EACL (1) | 3 |
| 2024 | Leveraging Large Language Models for Building Interpretable Rule-Based Data-to-Text SystemsabstractWe introduce a simple approach that uses a large language model (LLM) to automatically implement a fully interpretable rule-based data-to-text system in pure Python.Experimental evaluation on the WebNLG dataset showed that such a constructed system produces text of better quality (according to the BLEU and BLEURT metrics) than the same LLM prompted to directly produce outputs, and produces fewer hallucinations than a BART language model fine-tuned on the same data.Furthermore, at runtime, the approach generates text in a fraction of the processing time required by neural approaches, using only a single CPU. Jedrzej Warczynski, Mateusz Lango, Ondrej Dusek |
INLG | 2 |
| 2024 | GMMSampling: a new model-based, data difficulty-driven resampling method for multi-class imbalanced dataabstractAbstract Learning from multi-class imbalanced data has still received limited research attention. Most of the proposed methods focus on the global class imbalance ratio only. In contrast, experimental studies demonstrated that the imbalance ratio itself is not the main difficulty in the imbalanced learning. It is the combination of the imbalance ratio with other data difficulty factors, such as class overlapping or minority class decomposition into various subconcepts, that significantly affects the classification performance. This paper presents GMMSampling—a new resampling method that exploits information about data difficulty factors to clear class overlapping regions from majority class instances and to simultaneously oversample each subconcept of the minority class. The experimental evaluation demonstrated that the proposed method achieves better results in terms of G-mean, balanced accuracy, macro-AP, MCC and F-score than other related methods. Iwo Naglik, Mateusz Lango |
Mach. Learn. | 2 |
| 2023 | The Problem of Coherence in Natural Language Explanations of RecommendationsabstractProviding natural language explanations for recommendations is particularly useful from the perspective of a non-expert user. Although several methods for providing such explanations have recently been proposed, we argue that an important aspect of explanation quality has been overlooked in their experimental evaluation. Specifically, the coherence between generated text and predicted rating, which is a necessary condition for an explanation to be useful, is not properly captured by currently used evaluation measures. In this paper, we highlight the issue of explanation and prediction coherence by 1) presenting results from a manual verification of explanations generated by one of the state-of-the-art approaches 2) proposing a method of automatic coherence evaluation 3) introducing a new transformer-based method that aims to produce more coherent explanations than the state-of-the-art approaches 4) performing an experimental evaluation which demonstrates that this method significantly improves the explanation coherence without affecting the other aspects of recommendation performance. Jakub Raczynski, Mateusz Lango, Jerzy Stefanowski |
ECAI | 2 |
| 2023 | Critic-Driven Decoding for Mitigating Hallucinations in Data-to-text GenerationabstractHallucination of text ungrounded in the input is a well-known problem in neural data-to-text generation.Many methods have been proposed to mitigate it, but they typically require altering model architecture or collecting additional data, and thus cannot be easily applied to an existing model.In this paper, we explore a new way to mitigate hallucinations by combining the probabilistic output of a generator language model (LM) with the output of a special "text critic" classifier, which guides the generation by assessing the match between the input data and the text generated so far.Our method does not need any changes to the underlying LM's architecture or training procedure and can thus be combined with any model and decoding operating on word probabilities.The critic does not need any additional training data, using the base LM's training data and synthetic negative examples.Our experimental results show that our method improves over the baseline on the WebNLG and OpenDialKG benchmarks. Mateusz Lango, Ondrej Dusek |
EMNLP | 1 |
| 2023 | Generating clickbait spoilers with an ensemble of large language modelsabstractClickbait posts are a widespread problem in the webspace.The generation of spoilers, i.e. short texts that neutralize clickbait by providing information that satisfies the curiosity induced by it, is one of the proposed solutions to the problem.Current state-of-the-art methods are based on passage retrieval or question answering approaches and are limited to generating spoilers only in the form of a phrase or a passage.In this work, we propose an ensemble of fine-tuned large language models for clickbait spoiler generation.Our approach is not limited to phrase or passage spoilers, but is also able to generate multipart spoilers that refer to several non-consecutive parts of text.Experimental evaluation demonstrates that the proposed ensemble model outperforms the baselines in terms of BLEU, ME-TEOR and BERTScore metrics. Mateusz Wozny, Mateusz Lango |
INLG | 2 |
| 2023 | Exploiting Phrase Interrelations in Span-level Neural Approaches for Aspect Sentiment Triplet Extraction
Iwo Naglik, Mateusz Lango |
PAKDD (4) | 2 |
| 2022 | What makes multi-class imbalanced problems difficult? An experimental study
Mateusz Lango, Jerzy Stefanowski |
Expert Syst. Appl. | 1 |
| 2021 | Time Aspect in Making an Actionable Prediction of a Conversation Breakdown
Piotr Janiszewski, Mateusz Lango, Jerzy Stefanowski |
ECML/PKDD (5) | 2 |
| 2020 | A Closer Look on Unsupervised Cross-lingual Word Embeddings MappingabstractIn this work, we study the unsupervised cross-lingual word embeddings mapping method presented by Artetxe et al. (2018). First, wesuccessfully reproduced the experiments performed in the original work, finding only minor differences. Furthermore, we verified themethod’s robustness on different embedding representations and new language pairs, particularly these involving Slavic languages likePolish or Czech. We also performed an experimental analysis of the impact of the method’s parameters on the final result. Finally, welooked for an alternative way of initialization, which directly relies on the isometric assumption. Our work confirms the results presentedearlier, at the same time pointing at interesting problems occurring while using the method with different types of embeddings or onless-common language pairs. Kamil Plucinski, Mateusz Lango, Michal Zimniewicz |
LREC | 2 |
| 2018 | Semi-Automatic Construction of Word-Formation Networks (for Polish and Spanish)
Mateusz Lango, Magda Sevcíková, Zdenek Zabokrtský |
LREC | 1 |
| 2018 | Multi-class and feature selection extensions of Roughly Balanced Bagging for imbalanced dataabstractRoughly Balanced Bagging is one of the most efficient ensembles specialized for class imbalanced data. In this paper, we study its basic properties that may influence its good classification performance. We experimentally analyze them with respect to bootstrap construction, deciding on the number of component classifiers, their diversity, and ability to deal with the most difficult types of the minority examples. Then, we introduce two generalizations of this ensemble for dealing with a higher number of attributes and for adapting it to handle multiple minority classes. Experiments with synthetic and real life data confirm usefulness of both proposals. Mateusz Lango, Jerzy Stefanowski |
J. Intell. Inf. Syst. | 1 |
| 2017 | Discovering Minority Sub-clusters and Local Difficulty Factors from Imbalanced Data
Mateusz Lango, Dariusz Brzezinski, Sebastian Firlik, Jerzy Stefanowski |
DS | 1 |
| 2017 | Evaluating Difficulty of Multi-class Imbalanced Data
Mateusz Lango, Krystyna Napierala, Jerzy Stefanowski |
ISMIS | 1 |