Michal Stefánik

dblp:255/9301 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0003-1766-5538ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Theory of computation · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Language Models Learn Universal Representations of Numbers and Here's Why You Should Care
abstract
Michal Štefánik, Timothee Mickus, Marek Kadlčík, Bertram Højer, Michal Spiegel, Raúl Vázquez, Aman Sinha, Josef Kuchař, Philipp Mondorf, Pontus Stenetorp. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Michal Stefánik, Timothee Mickus, Marek Kadlcík, Bertram Højer, Michal Spiegel, Raúl Vázquez, Aman Sinha 0002, Josef Kuchar, Philipp Mondorf, Pontus Stenetorp
ACL (1)1
2026 VectorEdits: A Dataset and Benchmark for Instruction-Based Editing of Vector Graphics
abstract
We introduce a large-scale dataset for instruction-guided vector image editing, consisting of over 270,000 pairs of SVG images paired with natural language edit instructions. Our dataset enables training and evaluation of models that modify vector graphics based on textual commands. We describe the data collection process, including image pairing via CLIP similarity and instruction generation with vision-language models. Initial experiments with state-of-the-art large language models reveal that current methods struggle to produce accurate and valid edits, underscoring the challenge of this task. To foster research in natural language-driven vector graphic generation and editing, we make our resources created within this work publicly available.
Josef Kuchar, Marek Kadlcík, Michal Spiegel, Michal Stefánik
LREC4
2025 Pre-trained Language Models Learn Remarkably Accurate Representations of Numbers
abstract
Pretrained language models (LMs) are prone to arithmetic errors.Existing work showed limited success in probing numeric values from models' representations, indicating that these errors can be attributed to the inherent unreliability of distributionally learned embeddings in representing exact quantities.However, we observe that previous probing methods are inadequate for the emergent structure of learned number embeddings with sinusoidal patterns.In response, we propose a novel probing technique that decodes numeric values from input embeddings with near-perfect accuracy across a range of open-source LMs.This proves that after the sole pre-training, LMs represent numbers with remarkable precision.Finally, we find that the embeddings' precision, judged by our probe's accuracy, explains a large portion of LM's errors in elementary arithmetic, and show that aligning the embeddings with the pattern our probes discover can mitigate these errors.
Marek Kadlcík, Michal Stefánik, Timothee Mickus, Josef Kuchar, Michal Spiegel
EMNLP2
2025 BenCzechMark : A Czech-Centric Multitask and Multimetric Benchmark for Large Language Models with Duel Scoring Mechanism
abstract
Abstract We present BenCzechMark (BCM), the first comprehensive Czech language benchmark designed for large language models, offering diverse tasks, multiple task formats, and multiple evaluation metrics. Its duel scoring system is grounded in statistical significance theory and uses aggregation across tasks inspired by social preference theory. Our benchmark encompasses 50 challenging tasks, with corresponding test datasets, primarily in native Czech, with 14 newly collected ones. These tasks span 8 categories and cover diverse domains, including historical Czech news, essays from pupils or language learners, and spoken word. Furthermore, we collect and clean BUT-Large Czech Collection, the largest publicly available clean Czech language corpus, and use it for (i) contamination analysis and (ii) continuous pretraining of the first Czech-centric 7B language model with Czech-specific tokenization. We use our model as a baseline for comparison with publicly available multilingual models. Lastly, we release and maintain a leaderboard with existing 50 model submissions, where new model submissions can be made at https://huggingface.co/spaces/CZLC/BenCzechMark.
Martin Fajcik, Martin Docekal, Jan Dolezal, Karel Ondrej, Karel Benes, Jan Kapsa, Pavel Smrz, Alexander Polok, Michal Hradis, Zuzana Neverilová, Ales Horák, Radoslav Sabol, Michal Stefánik, Adam Jirkovsky, David Adamczyk, Petr Hyner, Jan Hula, Hynek Kydlícek
Trans. Assoc. Comput. Linguistics13
2024 Think Twice: Measuring the Efficiency of Eliminating Prediction Shortcuts of Question Answering Models
abstract
While the Large Language Models (LLMs) dominate a majority of language understanding tasks, previous work shows that some of these results are supported by modelling spurious correlations of training datasets.Authors commonly assess model robustness by evaluating their models on out-of-distribution (OOD) datasets of the same task, but these datasets might share the bias of the training dataset.We propose a simple method for measuring a scale of models' reliance on any identified spurious feature and assess the robustness towards a large set of known and newly found prediction biases for various pre-trained models and debiasing methods in Question Answering (QA).We find that the while existing debiasing methods can mitigate reliance on a chosen spurious feature, the OOD performance gains of these methods can not be explained by mitigated reliance on biased features, suggesting that biases are shared among different QA datasets.Finally, we evidence this to be the case by measuring that performance of models trained on different QA datasets rely on bias features comparably to the ID model.We hope these results will motivate future work to refine the reports of LMs' robustness to a level of adversarial samples addressing specific spurious features.
Lukás Mikula, Michal Stefánik, Marek Petrovic, Petr Sojka
EACL (1)2
2023 Soft Alignment Objectives for Robust Adaptation of Language Generation
abstract
Domain adaptation allows generative language models to address specific flaws caused by the domain shift of their application.However, the traditional adaptation by further training on indomain data rapidly weakens the model's ability to generalize to other domains, making the openended deployments of the adapted models prone to errors.This work introduces novel training objectives built upon a semantic similarity of the predicted tokens to the reference.Our results show that (1) avoiding the common assumption of a single correct prediction by constructing the training target from tokens' semantic similarity can largely mitigate catastrophic forgetting of adaptation, while (2) preserving the adaptation in-domain quality, (3) with negligible additions to compute costs.In the broader context, the objectives grounded in a continuous token similarity pioneer the exploration of the middle ground between the efficient but naïve exact-match token-level objectives and expressive but computationally-and resourceintensive sequential objectives.
Michal Stefánik, Marek Kadlcík, Petr Sojka
ACL (1)1
2023 Calc-X and Calcformers: Empowering Arithmetical Chain-of-Thought through Interaction with Symbolic Systems
abstract
Despite outstanding performance in many tasks, language models are notoriously inclined to make factual errors in tasks requiring arithmetic computation.We address this deficiency by creating Calc-X, a collection of datasets that demonstrates the appropriate use of a calculator in reasoning chains.Calc-X is suitable for teaching language models to offload computations to a symbolic system.We survey and unify several existing chain-of-thought datasets into a proposed format, resulting in a standard collection of over 300,000 samples requiring arithmetic reasoning.Finally, we use the new Calc-X collection to train open-source calculator-using models we call Calcformers and show that these models approximately double the accuracy of generating correct results compared to vanilla language model baselines.We make all Calc-X datasets, source code and Calcformers models publicly available.1
Marek Kadlcík, Michal Stefánik, Ondrej Sotolár, Vlastimil Martinek
EMNLP2
2021 CICM'21 Systems Entries
Martin Líska, Dávid Lupták, Vit Novotny, Michal Ruzicka, Boris Shminke, Petr Sojka, Michal Stefánik, Markus Wenzel 0001
CICM7
2021 WebMIaS on Docker - Deploying Math-Aware Search in a Single Line of Code
Dávid Lupták, Vit Novotny, Michal Stefánik, Petr Sojka
CICM3