VLDB 2026 Research / reviewers in the wild / expert
Vilém Zouhar
dblp:254/1832
· DBLP profile ↗
24ranked-venue papers
8as first author
23since 2021 · last 2026
0000-0001-9874-2069ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 8 first-author · 21 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Augmenting Text to Increase Translation DifficultyabstractAs state-of-the-art machine translation models saturate standard benchmarks, the field needs more challenging evaluations to distinguish between models of varying quality. We propose augmenting existing benchmarks to increase translation difficulty by combining adversarial optimization with a differentiable translation difficulty estimator. Our Adversarial Translation Optimization (ATO) uses gradients from a combined difficulty and fluency objective to iteratively replace tokens. Because each step branches over candidate substitutions at every position, optimization becomes a tree search problem, which we address with Beam Search. ATO offers a gradient-based alternative to LLM-based dataset creation without LLM prompting, expensive human curation, or task-specific model training. Our ATO-modified benchmark lowers average translation quality (xCOMET) from 0.93 to 0.82, compared to 0.88 for paraphrasing and 0.86 for a zero-shot baseline. Through human evaluation we confirm that the modified texts remain reasonably natural while being substantially harder to translate. We release two datasets of 200 English texts each, generated by our methods, as well as the code. William Kalikman, Simon Sukup, Michal Tesnar, Vilém Zouhar |
EAMT (1) | 4 |
| 2026 | Understanding Large Language Model Behaviors Through Interactive Counterfactual Generation and AnalysisabstractUnderstanding the behavior of large language models (LLMs) is crucial for ensuring their safe and reliable use. However, existing explainable AI (XAI) methods for LLMs primarily rely on word-level explanations, which are often computationally inefficient and misaligned with human reasoning processes. Moreover, these methods often treat explanation as a one-time output, overlooking its inherently interactive and iterative nature. In this paper, we present LLM Analyzer, an interactive visualization system that addresses these limitations by enabling intuitive and efficient exploration of LLM behaviors through counterfactual analysis. Our system features a novel algorithm that generates fluent and semantically meaningful counterfactuals via targeted removal and replacement operations at user-defined levels of granularity. These counterfactuals are used to compute feature attribution scores, which are then integrated with concrete examples in a table-based visualization, supporting dynamic analysis of model behavior. A user study with LLM practitioners and interviews with experts demonstrate the system's usability and effectiveness, emphasizing the importance of involving humans in the explanation process as active participants rather than passive recipients. Furui Cheng, Vilém Zouhar, Robin Shing Moon Chan, Daniel Fürst, Hendrik Strobelt, Mennatallah El-Assady |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | Biased Tales: Cultural and Topic Bias in Generating Children's StoriesabstractStories play a pivotal role in human communication, shaping beliefs and morals, particularly in children.As parents increasingly rely on large language models (LLMs) to craft bedtime stories, the presence of cultural and gender stereotypes in these narratives raises significant concerns.To address this issue, we present Biased Tales, a comprehensive dataset designed to analyze how biases influence protagonists' attributes and story elements in LLM-generated stories.Our analysis uncovers striking disparities.When the protagonist is described as a girl (as compared to a boy), appearance-related attributes increase by 55.26%.Stories featuring non-Western children disproportionately emphasize cultural heritage, tradition, and family themes far more than those for Western children.Our findings highlight the role of sociocultural bias in making creative AI use more equitable and diverse. Donya Rooein, Vilém Zouhar, Debora Nozza, Dirk Hovy |
EMNLP | 2 |
| 2025 | Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreementabstractWord-level quality estimation (WQE) aims to automatically identify fine-grained error spans in machine-translated outputs and has found many uses, including assisting translators during post-editing.Modern WQE techniques are often expensive, involving prompting of large language models or ad-hoc training on large amounts of human-labeled data.In this work, we investigate efficient alternatives exploiting recent advances in language model interpretability and uncertainty quantification to identify translation errors from the inner workings of translation models.In our evaluation spanning 14 metrics across 12 translation directions, we quantify the impact of human label variation on metric performance by using multiple sets of human labels.Our results highlight the untapped potential of unsupervised metrics, the shortcomings of supervised methods when faced with label uncertainty, and the brittleness of single-annotator evaluation practices. Gabriele Sarti, Vilém Zouhar, Malvina Nissim, Arianna Bisazza |
EMNLP | 2 |
| 2025 | Do My Eyes Deceive Me? A Survey of Human Evaluations of Hallucinations in NLGabstractHallucinations are one of the most pressing challenges for large language models (LLMs). While numerous methods have been proposed to detect and mitigate them automatically, human evaluation continues to serve as the gold standard. However, these human evaluations of hallucinations show substantial variation in definitions, terminology, and evaluation practices. In this paper, we survey 64 studies involving human evaluation of hallucination published between 2019 and 2024, to investigate how hallucinations are currently defined and assessed. Our analysis reveals a lack of consistency in definitions and exposes several concerning methodological shortcomings. Crucial details, such as evaluation guidelines, user interface design, inter-annotator agreement metrics, and annotator demographics, are frequently under-reported or omitted altogether. Patrícia Schmidtová, Eduardo Calò, Simone Balloccu, Dimitra Gkatzia, Rudali Huidrom, Mateusz Lango, Fahime Same, Vilém Zouhar, Saad Mahamood, Ondrej Dusek |
INLG | 8 |
| 2025 | A Bayesian Optimization Approach to Machine Translation RerankingabstractJulius Cheng, Maike Züfle, Vilém Zouhar, Andreas Vlachos. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Julius Cheng, Maike Züfle, Vilém Zouhar, Andreas Vlachos 0001 |
NAACL (Long Papers) | 3 |
| 2025 | AI-Assisted Human Evaluation of Machine TranslationabstractVilém Zouhar, Tom Kocmi, Mrinmaya Sachan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Vilém Zouhar, Tom Kocmi, Mrinmaya Sachan |
NAACL (Long Papers) | 1 |
| 2025 | QE4PE: Word-level Quality Estimation for Human Post-Editing
Gabriele Sarti, Vilém Zouhar, Grzegorz Chrupala, Ana Guerberof Arenas, Malvina Nissim, Arianna Bisazza |
Trans. Assoc. Comput. Linguistics | 2 |
| 2025 | How to Select Datapoints for Efficient Human Evaluation of NLG Models?abstractAbstract Human evaluation is the gold standard for evaluating text generation models. However, it is expensive. In order to fit budgetary constraints, a random subset of the test data is often chosen in practice for human evaluation. However, randomly selected data may not accurately represent test performance, making this approach economically inefficient for model comparison. Thus, in this work, we develop and analyze a suite of selectors to get the most informative datapoints for human evaluation, taking the evaluation costs into account. We show that selectors based on variance in automated metric scores, diversity in model outputs, or Item Response Theory outperform random selection. We further develop an approach to distill these selectors to the scenario where the model outputs are not yet available. In particular, we introduce source-based estimators, which predict item usefulness for human evaluation just based on the source texts. We demonstrate the efficacy of our selectors in two common NLG tasks, machine translation and summarization, and show that only ∼70% of the test data is needed to produce the same evaluation result as the entire data. Vilém Zouhar, Peng Cui 0006, Mrinmaya Sachan |
Trans. Assoc. Comput. Linguistics | 1 |
| 2024 | How to Engage your Readers? Generating Guiding Questions to Promote Active ReadingabstractUsing questions in written text is an effective strategy to enhance readability.However, what makes an active reading question good, what the linguistic role of these questions is, and what is their impact on human reading remains understudied.We introduce GUIDINGQ, a dataset of 10K in-text questions from textbooks and scientific articles.By analyzing the dataset, we present a comprehensive understanding of the use, distribution, and linguistic characteristics of these questions.Then, we explore various approaches to generate such questions using language models.Our results highlight the importance of capturing inter-question relationships and the challenge of question position identification in generating these questions.Finally, we conduct a human study to understand the implication of such questions on reading comprehension.We find that the generated questions are of high quality and are almost as effective as human-written questions in terms of improving readers' memorization and comprehension.github.com/eth-lre/engage-your-readers Questions in titles:How do Philosophers arrive at truth?Is there no quantum form of Einstein Gravity?Why do house-hunting ants recruit in both directions? Peng Cui 0006, Vilém Zouhar, Xiaoyu Zhang 0014, Mrinmaya Sachan |
ACL (1) | 2 |
| 2024 | Navigating the Metrics Maze: Reconciling Score Magnitudes and AccuraciesabstractTen years ago, a single metric, BLEU, governed progress in machine translation research.For better or worse, there is no such consensus today, and consequently it is difficult for researchers to develop and retain intuitions about metric deltas that drove earlier research and deployment decisions.This paper investigates the "dynamic range" of a number of modern metrics in an effort to provide a collective understanding of the meaning of differences in scores both within and among metrics; in other words, we ask what point difference x in metric y is required between two systems for humans to notice?We conduct our evaluation on a new large dataset, ToShip23, using it to discover deltas at which metrics achieve system-level differences that are meaningful to humans, which we measure by pairwise system accuracy.We additionally show that this method of establishing delta-accuracy is more stable than the standard use of statistical p-values in regards to testset size.Where data size permits, we also explore the effect of metric deltas and accuracy across finer-grained features such as translation direction, domain, and system closeness. Tom Kocmi, Vilém Zouhar, Christian Federmann, Matt Post |
ACL (1) | 2 |
| 2024 | RELIC: Investigating Large Language Model Responses using Self-ConsistencyabstractLarge Language Models (LLMs) are notorious for blending fact with fiction and generating non-factual content, known as hallucinations. To address this challenge, we propose an interactive system that helps users gain insight into the reliability of the generated text. Our approach is based on the idea that the self-consistency of multiple samples generated by the same LLM relates to its confidence in individual claims in the generated texts. Using this idea, we design RELIC, an interactive system that enables users to investigate and verify semantic-level variations in multiple long-form responses. This allows users to recognize potentially inaccurate information in the generated text and make necessary corrections. From a user study with ten participants, we demonstrate that our approach helps users better verify the reliability of the generated text. We further summarize the design implications and lessons learned from this research for future studies of reliable human-LLM interactions. Furui Cheng, Vilém Zouhar, Simran Arora, Mrinmaya Sachan, Hendrik Strobelt, Mennatallah El-Assady |
CHI | 2 |
| 2024 | Two Counterexamples to Tokenization and the Noiseless ChannelabstractIn Tokenization and the Noiseless Channel (Zouhar et al., 2023), Rényi efficiency is suggested as an intrinsic mechanism for evaluating a tokenizer: for NLP tasks, the tokenizer which leads to the highest Rényi efficiency of the unigram distribution should be chosen. The Rényi efficiency is thus treated as a predictor of downstream performance (e.g., predicting BLEU for a machine translation task), without the expensive step of training multiple models with different tokenizers. Although useful, the predictive power of this metric is not perfect, and the authors note there are additional qualities of a good tokenization scheme that Rényi efficiency alone cannot capture. We describe two variants of BPE tokenization which can arbitrarily increase Rényi efficiency while decreasing the downstream model performance. These counterexamples expose cases where Rényi efficiency fails as an intrinsic tokenization metric and thus give insight for building more accurate predictors. Marco Cognetta, Vilém Zouhar, Sangwhan Moon, Naoaki Okazaki |
LREC/COLING | 2 |
| 2024 | PWESuite: Phonetic Word Embeddings and Tasks They FacilitateabstractMapping words into a fixed-dimensional vector space is the backbone of modern NLP. While most word embedding methods successfully encode semantic information, they overlook phonetic information that is crucial for many tasks. We develop three methods that use articulatory features to build phonetically informed word embeddings. To address the inconsistent evaluation of existing phonetic word embedding methods, we also contribute a task suite to fairly evaluate past, current, and future methods. We evaluate both (1) intrinsic aspects of phonetic word embeddings, such as word retrieval and correlation with sound similarity, and (2) extrinsic performance on tasks such as rhyme and cognate detection and sound analogies. We hope our task suite will promote reproducibility and inspire future phonetic embedding research. Vilém Zouhar, Kalvin Chang, Chenxuan Cui, Nate B. Carlson, Nathaniel R. Robinson, Mrinmaya Sachan, David R. Mortensen |
LREC/COLING | 1 |
| 2024 | Distributional Properties of Subword RegularizationabstractSubword regularization, used widely in NLP, improves model performance by reducing the dependency on exact tokenizations, augmenting the training corpus, and exposing the model to more unique contexts during training.BPE and MaxMatch, two popular subword tokenization schemes, have stochastic dropout regularization variants.However, there has not been an analysis of the distributions formed by them.We show that these stochastic variants are heavily biased towards a small set of tokenizations per word.If the benefits of subword regularization are as mentioned, we hypothesize that biasedness artificially limits the effectiveness of these schemes.Thus, we propose an algorithm to uniformly sample tokenizations that we use as a drop-in replacement for the stochastic aspects of existing tokenizers, and find that it improves machine translation quality. Marco Cognetta, Vilém Zouhar, Naoaki Okazaki |
EMNLP | 2 |
| 2024 | AutoTutor meets Large Language Models: A Language Model Tutor with Rich Pedagogy and GuardrailsabstractLarge Language Models (LLMs) have found several use cases in education, ranging from automatic question generation to essay evaluation. In this paper, we explore the potential of using LLMs to author Intelligent Tutoring Systems. A common pitfall of using LLMs as tutors is their straying from desired pedagogical strategies such as leaking the answer to the student, and in general, providing no guarantees on the validity or appropriateness of the tutor assistance. We argue that while LLMs with certain guardrails can take the place of subject experts, the overall pedagogical design still needs to be handcrafted for the best learning results. Based on this principle, we create a sample end-to-end tutoring system named MWPTutor, which uses LLMs to fill in the state space of a predefined finite state transducer. This approach retains the structure and the pedagogy of traditional tutoring systems that has been developed over the years by learning scientists but brings in additional flexibility of LLM-based approaches. Through a human evaluation study on two datasets with math word problems, we show that our hybrid approach achieves a better overall tutoring score than an instructed, but otherwise free-form, GPT-4. MWPTutor is completely modular and opens up the scope for the community to improve its performance by refining its individual modules or using different teaching strategies that it can follow. Sankalan Pal Chowdhury, Vilém Zouhar, Mrinmaya Sachan |
L@S | 2 |
| 2023 | Tokenization and the Noiseless ChannelabstractVilém Zouhar, Clara Meister, Juan Gastaldi, Li Du, Mrinmaya Sachan, Ryan Cotterell. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Vilém Zouhar, Clara Meister, Juan Luis Gastaldi, Mrinmaya Sachan, Ryan Cotterell |
ACL (1) | 1 |
| 2023 | Poor Man's Quality Estimation: Predicting Reference-Based MT Metrics Without the ReferenceabstractVilém Zouhar, Shehzaad Dhuliawala, Wangchunshu Zhou, Nico Daheim, Tom Kocmi, Yuchen Eleanor Jiang, Mrinmaya Sachan. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Vilém Zouhar, Shehzaad Dhuliawala, Wangchunshu Zhou, Nico Daheim, Tom Kocmi, Yuchen Eleanor Jiang, Mrinmaya Sachan |
EACL | 1 |
| 2023 | A Diachronic Perspective on User Trust in AI under UncertaintyabstractIn a human-AI collaboration, users build a mental model of the AI system based on its reliability and how it presents its decision, e.g. its presentation of system confidence and an explanation of the output.Modern NLP systems are often uncalibrated, resulting in confidently incorrect predictions that undermine user trust.In order to build trustworthy AI, we must understand how user trust is developed and how it can be regained after potential trust-eroding events.We study the evolution of user trust in response to these trust-eroding events using a betting game.We find that even a few incorrect instances with inaccurate confidence estimates damage user trust and performance, with very slow recovery.We also show that this degradation in trust reduces the success of human-AI collaboration and that different types of miscalibration-unconfidently correct and confidently incorrect-have different negative effects on user trust.Our findings highlight the importance of calibration in user-facing AI applications and shed light on what aspects help users decide whether to trust the AI system. Shehzaad Dhuliawala, Vilém Zouhar, Mennatallah El-Assady, Mrinmaya Sachan |
EMNLP | 2 |
| 2023 | Enhancing Textbooks with Visuals from the Web for Improved LearningabstractTextbooks are one of the main mediums for delivering high-quality education to students.In particular, explanatory and illustrative visuals play a key role in retention, comprehension and general transfer of knowledge.However, many textbooks lack these interesting visuals to support student learning.In this paper, we investigate the effectiveness of vision-language models to automatically enhance textbooks with images from the web.We collect a dataset of e-textbooks in the math, science, social science and business domains.We then set up a text-image matching task that involves retrieving and appropriately assigning web images to textbooks, which we frame as a matching optimization problem.Through a crowd-sourced evaluation, we verify that (1) while the original textbook images are rated higher, automatically assigned ones are not far behind, and (2) the precise formulation of the optimization problem matters.We release the dataset of textbooks with an associated image bank to inspire further research in this intersectional area of computer vision and NLP for education. Janvijay Singh, Vilém Zouhar, Mrinmaya Sachan |
EMNLP | 2 |
| 2023 | Revisiting Automated Topic Model Evaluation with Large Language ModelsabstractTopic models help make sense of large text collections.Automatically evaluating their output and determining the optimal number of topics are both longstanding challenges, with no effective automated solutions to date.This paper evaluates the effectiveness of large language models (LLMs) for these tasks.We find that LLMs appropriately assess the resulting topics, correlating more strongly with human judgments than existing automated metrics.However, the type of evaluation task matters -LLMs correlate better with coherence ratings of word sets than on a word intrusion task.We find that LLMs can also guide users toward a reasonable number of topics.In actual applications, topic models are typically used to answer a research question related to a collection of texts.We can incorporate this research question in the prompt to the LLM, which helps estimate the optimal number of topics.github.com/dominiksinsaarland/ evaluating-topic-model-output Dominik Stammbach, Vilém Zouhar, Alexander Miserlis Hoyle, Mrinmaya Sachan, Elliott Ash |
EMNLP | 2 |
| 2021 | Neural Machine Translation Quality and Post-Editing PerformanceabstractWe test the natural expectation that using MT in professional translation saves human processing time.The last such study was carried out by Sanchez-Torron and Koehn (2016) with phrase-based MT, artificially reducing the translation quality.In contrast, we focus on neural MT (NMT) of high quality, which has become the state-of-the-art approach since then and also got adopted by most translation companies.Through an experimental study involving over 30 professional translators for English→Czech translation, we examine the relationship between NMT performance and post-editing time and quality.Across all models, we found that better MT systems indeed lead to fewer changes in the sentences in this industry setting.The relation between system quality and post-editing time is however not straightforward and, contrary to the results on phrase-based MT, BLEU is definitely not a stable predictor of the time or final output quality. Vilém Zouhar, Martin Popel, Ondrej Bojar, Ales Tamchyna |
EMNLP (1) | 1 |
| 2021 | Backtranslation Feedback Improves User Confidence in MT, Not QualityabstractVilém Zouhar, Michal Novák, Matúš Žilinec, Ondřej Bojar, Mateo Obregón, Robin L. Hill, Frédéric Blain, Marina Fomicheva, Lucia Specia, Lisa Yankovskaya. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Vilém Zouhar, Michal Novák 0001, Matús Zilinec, Ondrej Bojar, Mateo Obregón, Robin L. Hill, Frédéric Blain, Marina Fomicheva, Lucia Specia, Lisa Yankovskaya |
NAACL-HLT | 1 |
| 2020 | Outbound Translation User Interface Ptakopet: A Pilot Study
Vilém Zouhar, Ondrej Bojar |
LREC | 1 |