EDBT 2026 Demo / reviewers in the wild / expert
Mario Giulianelli
dblp:205/2569
· DBLP profile ↗
24ranked-venue papers
8as first author
22since 2021 · last 2026
0009-0004-1281-9686ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 8 first-author · 22 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Probing for Reading TimesabstractEleftheria Tsipidi, Samuel Kiegeland, Francesco Ignazio Re, Tianyang Xu, Mario Giulianelli, Karolina Stanczak, Ryan Cotterell. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Eleftheria Tsipidi, Samuel Kiegeland, Francesco Ignazio Re, Tianyang Xu 0002, Mario Giulianelli, Karolina Stanczak, Ryan Cotterell |
ACL (1) | 5 |
| 2026 | Surprisal Minimisation over Goal-directed Alternatives Predicts Production Choice in DialogueabstractWe model utterance production as probabilistic cost-sensitive choice over contextual alternatives, using information-theoretic notions of cost.We distinguish between goal-directed alternatives that realise a fixed communicative intent and goal-agnostic alternatives defined only by contextual plausibility, allowing us to derive speaker-and listener-oriented interpretations of different cost measures.We present a procedure to generate both types of alternative sets using language models.Analysing production choices in open-ended dialogue under both deterministic and probabilistic cost minimisation, we find that surprisal minimisation relative to goal-directed alternatives provides the strongest predictive account under both analyses.By contrast, uniform information density and length-based costs exhibit weaker and less consistent predictive power across conditions.More broadly, our study suggests that alternative-conditioned optimisation with LMgenerated alternatives provides a principled framework for studying speaker and listener pressures in naturalistic language production. 1 Thomas P. Utting, Mario Giulianelli, Arabella Sinclair |
ACL (1) | 2 |
| 2025 | A Spatio-Temporal Point Process for Fine-Grained Modeling of Reading BehaviorabstractFrancesco Ignazio Re, Andreas Opedal, Glib Manaiev, Mario Giulianelli, Ryan Cotterell. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Francesco Ignazio Re, Andreas Opedal, Glib Manaiev, Mario Giulianelli, Ryan Cotterell |
ACL (1) | 4 |
| 2025 | Information Locality as an Inductive Bias for Neural Language ModelsabstractTaiga Someya, Anej Svete, Brian DuSell, Timothy J. O’Donnell, Mario Giulianelli, Ryan Cotterell. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Taiga Someya, Anej Svete, Brian DuSell, Timothy J. O'Donnell, Mario Giulianelli, Ryan Cotterell |
ACL (1) | 5 |
| 2025 | The Harmonic Structure of Information ContoursabstractEleftheria Tsipidi, Samuel Kiegeland, Franz Nowak, Tianyang Xu, Ethan Wilcox, Alex Warstadt, Ryan Cotterell, Mario Giulianelli. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Eleftheria Tsipidi, Samuel Kiegeland, Franz Nowak, Tianyang Xu 0002, Ethan Wilcox, Alex Warstadt, Ryan Cotterell, Mario Giulianelli |
ACL (1) | 8 |
| 2025 | Structure-Conditional Minimum Bayes Risk DecodingabstractMinimum Bayes Risk (MBR) decoding has seen renewed interest as an alternative to traditional generation strategies.While MBR has proven effective in machine translation, where the variability of a language model's outcome space is naturally constrained, it may face challenges in more open-ended tasks such as dialogue or instruction-following.We hypothesise that in such settings, applying MBR with standard similarity-based utility functions may result in selecting responses that are broadly representative of the model's distribution, yet sub-optimal with respect to any particular grouping of generations that share an underlying latent structure.In this work, we introduce three lightweight adaptations to the utility function, designed to make MBR more sensitive to structural variability in the outcome space.To test our hypothesis, we curate a dataset capturing three representative types of latent structuredialogue act, emotion, and response structure (e.g., a sentence, a paragraph, or a list)-and we propose two metrics to evaluate the structural optimality of MBR.Our analysis demonstrates that common similarity-based utility functions fall short by these metrics.In contrast, our proposed adaptations considerably improve structural optimality.Finally, we evaluate our approaches on real-world instruction-following benchmarks, AlpacaEval and MT-Bench, and show that increased structural sensitivity improves generation quality by up to 13.7 percentage points in win rate. 1 Bryan Eikema, Anna Rutkiewicz, Mario Giulianelli |
EMNLP | 3 |
| 2025 | Playpen: An Environment for Exploring Learning From Dialogue Game FeedbackabstractNicola Horst, Davide Mazzaccara, Antonia Schmidt, Michael Sullivan, Filippo Momentè, Luca Franceschetti, Philipp Sadler, Sherzod Hakimov, Alberto Testoni, Raffaella Bernardi, Raquel Fernández, Alexander Koller, Oliver Lemon, David Schlangen, Mario Giulianelli, Alessandro Suglia. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Nicola Horst, Davide Mazzaccara, Antonia Schmidt, Michael Sullivan, Filippo Momentè, Luca Franceschetti, Philipp Sadler, Sherzod Hakimov, Alberto Testoni, Raffaella Bernardi, Raquel Fernández, Alexander Koller, Oliver Lemon, David Schlangen, Mario Giulianelli, Alessandro Suglia |
EMNLP | 15 |
| 2025 | Language Models over Canonical Byte-Pair EncodingsabstractModern language models represent probability distributions over character strings as distributions over (shorter) token strings derived via a deterministic tokenizer, such as byte-pair encoding. While this approach is highly effective at scaling up language models to large corpora, its current incarnations have a concerning property: the model assigns nonzero probability mass to an exponential number of noncanonical token encodings of each character string—these are token strings that decode to valid character strings but are impossible under the deterministic tokenizer (i.e., they will never be seen in any training corpus, no matter how large). This misallocation is both erroneous, as noncanonical strings never appear in training data, and wasteful, diverting probability mass away from plausible outputs. These are avoidable mistakes! In this work, we propose methods to enforce canonicality in token-level language models, ensuring that only canonical token strings are assigned positive probability. We present two approaches: (1) canonicality by conditioning, leveraging test-time inference strategies without additional training, and (2) canonicality by construction, a model parameterization that guarantees canonical outputs but requires training. We demonstrate that fixing canonicality mistakes improves the likelihood of held-out data for several models and corpora. Tim Vieira, Tianyu Liu 0004, Clemente Pasti, Yahya Emara, Brian DuSell, Benjamin LeBrun, Mario Giulianelli, Juan Luis Gastaldi, Timothy J. O'Donnell, Ryan Cotterell |
ICML | 7 |
| 2025 | From Language Models over Tokens to Language Models over CharactersabstractModern language models are internally—and mathematically—distributions over token strings rather than character strings, posing numerous challenges for programmers building user applications on top of them. For example, if a prompt is specified as a character string, it must be tokenized before passing it to the token-level language model. Thus, the tokenizer and consequent processing are very sensitive to the specification of the prompt (e.g., whether the prompt ends with a space or not). This paper presents algorithms for converting token-level language models to character-level ones. We present both exact and approximate algorithms. In the empirical portion of the paper, we benchmark the practical runtime and approximation quality. Across four publicly available language models, we find that—even with a small computation budget—our method is able to accurately approximate the character-level distribution at reasonably fast speeds, and that a significant improvement in the language model’s compression rate (bits/byte) is achieved. Tim Vieira, Benjamin LeBrun, Mario Giulianelli, Juan Luis Gastaldi, Brian DuSell, John Terilla, Timothy J. O'Donnell, Ryan Cotterell |
ICML | 3 |
| 2025 | Establishing Best Practices in Building Rigorous Agentic BenchmarksabstractBenchmarks are essential for quantitatively tracking progress in AI. As AI agents become increasingly capable, researchers and practitioners have introduced agentic benchmarks to evaluate agents on complex, real-world tasks. These benchmarks typically measure agent capabilities by evaluating task outcomes via specific reward designs. However, we show that many agentic benchmarks have issues in task setup or reward design. For example, SWE-bench-Verified uses insufficient test cases, while $\tau$-bench counts empty responses as successes. Such issues can lead to under- or overestimation of agents’ performance by up to 100% in relative terms. To make agentic evaluation rigorous, we introduce the Agentic Benchmark Checklist (ABC), a set of guidelines that we synthesized from our benchmark-building experience, a survey of best practices, and previously reported issues. When applied to CVE-Bench, a benchmark with a particularly complex evaluation design, ABC reduces performance overestimation by 33%. Yuxuan Zhu 0003, Tengjun Jin, Yada Pruksachatkun, Andy Zhang, Sasha Cui, Sayash Kapoor, Shayne Longpre, Kevin Meng, Rebecca Weiss, Fazl Barez, Rahul Gupta 0001, Jwala Dhamala, Jacob Merizian, Mario Giulianelli, Harry Coppock, Cozmin Ududec, Antony Kellermann, Jasjeet S. Sekhon, Jacob Steinhardt, Sarah Schwettmann, Arvind Narayanan, Matei Zaharia, Ion Stoica, Percy Liang, Daniel Kang 0001 |
NeurIPS | 15 |
| 2024 | Efficiency and Effectiveness in Task-Oriented Dialogue: On Construction Repetition, Information Rate, and Task SuccessabstractWe investigate the roles that efficiency and effectiveness play in speakers’ repetition of shared word sequences, or constructions, in task-oriented dialogue. We find that repeating constructions has negative effects on information rate and positive effects on rate of delivery, that information rate managing strategies are predictive of task success, and that this varies by the communicative function of the constructions being repeated. More effective dialogue is characterised by greater levels of shared construction usage and more efficient task-related repetition; while task-agnostic repetition can seem redundant, it can serve important efficiency and effectiveness functions. Our results provide a nuanced picture of the importance of repetition and of developing a shared lexicon for both efficiency and effectiveness in task-oriented dialogue. Jun Sen Yee, Mario Giulianelli, Arabella Sinclair |
LREC/COLING | 2 |
| 2024 | On the Proper Treatment of Tokenization in PsycholinguisticsabstractLanguage models are widely used in computational psycholinguistics to test theories that relate the negative log probability (the surprisal) of a region of interest (a substring of characters) under a language model to its cognitive cost experienced by readers, as operationalized, for example, by gaze duration on the region.However, the application of modern language models to psycholinguistic studies is complicated by the practice of using tokenization as an intermediate step in training a model.Doing so results in a language model over token strings rather than one over character strings.Vexingly, regions of interest are generally misaligned with these token strings.The paper argues that token-level language models should be (approximately) marginalized into character-level language models before they are used in psycholinguistic studies to compute the surprisal of a region of interest; then, the marginalized character-level language model can be used to compute the surprisal of an arbitrary character substring, which we term a focal area, that the experimenter may wish to use as a predictor.Our proposal of marginalizing a token-level model into a character-level one solves this misalignment issue independently of the tokenization scheme.Empirically, we discover various focal areas whose surprisal is a better psychometric predictor than the surprisal of the region of interest itself. Mario Giulianelli, Luca Malagutti, Juan Luis Gastaldi, Brian DuSell, Tim Vieira, Ryan Cotterell |
EMNLP | 1 |
| 2024 | Towards a Similarity-adjusted Surprisal TheoryabstractSurprisal theory posits that the cognitive effort required to comprehend a word is determined by its contextual predictability, quantified as surprisal.Traditionally, surprisal theory treats words as distinct entities, overlooking any potential similarity between them.Giulianelli et al. (2023) address this limitation by introducing information value, a measure of predictability designed to account for similarities between communicative units.Our work leverages Ricotta and Szeidl's (2006) diversity index to extend surprisal into a metric that we term similarity-adjusted surprisal, exposing a mathematical relationship between surprisal and information value.Similarity-adjusted surprisal aligns with information value when considering graded similarities and reduces to standard surprisal when words are treated as distinct.Experimental results with reading time data indicate that similarity-adjusted surprisal adds predictive power beyond standard surprisal for certain datasets, suggesting it serves as a complementary measure of comprehension effort. Clara Meister, Mario Giulianelli, Tiago Pimentel |
EMNLP | 2 |
| 2024 | Surprise! Uniform Information Density Isn't the Whole Story: Predicting Surprisal Contours in Long-form DiscourseabstractThe Uniform Information Density (UID) hypothesis posits that speakers tend to distribute information evenly across linguistic units to achieve efficient communication.Of course, information rate in texts and discourses is not perfectly uniform.While these fluctuations can be viewed as theoretically uninteresting noise on top of a uniform target, another explanation is that UID is not the only functional pressure regulating information content in a language.Speakers may also seek to maintain interest, adhere to writing conventions, and build compelling arguments.In this paper, we propose one such functional pressure; namely that speakers modulate information rate based on location within a hierarchically-structured model of discourse.We term this the Structured Context Hypothesis and test it by predicting the surprisal contours of naturally occurring discourses extracted from large language models using predictors derived from discourse structure.We find that hierarchical predictors are significant predictors of a discourse's information contour and that deeply nested hierarchical predictors are more predictive than shallow ones.This work takes an initial step beyond UID to propose testable hypotheses for why the information rate fluctuates in predictable ways.https://github.com/rycolab/ surprisal-discourse Eleftheria Tsipidi, Franz Nowak, Ryan Cotterell, Ethan Wilcox, Mario Giulianelli, Alex Warstadt |
EMNLP | 5 |
| 2023 | Interpretable Word Sense Representations via Definition Generation: The Case of Semantic Change AnalysisabstractWe propose using automatically generated natural language definitions of contextualised word usages as interpretable word and word sense representations.Given a collection of usage examples for a target word, and the corresponding data-driven usage clusters (i.e., word senses), a definition is generated for each usage with a specialised Flan-T5 language model, and the most prototypical definition in a usage cluster is chosen as the sense label.We demonstrate how the resulting sense labels can make existing approaches to semantic change analysis more interpretable, and how they can allow users-historical linguists, lexicographers, or social scientists-to explore and intuitively explain diachronic trajectories of word meaning.Semantic change analysis is only one of many possible applications of the 'definitions as representations' paradigm.Beyond being human-readable, contextualised definitions also outperform token or usage sentence embeddings in word-in-context semantic similarity judgements, making them a new promising type of lexical representation for NLP. Mario Giulianelli, Iris Luden, Raquel Fernández, Andrey Kutuzov |
ACL (1) | 1 |
| 2023 | Attribution and Alignment: Effects of Local Context Repetition on Utterance Production and Comprehension in DialogueabstractLanguage models are often used as the backbone of modern dialogue systems.These models are pre-trained on large amounts of written fluent language.Repetition is typically penalised when evaluating language model generations.However, it is a key component of dialogue.Humans use local and partner specific repetitions; these are preferred by human users and lead to more successful communication in dialogue.In this study, we evaluate (a) whether language models produce humanlike levels of repetition in dialogue, and (b) what are the processing mechanisms related to lexical re-use they use during comprehension.We believe that such joint analysis of model production and comprehension behaviour can inform the development of cognitively inspired dialogue generation systems. Aron Molnar, Jaap Jumelet, Mario Giulianelli, Arabella Sinclair |
CoNLL | 3 |
| 2023 | What Comes Next? Evaluating Uncertainty in Neural Text Generators Against Human Production VariabilityabstractIn Natural Language Generation (NLG) tasks, for any input, multiple communicative goals are plausible, and any goal can be put into words, or produced, in multiple ways.We characterise the extent to which human production varies lexically, syntactically, and semantically across four NLG tasks, connecting human production variability to aleatoric or data uncertainty.We then inspect the space of output strings shaped by a generation system's predicted probability distribution and decoding algorithm to probe its uncertainty.For each test input, we measure the generator's calibration to human production variability.Following this instance-level approach, we analyse NLG models and decoding strategies, demonstrating that probing a generator with multiple samples and, when possible, multiple references, provides the level of detail necessary to gain understanding of a model's representation of uncertainty. 1 * Equal contribution. 1 https://github.com/dmg-illc/nlg-uncertainty-probes Mario Giulianelli, Joris Baan, Wilker Aziz, Raquel Fernández, Barbara Plank |
EMNLP | 1 |
| 2023 | Information Value: Measuring Utterance Predictability as Distance from Plausible AlternativesabstractWe present information value, a measure which quantifies the predictability of an utterance relative to a set of plausible alternatives.We introduce a method to obtain interpretable estimates of information value using neural text generators, and exploit their psychometric predictive power to investigate the dimensions of predictability that drive human comprehension behaviour.Information value is a stronger predictor of utterance acceptability in written and spoken dialogue than aggregates of token-level surprisal and it is complementary to surprisal for predicting eye-tracked reading times. 1 Mario Giulianelli, Sarenne Wallbridge, Raquel Fernández |
EMNLP | 1 |
| 2022 | Towards Pragmatic Production Strategies for Natural Language Generation TasksabstractThis position paper proposes a conceptual framework for the design of Natural Language Generation (NLG) systems that follow efficient and effective production strategies in order to achieve complex communicative goals.In this general framework, efficiency is characterised as the parsimonious regulation of production and comprehension costs while effectiveness is measured with respect to task-oriented and contextually grounded communicative goals.We provide concrete suggestions for the estimation of goals, costs, and utility via modern statistical methods, demonstrating applications of our framework to the classic pragmatic task of visually grounded referential games and to abstractive text summarisation, two popular generation tasks with real-world applications.In sum, we advocate for the development of NLG systems that learn to make pragmatic production decisions from experience, by reasoning about goals, costs, and utility in a human-like way. Mario Giulianelli |
EMNLP | 1 |
| 2021 | Analysing Human Strategies of Information Transmission as a Function of Discourse ContextabstractSpeakers are thought to use rational information transmission strategies for efficient communication (Genzel and Charniak, 2002;Aylett and Turk, 2004;Jaeger and Levy, 2007).Previous work analysing these strategies in sentence production has failed to take into account how the information content of sentences varies as a function of the available discourse context.In this study, we estimate sentence information content within discourse context.We find that speakers transmit information at a stable rate-i.e., rationally-in English newspaper articles but that this rate decreases in spoken open domain and written task-oriented dialogues.We also observe that speakers' choices are not oriented towards local uniformity of information, which is another hypothesised rational strategy.We suggest that a more faithful model of communication should explicitly include production costs and goal-oriented rewards. Mario Giulianelli, Raquel Fernández |
CoNLL | 1 |
| 2021 | Grammatical Profiling for Semantic Change DetectionabstractSemantics, morphology and syntax are strongly interdependent. However, the majority of computational methods for semantic change detection use distributional word representations which encode mostly semantics. We investigate an alternative method, grammatical profiling, based entirely on changes in the morphosyntactic behaviour of words. We demonstrate that it can be used for semantic change detection and even outperforms some distributional semantic methods. We present an in-depth qualitative and quantitative analysis of the predictions made by our grammatical profiling system, showing that they are plausible and interpretable. Andrey Kutuzov, Lidia Pivovarova, Mario Giulianelli |
CoNLL | 3 |
| 2021 | Is Information Density Uniform in Task-Oriented Dialogues?abstractThe Uniform Information Density principle states that speakers plan their utterances to reduce fluctuations in the density of the information transmitted.In this paper, we test whether, and within which contextual units this principle holds in task-oriented dialogues.We show that there is evidence supporting the principle in written dialogues where participants play a cooperative reference game as well as in spoken dialogues involving instruction giving and following.Our study underlines the importance of identifying the relevant contextual components, showing that information content increases particularly within topically and referentially related contextual units. Mario Giulianelli, Arabella Sinclair, Raquel Fernández |
EMNLP (1) | 1 |
| 2020 | Analysing Lexical Semantic Change with Contextualised Word RepresentationsabstractThis paper presents the first unsupervised approach to lexical semantic change that makes use of contextualised word representations.We propose a novel method that exploits the BERT neural language model to obtain representations of word usages, clusters these representations into usage types, and measures change along time with three proposed metrics.We create a new evaluation dataset and show that the model representations and the detected semantic shifts are positively correlated with human judgements.Our extensive qualitative analysis demonstrates that our method captures a variety of synchronic and diachronic linguistic phenomena.We expect our work to inspire further research in this direction. Mario Giulianelli, Marco Del Tredici, Raquel Fernández |
ACL | 1 |
| 2020 | Refer, Reuse, Reduce: Generating Subsequent References in Visual and Conversational ContextsabstractDialogue participants often refer to entities or situations repeatedly within a conversation, which contributes to its cohesiveness.Subsequent references exploit the common ground accumulated by the interlocutors and hence have several interesting properties, namely, they tend to be shorter and reuse expressions that were effective in previous mentions.In this paper, we tackle the generation of first and subsequent references in visually grounded dialogue.We propose a generation model that produces referring utterances grounded in both the visual and the conversational context.To assess the referring effectiveness of its output, we also implement a reference resolution system.Our experiments and analyses show that the model produces better, more effective referring utterances than a model not grounded in the dialogue context, and generates subsequent references that exhibit linguistic patterns akin to humans. Ece Takmaz, Mario Giulianelli, Sandro Pezzelle, Arabella Sinclair, Raquel Fernández |
EMNLP (1) | 2 |