Marianna Apidianaki

dblp:53/1712 · DBLP profile ↗
← Back
30ranked-venue papers
10as first author
12since 2021 · last 2025
0000-0002-8785-7722ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 10 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Calibrating Large Language Models with Sample Consistency
abstract
Accurately gauging the confidence level of Large Language Models' (LLMs) predictions is pivotal for their reliable application. However, LLMs are often uncalibrated inherently and elude conventional calibration techniques due to their proprietary nature and massive scale. In this work, we derive model confidence from the distribution of multiple randomly sampled generations, using three measures of consistency. We extensively evaluate eleven open and closed-source models on nine reasoning datasets. Results show that consistency-based calibration methods outperform existing post-hoc approaches in terms of calibration error. Meanwhile, we find that factors such as intermediate explanations, model scaling, and larger sample sizes enhance calibration, while instruction-tuning makes calibration more difficult. Moreover, confidence scores obtained from consistency can potentially enhance model performance. Finally, we offer guidance on choosing suitable consistency metrics for calibration, tailored to model characteristics such as the exposure to instruction-tuning and RLHF.
Qing Lyu 0001, Kumar Shridhar, Chaitanya Malaviya, Li Zhang 0039, Yanai Elazar, Niket Tandon, Marianna Apidianaki, Mrinmaya Sachan, Chris Callison-Burch
AAAI7
2025 Latent Space Interpretation for Stylistic Analysis and Explainable Authorship Attribution
abstract
Recent state-of-the-art authorship attribution methods learn authorship representations of text in a latent, uninterpretable space, which hinders their usability in real-world applications. We propose a novel approach for interpreting learned embeddings by identifying representative points in the latent space and leveraging large language models to generate informative natural language descriptions of the writing style associated with each point. We evaluate the alignment between our interpretable and latent spaces and demonstrate superior prediction agreement over baseline methods. Additionally, we conduct a human evaluation to assess the quality of these style descriptions and validate their utility in explaining the latent space. Finally, we show that human performance on the challenging authorship attribution task improves by +20% on average when aided with explanations from our method.
Milad Alshomary, Narutatsu Ri, Marianna Apidianaki, Ajay Patel, Smaranda Muresan, Kathy McKeown
COLING3
2025 StyleDistance: Stronger Content-Independent Style Embeddings with Synthetic Parallel Examples
abstract
Ajay Patel, Jiacheng Zhu, Justin Qiu, Zachary Horvitz, Marianna Apidianaki, Kathleen McKeown, Chris Callison-Burch. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Ajay Patel, Justin Qiu, Zachary Horvitz, Marianna Apidianaki, Kathy McKeown, Chris Callison-Burch
NAACL (Long Papers)5
2024 Adjusting Interpretable Dimensions in Embedding Space with Human Judgments
abstract
Katrin Erk, Marianna Apidianaki. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Katrin Erk, Marianna Apidianaki
NAACL-HLT2
2024 Language Learning, Representation, and Processing in Humans and Machines: Introduction to the Special Issue
abstract
Abstract Large Language Models (LLMs) and humans acquire knowledge about language without direct supervision. LLMs do so by means of specific training objectives, while humans rely on sensory experience and social interaction. This parallelism has created a feeling in NLP and cognitive science that a systematic understanding of how LLMs acquire and use the encoded knowledge could provide useful insights for studying human cognition. Conversely, methods and findings from the field of cognitive science have occasionally inspired language model development. Yet, the differences in the way that language is processed by machines and humans—in terms of learning mechanisms, amounts of data used, grounding and access to different modalities—make a direct translation of insights challenging. The aim of this edited volume has been to create a forum of exchange and debate along this line of research, inviting contributions that further elucidate similarities and differences between humans and LLMs.
Marianna Apidianaki, Abdellah Fourtassi, Sebastian Padó
Comput. Linguistics1
2024 Towards Faithful Model Explanation in NLP: A Survey
abstract
Abstract End-to-end neural Natural Language Processing (NLP) models are notoriously difficult to understand. This has given rise to numerous efforts towards model explainability in recent years. One desideratum of model explanation is faithfulness, that is, an explanation should accurately represent the reasoning process behind the model’s prediction. In this survey, we review over 110 model explanation methods in NLP through the lens of faithfulness. We first discuss the definition and evaluation of faithfulness, as well as its significance for explainability. We then introduce recent advances in faithful explanation, grouping existing approaches into five categories: similarity-based methods, analysis of model-internal structures, backpropagation-based methods, counterfactual intervention, and self-explanatory models. For each category, we synthesize its representative studies, strengths, and weaknesses. Finally, we summarize their common virtues and remaining challenges, and reflect on future work directions towards faithful explainability in NLP.
Qing Lyu 0001, Marianna Apidianaki, Chris Callison-Burch
Comput. Linguistics2
2023 Explanation-based Finetuning Makes Models More Robust to Spurious Cues
abstract
Josh Magnus Ludan, Yixuan Meng, Tai Nguyen, Saurabh Shah, Qing Lyu, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Josh Magnus Ludan, Yixuan Meng, Tai Nguyen 0005, Saurabh Shah, Qing Lyu 0001, Marianna Apidianaki, Chris Callison-Burch
ACL (1)6
2023 Faithful Chain-of-Thought Reasoning
abstract
Qing Lyu, Shreya Havaldar, Adam Stein, Li Zhang, Delip Rao, Eric Wong, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Qing Lyu 0001, Shreya Havaldar, Adam Stein, Li Zhang 0039, Delip Rao, Eric Wong 0001, Marianna Apidianaki, Chris Callison-Burch
IJCNLP (1)7
2023 From Word Types to Tokens and Back: A Survey of Approaches to Word Meaning Representation and Interpretation
abstract
Abstract Vector-based word representation paradigms situate lexical meaning at different levels of abstraction. Distributional and static embedding models generate a single vector per word type, which is an aggregate across the instances of the word in a corpus. Contextual language models, on the contrary, directly capture the meaning of individual word instances. The goal of this survey is to provide an overview of word meaning representation methods, and of the strategies that have been proposed for improving the quality of the generated vectors. These often involve injecting external knowledge about lexical semantic relationships, or refining the vectors to describe different senses. The survey also covers recent approaches for obtaining word type-level representations from token-level ones, and for combining static and contextualized representations. Special focus is given to probing and interpretation studies aimed at discovering the lexical semantic knowledge that is encoded in contextualized representations. The challenges posed by this exploration have motivated the interest towards static embedding derivation from contextualized embeddings, and for methods aimed at improving the similarity estimates that can be drawn from the space of contextual language models.
Marianna Apidianaki
Comput. Linguistics1
2022 Is "My Favorite New Movie" My Favorite Movie? Probing the Understanding of Recursive Noun Phrases
abstract
Qing Lyu, Zheng Hua, Daoxin Li, Li Zhang, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Qing Lyu 0001, Daoxin Li, Li Zhang 0039, Marianna Apidianaki, Chris Callison-Burch
NAACL-HLT5
2021 Scalar Adjective Identification and Multilingual Ranking
abstract
The intensity relationship that holds between scalar adjectives (e.g., nice < great < wonderful) is highly relevant for natural language inference and common-sense reasoning.Previous research on scalar adjective ranking has focused on English, mainly due to the availability of datasets for evaluation.We introduce a new multilingual dataset in order to promote research on scalar adjectives in new languages.We perform a series of experiments and set performance baselines on this dataset, using monolingual and multilingual contextual language models.Additionally, we introduce a new binary classification task for English scalar adjective identification which examines the models' ability to distinguish scalar from relational adjectives.We probe contextualised representations and report baseline results for future comparison on this task.
Aina Garí Soler, Marianna Apidianaki
NAACL-HLT2
2021 Let's Play Mono-Poly: BERT Can Reveal Words' Polysemy Level and Partitionability into Senses
abstract
Pre-trained language models (LMs) encode rich information about linguistic structure but their knowledge about lexical polysemy remains unclear. We propose a novel experimental setup for analyzing this knowledge in LMs specifically trained for different languages (English, French, Spanish, and Greek) and in multilingual BERT. We perform our analysis on datasets carefully designed to reflect different sense distributions, and control for parameters that are highly correlated with polysemy such as frequency and grammatical category. We demonstrate that BERT-derived representations reflect words’ polysemy level and their partitionability into senses. Polysemy-related information is more clearly present in English BERT embeddings, but models in other languages also manage to establish relevant distinctions between words at different polysemy levels. Our results contribute to a better understanding of the knowledge encoded in contextualized representations and open up new avenues for multilingual lexical semantics research.
Aina Garí Soler, Marianna Apidianaki
Trans. Assoc. Comput. Linguistics2
2020 BERT Knows Punta Cana is not just beautiful, it's gorgeous: Ranking Scalar Adjectives with Contextualised Representations
abstract
Adjectives like pretty, beautiful and gorgeous describe positive properties of the nouns they modify but with different intensity. These differences are important for natural language understanding and reasoning. We propose a novel BERT-based approach to intensity detection for scalar adjectives. We model intensity by vectors directly derived from contextualised representations and show they can successfully rank scalar adjectives. We evaluate our models both intrinsically, on gold standard datasets, and on an Indirect Question Answering task. Our results demonstrate that BERT encodes rich knowledge about the semantics of scalar adjectives, and is able to provide better quality intensity rankings than static embeddings and previous models with access to dedicated resources.
Aina Garí Soler, Marianna Apidianaki
EMNLP (1)2
2019 SUM-QE: a BERT-based Summary Quality Estimation Model
abstract
Stratos Xenouleas, Prodromos Malakasiotis, Marianna Apidianaki, Ion Androutsopoulos. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Stratos Xenouleas, Prodromos Malakasiotis, Marianna Apidianaki, Ion Androutsopoulos
EMNLP/IJCNLP (1)3
2018 Learning Scalar Adjective Intensity from Paraphrases
abstract
Adjectives like warm, hot, and scalding all describe temperature but differ in intensity.Understanding these differences between adjectives is a necessary part of reasoning about natural language.We propose a new paraphrasebased method to automatically learn the relative intensity relation that holds between a pair of scalar adjectives.Our approach analyzes over 36k adjectival pairs from the Paraphrase Database under the assumption that, for example, paraphrase pair really hot ↔ scalding suggests that hot < scalding.We show that combining this paraphrase evidence with existing, complementary pattern-and lexicon-based approaches improves the quality of systems for automatically ordering sets of scalar adjectives and inferring the polarity of indirect answers to yes/no questions.
Anne Cocos, Veronica Wharton, Ellie Pavlick, Marianna Apidianaki, Chris Callison-Burch
EMNLP4
2018 Comparing Constraints for Taxonomic Organization
abstract
Anne Cocos, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Anne Cocos, Marianna Apidianaki, Chris Callison-Burch
NAACL-HLT2
2018 Simplification Using Paraphrases and Context-Based Lexical Substitution
abstract
Reno Kriz, Eleni Miltsakaki, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Reno Kriz, Eleni Miltsakaki, Marianna Apidianaki, Chris Callison-Burch
NAACL-HLT3
2017 Learning Translations via Matrix Completion
abstract
Derry Tanti Wijaya, Brendan Callahan, John Hewitt, Jie Gao, Xiao Ling, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. 2017.
Derry Wijaya, Brendan Callahan, John Hewitt, Marianna Apidianaki, Chris Callison-Burch
EMNLP6
2016 Vector-space models for PPDB paraphrase ranking in context
abstract
The PPDB is an automatically built database which contains millions of paraphrases in different languages.Paraphrases in this resource are associated with features that serve to their ranking and reflect paraphrase quality.This context-unaware ranking captures the semantic similarity of paraphrases but cannot serve to estimate their adequacy in specific contexts.We propose to use vector-space semantic models for selecting PPDB paraphrases that preserve the meaning of specific text fragments.This is the first work that addresses the substitutability of PPDB paraphrases in context.We show that vector-space models of meaning can be successfully applied to this task and increase the benefit brought by the use of the PPDB resource in applications.
Marianna Apidianaki
EMNLP1
2016 Datasets for Aspect-Based Sentiment Analysis in French
Marianna Apidianaki, Xavier Tannier, Cécile Richart
LREC1
2016 Word Sense Clustering and Clusterability
abstract
Word sense disambiguation and the related field of automated word sense induction traditionally assume that the occurrences of a lemma can be partitioned into senses. But this seems to be a much easier task for some lemmas than others. Our work builds on recent work that proposes describing word meaning in a graded fashion rather than through a strict partition into senses; in this article we argue that not all lemmas may need the more complex graded analysis, depending on their partitionability. Although there is plenty of evidence from previous studies and from the linguistics literature that there is a spectrum of partitionability of word meanings, this is the first attempt to measure the phenomenon and to couple the machine learning literature on clusterability with word usage data used in computational linguistics. We propose to operationalize partitionability as clusterability, a measure of how easy the occurrences of a lemma are to cluster. We test two ways of measuring clusterability: (1) existing measures from the machine learning literature that aim to measure the goodness of optimal k-means clusterings, and (2) the idea that if a lemma is more clusterable, two clusterings based on two different “views” of the same data points will be more congruent. The two views that we use are two different sets of manually constructed lexical substitutes for the target lemma, on the one hand monolingual paraphrases, and on the other hand translations. We apply automatic clustering to the manual annotations. We use manual annotations because we want the representations of the instances that we cluster to be as informative and “clean” as possible. We show that when we control for polysemy, our measures of clusterability tend to correlate with partitionability, in particular some of the type-(1) clusterability measures, and that these measures outperform a baseline that relies on the amount of overlap in a soft clustering.
Diana McCarthy, Marianna Apidianaki, Katrin Erk
Comput. Linguistics2
2014 Global Methods for Cross-lingual Semantic Role and Predicate Labelling
Lonneke van der Plas, Marianna Apidianaki, Chenhua Chen
COLING2
2014 Semantic Clustering of Pivot Paraphrases
Marianna Apidianaki, Emilia Verzeni, Diana McCarthy
LREC1
2012 Applying cross-lingual WSD to wordnet development
Marianna Apidianaki, Benoît Sagot
LREC1
2012 Boosting the Coverage of a Semantic Lexicon by Automatically Extracted Event Nominalizations
Kata Gábor, Marianna Apidianaki, Benoît Sagot, Éric Villemonte de la Clergerie
LREC2
2011 Latent Semantic Word Sense Induction and Disambiguation
Tim Van de Cruys, Marianna Apidianaki
ACL2
2011 A Quantitative Evaluation of Global Word Sense Induction
Marianna Apidianaki, Tim Van de Cruys
CICLing (1)1
2009 Data-Driven Semantic Analysis for Multilingual WSD and Lexical Selection in Translation
Marianna Apidianaki
EACL1
2009 Capturing Lexical Variation in MT Evaluation Using Automatically Built Sense-Cluster Inventories
Marianna Apidianaki, Yifan He 0007, Andy Way
PACLIC1
2008 Translation-oriented Word Sense Induction Based on Parallel Corpora
Marianna Apidianaki
LREC1