VLDB 2026 Research / reviewers in the wild / expert
Marianna Apidianaki
dblp:53/1712
· DBLP profile ↗
30ranked-venue papers
10as first author
12since 2021 · last 2025
0000-0002-8785-7722ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 10 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Calibrating Large Language Models with Sample ConsistencyabstractAccurately gauging the confidence level of Large Language Models' (LLMs) predictions is pivotal for their reliable application. However, LLMs are often uncalibrated inherently and elude conventional calibration techniques due to their proprietary nature and massive scale. In this work, we derive model confidence from the distribution of multiple randomly sampled generations, using three measures of consistency. We extensively evaluate eleven open and closed-source models on nine reasoning datasets. Results show that consistency-based calibration methods outperform existing post-hoc approaches in terms of calibration error. Meanwhile, we find that factors such as intermediate explanations, model scaling, and larger sample sizes enhance calibration, while instruction-tuning makes calibration more difficult. Moreover, confidence scores obtained from consistency can potentially enhance model performance. Finally, we offer guidance on choosing suitable consistency metrics for calibration, tailored to model characteristics such as the exposure to instruction-tuning and RLHF. Qing Lyu 0001, Kumar Shridhar, Chaitanya Malaviya, Li Zhang 0039, Yanai Elazar, Niket Tandon, Marianna Apidianaki, Mrinmaya Sachan, Chris Callison-Burch |
AAAI | 7 |
| 2025 | Latent Space Interpretation for Stylistic Analysis and Explainable Authorship AttributionabstractRecent state-of-the-art authorship attribution methods learn authorship representations of text in a latent, uninterpretable space, which hinders their usability in real-world applications. We propose a novel approach for interpreting learned embeddings by identifying representative points in the latent space and leveraging large language models to generate informative natural language descriptions of the writing style associated with each point. We evaluate the alignment between our interpretable and latent spaces and demonstrate superior prediction agreement over baseline methods. Additionally, we conduct a human evaluation to assess the quality of these style descriptions and validate their utility in explaining the latent space. Finally, we show that human performance on the challenging authorship attribution task improves by +20% on average when aided with explanations from our method. Milad Alshomary, Narutatsu Ri, Marianna Apidianaki, Ajay Patel, Smaranda Muresan, Kathy McKeown |
COLING | 3 |
| 2025 | StyleDistance: Stronger Content-Independent Style Embeddings with Synthetic Parallel ExamplesabstractAjay Patel, Jiacheng Zhu, Justin Qiu, Zachary Horvitz, Marianna Apidianaki, Kathleen McKeown, Chris Callison-Burch. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Ajay Patel, Justin Qiu, Zachary Horvitz, Marianna Apidianaki, Kathy McKeown, Chris Callison-Burch |
NAACL (Long Papers) | 5 |
| 2024 | Adjusting Interpretable Dimensions in Embedding Space with Human JudgmentsabstractKatrin Erk, Marianna Apidianaki. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Katrin Erk, Marianna Apidianaki |
NAACL-HLT | 2 |
| 2024 | Language Learning, Representation, and Processing in Humans and Machines: Introduction to the Special IssueabstractAbstract Large Language Models (LLMs) and humans acquire knowledge about language without direct supervision. LLMs do so by means of specific training objectives, while humans rely on sensory experience and social interaction. This parallelism has created a feeling in NLP and cognitive science that a systematic understanding of how LLMs acquire and use the encoded knowledge could provide useful insights for studying human cognition. Conversely, methods and findings from the field of cognitive science have occasionally inspired language model development. Yet, the differences in the way that language is processed by machines and humans—in terms of learning mechanisms, amounts of data used, grounding and access to different modalities—make a direct translation of insights challenging. The aim of this edited volume has been to create a forum of exchange and debate along this line of research, inviting contributions that further elucidate similarities and differences between humans and LLMs. Marianna Apidianaki, Abdellah Fourtassi, Sebastian Padó |
Comput. Linguistics | 1 |
| 2024 | Towards Faithful Model Explanation in NLP: A SurveyabstractAbstract End-to-end neural Natural Language Processing (NLP) models are notoriously difficult to understand. This has given rise to numerous efforts towards model explainability in recent years. One desideratum of model explanation is faithfulness, that is, an explanation should accurately represent the reasoning process behind the model’s prediction. In this survey, we review over 110 model explanation methods in NLP through the lens of faithfulness. We first discuss the definition and evaluation of faithfulness, as well as its significance for explainability. We then introduce recent advances in faithful explanation, grouping existing approaches into five categories: similarity-based methods, analysis of model-internal structures, backpropagation-based methods, counterfactual intervention, and self-explanatory models. For each category, we synthesize its representative studies, strengths, and weaknesses. Finally, we summarize their common virtues and remaining challenges, and reflect on future work directions towards faithful explainability in NLP. Qing Lyu 0001, Marianna Apidianaki, Chris Callison-Burch |
Comput. Linguistics | 2 |
| 2023 | Explanation-based Finetuning Makes Models More Robust to Spurious CuesabstractJosh Magnus Ludan, Yixuan Meng, Tai Nguyen, Saurabh Shah, Qing Lyu, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Josh Magnus Ludan, Yixuan Meng, Tai Nguyen 0005, Saurabh Shah, Qing Lyu 0001, Marianna Apidianaki, Chris Callison-Burch |
ACL (1) | 6 |
| 2023 | Faithful Chain-of-Thought ReasoningabstractQing Lyu, Shreya Havaldar, Adam Stein, Li Zhang, Delip Rao, Eric Wong, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Qing Lyu 0001, Shreya Havaldar, Adam Stein, Li Zhang 0039, Delip Rao, Eric Wong 0001, Marianna Apidianaki, Chris Callison-Burch |
IJCNLP (1) | 7 |
| 2023 | From Word Types to Tokens and Back: A Survey of Approaches to Word Meaning Representation and InterpretationabstractAbstract Vector-based word representation paradigms situate lexical meaning at different levels of abstraction. Distributional and static embedding models generate a single vector per word type, which is an aggregate across the instances of the word in a corpus. Contextual language models, on the contrary, directly capture the meaning of individual word instances. The goal of this survey is to provide an overview of word meaning representation methods, and of the strategies that have been proposed for improving the quality of the generated vectors. These often involve injecting external knowledge about lexical semantic relationships, or refining the vectors to describe different senses. The survey also covers recent approaches for obtaining word type-level representations from token-level ones, and for combining static and contextualized representations. Special focus is given to probing and interpretation studies aimed at discovering the lexical semantic knowledge that is encoded in contextualized representations. The challenges posed by this exploration have motivated the interest towards static embedding derivation from contextualized embeddings, and for methods aimed at improving the similarity estimates that can be drawn from the space of contextual language models. Marianna Apidianaki |
Comput. Linguistics | 1 |
| 2022 | Is "My Favorite New Movie" My Favorite Movie? Probing the Understanding of Recursive Noun PhrasesabstractQing Lyu, Zheng Hua, Daoxin Li, Li Zhang, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Qing Lyu 0001, Daoxin Li, Li Zhang 0039, Marianna Apidianaki, Chris Callison-Burch |
NAACL-HLT | 5 |
| 2021 | Scalar Adjective Identification and Multilingual RankingabstractThe intensity relationship that holds between scalar adjectives (e.g., nice < great < wonderful) is highly relevant for natural language inference and common-sense reasoning.Previous research on scalar adjective ranking has focused on English, mainly due to the availability of datasets for evaluation.We introduce a new multilingual dataset in order to promote research on scalar adjectives in new languages.We perform a series of experiments and set performance baselines on this dataset, using monolingual and multilingual contextual language models.Additionally, we introduce a new binary classification task for English scalar adjective identification which examines the models' ability to distinguish scalar from relational adjectives.We probe contextualised representations and report baseline results for future comparison on this task. Aina Garí Soler, Marianna Apidianaki |
NAACL-HLT | 2 |
| 2021 | Let's Play Mono-Poly: BERT Can Reveal Words' Polysemy Level and Partitionability into SensesabstractPre-trained language models (LMs) encode rich information about linguistic structure but their knowledge about lexical polysemy remains unclear. We propose a novel experimental setup for analyzing this knowledge in LMs specifically trained for different languages (English, French, Spanish, and Greek) and in multilingual BERT. We perform our analysis on datasets carefully designed to reflect different sense distributions, and control for parameters that are highly correlated with polysemy such as frequency and grammatical category. We demonstrate that BERT-derived representations reflect words’ polysemy level and their partitionability into senses. Polysemy-related information is more clearly present in English BERT embeddings, but models in other languages also manage to establish relevant distinctions between words at different polysemy levels. Our results contribute to a better understanding of the knowledge encoded in contextualized representations and open up new avenues for multilingual lexical semantics research. Aina Garí Soler, Marianna Apidianaki |
Trans. Assoc. Comput. Linguistics | 2 |
| 2020 | BERT Knows Punta Cana is not just beautiful, it's gorgeous: Ranking Scalar Adjectives with Contextualised RepresentationsabstractAdjectives like pretty, beautiful and gorgeous describe positive properties of the nouns they modify but with different intensity. These differences are important for natural language understanding and reasoning. We propose a novel BERT-based approach to intensity detection for scalar adjectives. We model intensity by vectors directly derived from contextualised representations and show they can successfully rank scalar adjectives. We evaluate our models both intrinsically, on gold standard datasets, and on an Indirect Question Answering task. Our results demonstrate that BERT encodes rich knowledge about the semantics of scalar adjectives, and is able to provide better quality intensity rankings than static embeddings and previous models with access to dedicated resources. Aina Garí Soler, Marianna Apidianaki |
EMNLP (1) | 2 |
| 2019 | SUM-QE: a BERT-based Summary Quality Estimation ModelabstractStratos Xenouleas, Prodromos Malakasiotis, Marianna Apidianaki, Ion Androutsopoulos. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Stratos Xenouleas, Prodromos Malakasiotis, Marianna Apidianaki, Ion Androutsopoulos |
EMNLP/IJCNLP (1) | 3 |
| 2018 | Learning Scalar Adjective Intensity from ParaphrasesabstractAdjectives like warm, hot, and scalding all describe temperature but differ in intensity.Understanding these differences between adjectives is a necessary part of reasoning about natural language.We propose a new paraphrasebased method to automatically learn the relative intensity relation that holds between a pair of scalar adjectives.Our approach analyzes over 36k adjectival pairs from the Paraphrase Database under the assumption that, for example, paraphrase pair really hot ↔ scalding suggests that hot < scalding.We show that combining this paraphrase evidence with existing, complementary pattern-and lexicon-based approaches improves the quality of systems for automatically ordering sets of scalar adjectives and inferring the polarity of indirect answers to yes/no questions. Anne Cocos, Veronica Wharton, Ellie Pavlick, Marianna Apidianaki, Chris Callison-Burch |
EMNLP | 4 |
| 2018 | Comparing Constraints for Taxonomic OrganizationabstractAnne Cocos, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Anne Cocos, Marianna Apidianaki, Chris Callison-Burch |
NAACL-HLT | 2 |
| 2018 | Simplification Using Paraphrases and Context-Based Lexical SubstitutionabstractReno Kriz, Eleni Miltsakaki, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Reno Kriz, Eleni Miltsakaki, Marianna Apidianaki, Chris Callison-Burch |
NAACL-HLT | 3 |
| 2017 | Learning Translations via Matrix CompletionabstractDerry Tanti Wijaya, Brendan Callahan, John Hewitt, Jie Gao, Xiao Ling, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. 2017. Derry Wijaya, Brendan Callahan, John Hewitt, Marianna Apidianaki, Chris Callison-Burch |
EMNLP | 6 |
| 2016 | Vector-space models for PPDB paraphrase ranking in contextabstractThe PPDB is an automatically built database which contains millions of paraphrases in different languages.Paraphrases in this resource are associated with features that serve to their ranking and reflect paraphrase quality.This context-unaware ranking captures the semantic similarity of paraphrases but cannot serve to estimate their adequacy in specific contexts.We propose to use vector-space semantic models for selecting PPDB paraphrases that preserve the meaning of specific text fragments.This is the first work that addresses the substitutability of PPDB paraphrases in context.We show that vector-space models of meaning can be successfully applied to this task and increase the benefit brought by the use of the PPDB resource in applications. Marianna Apidianaki |
EMNLP | 1 |
| 2016 | Datasets for Aspect-Based Sentiment Analysis in French
Marianna Apidianaki, Xavier Tannier, Cécile Richart |
LREC | 1 |
| 2016 | Word Sense Clustering and ClusterabilityabstractWord sense disambiguation and the related field of automated word sense induction traditionally assume that the occurrences of a lemma can be partitioned into senses. But this seems to be a much easier task for some lemmas than others. Our work builds on recent work that proposes describing word meaning in a graded fashion rather than through a strict partition into senses; in this article we argue that not all lemmas may need the more complex graded analysis, depending on their partitionability. Although there is plenty of evidence from previous studies and from the linguistics literature that there is a spectrum of partitionability of word meanings, this is the first attempt to measure the phenomenon and to couple the machine learning literature on clusterability with word usage data used in computational linguistics. We propose to operationalize partitionability as clusterability, a measure of how easy the occurrences of a lemma are to cluster. We test two ways of measuring clusterability: (1) existing measures from the machine learning literature that aim to measure the goodness of optimal k-means clusterings, and (2) the idea that if a lemma is more clusterable, two clusterings based on two different “views” of the same data points will be more congruent. The two views that we use are two different sets of manually constructed lexical substitutes for the target lemma, on the one hand monolingual paraphrases, and on the other hand translations. We apply automatic clustering to the manual annotations. We use manual annotations because we want the representations of the instances that we cluster to be as informative and “clean” as possible. We show that when we control for polysemy, our measures of clusterability tend to correlate with partitionability, in particular some of the type-(1) clusterability measures, and that these measures outperform a baseline that relies on the amount of overlap in a soft clustering. Diana McCarthy, Marianna Apidianaki, Katrin Erk |
Comput. Linguistics | 2 |
| 2014 | Global Methods for Cross-lingual Semantic Role and Predicate Labelling
Lonneke van der Plas, Marianna Apidianaki, Chenhua Chen |
COLING | 2 |
| 2014 | Semantic Clustering of Pivot Paraphrases
Marianna Apidianaki, Emilia Verzeni, Diana McCarthy |
LREC | 1 |
| 2012 | Applying cross-lingual WSD to wordnet development
Marianna Apidianaki, Benoît Sagot |
LREC | 1 |
| 2012 | Boosting the Coverage of a Semantic Lexicon by Automatically Extracted Event Nominalizations
Kata Gábor, Marianna Apidianaki, Benoît Sagot, Éric Villemonte de la Clergerie |
LREC | 2 |
| 2011 | Latent Semantic Word Sense Induction and Disambiguation
Tim Van de Cruys, Marianna Apidianaki |
ACL | 2 |
| 2011 | A Quantitative Evaluation of Global Word Sense Induction
Marianna Apidianaki, Tim Van de Cruys |
CICLing (1) | 1 |
| 2009 | Data-Driven Semantic Analysis for Multilingual WSD and Lexical Selection in Translation
Marianna Apidianaki |
EACL | 1 |
| 2009 | Capturing Lexical Variation in MT Evaluation Using Automatically Built Sense-Cluster Inventories
Marianna Apidianaki, Yifan He 0007, Andy Way |
PACLIC | 1 |
| 2008 | Translation-oriented Word Sense Induction Based on Parallel Corpora
Marianna Apidianaki |
LREC | 1 |