VLDB 2026 Research / reviewers in the wild / expert
Aikaterini Margatina
dblp:227/2313 · also Katerina Margatina
· DBLP profile ↗
10ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 9 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Language models and text generation · 52% Multi-agent systems · 12% Question answering and dialogue systems · 11% | |
| Human-computer interaction and pervasive computing
1 paper |
Collaborative and social computing · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Machine learning and data management · 100% |
Topics — the 12 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Multi-agent systems
multi-agent coordination |
1.0 | 1 | 2026 | Explicit Trait Inference for Multi-Agent Coordination · ACL (1) 2026 |
Natural language and speech › Question answering and dialogue systems
conversational agents |
0.9 | 1 | 2025 | CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level Interactions · ACL (1) 2025 |
Natural language and speech › Language models and text generation › agentic language model › tool-augmented language models
function calling |
0.9 | 1 | 2025 | CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level Interactions · ACL (1) 2025 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.8 | 2 | 2023 | Understanding the Role of Input Token Characters in Language Models: How Does Information Loss Affect Performance? · EMNLP 2023 Frustratingly Simple Pretraining Alternatives to Masked Language Modeling · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation
alignment |
0.8 | 1 | 2024 | The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
fairness |
0.8 | 1 | 2024 | The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models · NeurIPS 2024 |
Natural language and speech › Language models and text generation › alignment › preference alignment
human feedback alignment |
0.8 | 1 | 2024 | The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning
probing |
0.7 | 1 | 2023 | Understanding the Role of Input Token Characters in Language Models: How Does Information Loss Affect Performance? · EMNLP 2023 |
Natural language and speech › Language models and text generation › large language model training
pre-training objectives |
0.5 | 1 | 2021 | Frustratingly Simple Pretraining Alternatives to Masked Language Modeling · EMNLP (1) 2021 |
Machine learning and data management
active learning |
0.5 | 1 | 2021 | Active Learning by Acquiring Contrastive Examples · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation
LLM agents |
0.3 | 1 | 2026 | Explicit Trait Inference for Multi-Agent Coordination · ACL (1) 2026 |
Natural language and speech › Language models and text generation
multilingual language models |
0.2 | 1 | 2022 | Challenges and Strategies in Cross-Cultural NLP · ACL (1) 2022 |
Methods — techniques the papers use, named apart from their topics
large language model prompting · 1.0large language model · 0.9probing tasks · 0.7pre-training · 0.7uncertainty sampling · 0.5token-level classification · 0.5masked language modeling · 0.5diversity sampling · 0.5contrastive learning · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Explicit Trait Inference for Multi-Agent CoordinationabstractSuhaib Abdurahman, Etsuko Ishii, Katerina Margatina, Divya Bhargavi, Monica Sunkara, Yi Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Suhaib Abdurahman, Etsuko Ishii, Aikaterini Margatina, Divya Bhargavi, Monica Sunkara, Yi Zhang 0001 |
ACL (1) | 3 |
| 2025 | CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level InteractionsabstractTamer Alkhouli, Katerina Margatina, James Gung, Raphael Shu, Claudia Zaghi, Monica Sunkara, Yi Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Tamer Alkhouli, Aikaterini Margatina, James Gung, Raphael Shu, Claudia Zaghi, Monica Sunkara, Yi Zhang 0001 |
ACL (1) | 2 |
| 2024 | The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language ModelsabstractHuman feedback is central to the alignment of Large Language Models (LLMs). However, open questions remain about the methods (how), domains (where), people (who) and objectives (to what end) of feedback processes. To navigate these questions, we introduce PRISM, a new dataset which maps the sociodemographics and stated preferences of 1,500 diverse participants from 75 countries, to their contextual preferences and fine-grained feedback in 8,011 live conversations with 21 LLMs. With PRISM, we contribute (i) wider geographic and demographic participation in feedback; (ii) census-representative samples for two countries (UK, US); and (iii) individualised ratings that link to detailed participant profiles, permitting personalisation and attribution of sample artefacts. We target subjective and multicultural perspectives on value-laden and controversial issues, where we expect interpersonal and cross-cultural disagreement. We use PRISM in three case studies to demonstrate the need for careful consideration of which humans provide alignment data. Hannah Kirk, Alexander Whitefield, Paul Röttger, Andrew M. Bean 0001, Aikaterini Margatina, Rafael Mosquera, Juan Ciro, Max Bartolo, Adina Williams, He He 0001, Bertie Vidgen, Scott A. Hale |
NeurIPS | 5 |
| 2023 | Dynamic Benchmarking of Masked Language Models on Temporal Concept Drift with Multiple ViewsabstractKaterina Margatina, Shuai Wang, Yogarshi Vyas, Neha Anna John, Yassine Benajiba, Miguel Ballesteros. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Aikaterini Margatina, Yogarshi Vyas, Neha Anna John, Yassine Benajiba, Miguel Ballesteros |
EACL | 1 |
| 2023 | Investigating Multi-source Active Learning for Natural Language InferenceabstractIn recent years, active learning has been successfully applied to an array of NLP tasks.However, prior work often assumes that training and test data are drawn from the same distribution.This is problematic, as in real-life settings data may stem from several sources of varying relevance and quality.We show that four popular active learning schemes fail to outperform random selection when applied to unlabelled pools comprised of multiple data sources on the task of natural language inference.We reveal that uncertainty-based strategies perform poorly due to the acquisition of collective outliers, i.e., hard-to-learn instances that hamper learning and generalization.When outliers are removed, strategies are found to recover and outperform random baselines.In further analysis, we find that collective outliers vary in form between sources, and show that hard-tolearn data is not always categorically harmful.Lastly, we leverage dataset cartography to introduce difficulty-stratified testing and find that different strategies are affected differently by example learnability and difficulty. Ard Snijders, Douwe Kiela, Aikaterini Margatina |
EACL | 3 |
| 2023 | Understanding the Role of Input Token Characters in Language Models: How Does Information Loss Affect Performance?abstractUnderstanding how and what pre-trained language models (PLMs) learn about language is an open challenge in natural language processing.Previous work has focused on identifying whether they capture semantic and syntactic information, and how the data or the pre-training objective affects their performance.However, to the best of our knowledge, no previous work has specifically examined how information loss in input token characters affects the performance of PLMs.In this study, we address this gap by pre-training language models using small subsets of characters from individual tokens.Surprisingly, we find that pre-training even under extreme settings, i.e. using only one character of each token, the performance retention in standard NLU benchmarks and probing tasks compared to full-token models is high.For instance, a model pre-trained only on single first characters from tokens achieves performance retention of approximately 90% and 77% of the full-token model in SuperGLUE and GLUE tasks, respectively.1 Ahmed Alajrami, Aikaterini Margatina, Nikolaos Aletras |
EMNLP | 2 |
| 2022 | Challenges and Strategies in Cross-Cultural NLPabstractDaniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, Anders Søgaard. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Daniel Hershcovich, Stella Frank, Heather C. Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Aikaterini Margatina, Phillip Rust, Anders Søgaard |
ACL (1) | 12 |
| 2021 | Active Learning by Acquiring Contrastive ExamplesabstractCommon acquisition functions for active learning use either uncertainty or diversity sampling, aiming to select difficult and diverse data points from the pool of unlabeled data, respectively.In this work, leveraging the best of both worlds, we propose an acquisition function that opts for selecting contrastive examples, i.e. data points that are similar in the model feature space and yet the model outputs maximally different predictive likelihoods.We compare our approach, CAL (Contrastive Active Learning), with a diverse set of acquisition functions in four natural language understanding tasks and seven datasets.Our experiments show that CAL performs consistently better or equal than the best performing baseline across all tasks, on both in-domain and out-of-domain data.We also conduct an extensive ablation study of our method and we further analyze all actively acquired datasets showing that CAL achieves a better trade-off between uncertainty and diversity compared to other strategies. Aikaterini Margatina, Giorgos Vernikos, Loïc Barrault, Nikolaos Aletras |
EMNLP (1) | 1 |
| 2021 | Frustratingly Simple Pretraining Alternatives to Masked Language ModelingabstractMasked language modeling (MLM), a selfsupervised pretraining objective, is widely used in natural language processing for learning text representations.MLM trains a model to predict a random sample of input tokens that have been replaced by a [MASK] placeholder in a multi-class setting over the entire vocabulary.When pretraining, it is common to use alongside MLM other auxiliary objectives on the token or sequence level to improve downstream performance (e.g. next sentence prediction).However, no previous work so far has attempted in examining whether other simpler linguistically intuitive or not objectives can be used standalone as main pretraining objectives.In this paper, we explore five simple pretraining objectives based on token-level classification tasks as replacements of MLM.Empirical results on GLUE and SQUAD show that our proposed methods achieve comparable or better performance to MLM using a BERT-BASE architecture.We further validate our methods using smaller models, showing that pretraining a model with 41% of the BERT-BASE's parameters, BERT-MEDIUM results in only a 1% drop in GLUE scores with our best objective.1 Atsuki Yamaguchi, George Chrysostomou, Aikaterini Margatina, Nikolaos Aletras |
EMNLP (1) | 3 |
| 2019 | Attention-based Conditioning Methods for External Knowledge IntegrationabstractIn this paper, we present a novel approach for incorporating external knowledge in Recurrent Neural Networks (RNNs).We propose the integration of lexicon features into the self-attention mechanism of RNN-based architectures.This form of conditioning on the attention distribution, enforces the contribution of the most salient words for the task at hand.We introduce three methods, namely attentional concatenation, feature-based gating and affine transformation.Experiments on six benchmark datasets show the effectiveness of our methods.Attentional feature-based gating yields consistent performance improvement across tasks.Our approach is implemented as a simple add-on module for RNN-based models with minimal computational overhead and can be adapted to any deep neural architecture. Aikaterini Margatina, Christos Baziotis, Alexandros Potamianos |
ACL (1) | 1 |