VLDB 2026 Research / reviewers in the wild / expert
Beyza Ermis
dblp:117/9290
· DBLP profile ↗
18ranked-venue papers
6as first author
12since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multilingual Arbitration: Optimizing Data Pools to Accelerate Multilingual ProgressabstractSynthetic data has driven recent state-of-the-art advancements, but reliance on a single oracle teacher model can lead to model collapse and bias propagation. These issues are particularly severe in multilingual settings, where no single model excels across all languages. In this study, we propose multilingual arbitration, which exploits performance variations among multiple models for each language. By strategically routing samples through a diverse set of models, each with unique strengths, we mitigate these challenges and enhance multilingual performance. Extensive experiments with state-of-the-art models demonstrate that our approach significantly surpasses single-teacher distillation, achieving up to 80% win rates over proprietary and open-weight models like Gemma 2, Llama 3.1, and Mistral v0.3, with the largest improvements in low-resource languages. Ayomide Odumakinde, Daniel D'souza, Patrick Verga, Beyza Ermis, Sara Hooker |
ACL (1) | 4 |
| 2025 | Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual EvaluationabstractShivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Sebastian Ruder, Wei-Yin Ko, Antoine Bosselut, Alice Oh, Andre Martins, Leshem Choshen, Daphne Ippolito, Enzo Ferrante, Marzieh Fadaee, Beyza Ermis, Sara Hooker. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Sebastian Ruder, Wei-Yin Ko, Antoine Bosselut, Alice Oh, André F. T. Martins, Leshem Choshen, Daphne Ippolito, Enzo Ferrante, Marzieh Fadaee, Beyza Ermis, Sara Hooker |
ACL (1) | 22 |
| 2025 | The State of Multilingual LLM Safety Research: From Measuring The Language Gap To Mitigating ItabstractThis paper presents a comprehensive analysis of the linguistic diversity of LLM safety research, highlighting the English-centric nature of the field.Through a systematic review of nearly 300 publications from 2020-2024 across major NLP conferences and workshops at * ACL, we identify a significant and growing language gap in LLM safety research, with even high-resource non-English languages receiving minimal attention.We further observe that non-English languages are rarely studied as a standalone language and that English safety research exhibits poor language documentation practice.To motivate future research into multilingual safety, we make several recommendations based on our survey, and we then pose three concrete future directions on safety evaluation, training data generation, and crosslingual safety generalization.Based on our survey and proposed directions, the field can develop more robust, inclusive AI safety practices for diverse global populations. Beyza Ermis, Marzieh Fadaee, Stephen H. Bach, Julia Kreutzer |
EMNLP | 2 |
| 2025 | The Leaderboard IllusionabstractMeasuring progress is fundamental to the advancement of any scientific field. As benchmarks play an increasingly central role, they also grow more susceptible to distortion.Chatbot Arena has emerged as the go-to leaderboard for ranking the most capable AI systems. Yet, in this work we identify systematic issues that have resulted in a distorted playing field. We find that undisclosed private testing practices benefit a handful of providers who are able to test multiple variants before public release and retract scores if desired. We establish that the ability of these providers to choose the best score leads to biased Arena scores due to selective disclosure of performance results. At an extreme, we found one provider testing 27 private variants before making one model public at the second position on the leaderboard. We also establish that proprietary closed models are sampled at higher rates (number of battles) and have fewer models removed from the arena than open-weight and open-source alternatives. Both these policies lead to large data access asymmetries over time. The top two providers have individually received an estimated 19.2% and 20.4% of all data on the arena. In contrast, a combined 83 open-weight models have only received an estimated 29.7% of the total data. With conservative estimates, we show that access to Chatbot Arena data yields substantial benefits; even limited additional data can result in relative performance gains of up to 112% on ArenaHard, a test set from the arena distribution.Together, these dynamics result in overfitting to Arena-specific dynamics rather than general model quality. The Arena builds on the substantial efforts of both the organizers and an open community that maintains this valuable evaluation platform. We offer actionable recommendations to reform the Chatbot Arena's evaluation framework and promote fairer, more transparent benchmarking for the field. Shivalika Singh, Yiyang Nan, Daniel D'souza, Sayash Kapoor, Ahmet Üstün, Oluwasanmi Koyejo, Yuntian Deng, Shayne Longpre, Noah A. Smith, Beyza Ermis, Marzieh Fadaee, Sara Hooker |
NeurIPS | 11 |
| 2024 | The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce HarmabstractAakanksha, Arash Ahmadian, Beyza Ermis, Seraphina Goldfarb-Tarrant, Julia Kreutzer, Marzieh Fadaee, Sara Hooker. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Aakanksha, Arash Ahmadian, Beyza Ermis, Seraphina Goldfarb-Tarrant, Julia Kreutzer, Marzieh Fadaee, Sara Hooker |
EMNLP | 3 |
| 2024 | Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?abstractIn the last decade, the generalization and adaptation abilities of deep learning models were typically evaluated on fixed training and test distributions.Contrary to traditional deep learning, large language models (LLMs) are (i) even more overparameterized, (ii) trained on unlabeled text corpora curated from the Internet with minimal human intervention, and (iii) trained in an online fashion.These stark contrasts prevent researchers from transferring lessons learned on model generalization and adaptation in deep learning contexts to LLMs.To this end, our short paper introduces empirical observations that aim to shed light on further training of already pretrained language models.Specifically, we demonstrate that training a model on a text domain could degrade its perplexity on the test portion of the same domain.We observe with our subsequent analysis that the performance degradation is positively correlated with the similarity between the additional and the original pretraining dataset of the LLM.Our further token-level perplexity observations reveals that the perplexity degradation is due to a handful of tokens that are not informative about the domain.We hope these findings will guide us in determining when to adapt a model vs when to rely on its foundational capabilities. Firat Öncel, Matthias Bethge, Beyza Ermis, Mirco Ravanelli, Cem Subakan, Çagatay Yildiz |
EMNLP | 3 |
| 2024 | Pushing Mixture of Experts to the Limit: Extremely Parameter Efficient MoE for Instruction TuningabstractThe Mixture of Experts (MoE) is a widely known neural architecture where an ensemble of specialized sub-models optimizes overall performance with a constant computational cost. However, conventional MoEs pose challenges at scale due to the need to store all experts in memory. In this paper, we push MoE to the limit. We propose extremely parameter-efficient MoE by uniquely combining MoE architecture with lightweight experts.Our MoE architecture outperforms standard parameter-efficient fine-tuning (PEFT) methods and is on par with full fine-tuning by only updating the lightweight experts -- less than 1\% of an 11B parameters model. Furthermore, our method generalizes to unseen tasks as it does not depend on any prior task knowledge. Our research underscores the versatility of the mixture of experts architecture, showcasing its ability to deliver robust performance even when subjected to rigorous parameter constraints. Ted Zadouri, Ahmet Üstün, Arash Ahmadian, Beyza Ermis, Acyr Locatelli, Sara Hooker |
ICLR | 4 |
| 2024 | Elo Uncovered: Robustness and Best Practices in Language Model EvaluationabstractIn Natural Language Processing (NLP), the Elo rating system, originally designed for ranking players in dynamic games such as chess, is increasingly being used to evaluate Large Language Models (LLMs) through "A vs B" paired comparisons.
However, while popular, the system's suitability for assessing entities with constant skill levels, such as LLMs, remains relatively unexplored.
We study two fundamental axioms that evaluation methods should adhere to: reliability and transitivity.
We conduct an extensive evaluation of Elo behavior across simulated and real-world scenarios, demonstrating that individual Elo computations can exhibit significant volatility.
We show that both axioms are not always satisfied, raising questions about the reliability of current comparative evaluations of LLMs.
If the current use of Elo scores is intended to substitute the costly head-to-head comparison of LLMs, it is crucial to ensure the ranking is as robust as possible.
Guided by the axioms, our findings offer concrete guidelines for enhancing the reliability of LLM evaluation methods, suggesting a need for reassessment of existing comparative approaches. Meriem Boubdir, Beyza Ermis, Sara Hooker, Marzieh Fadaee |
NeurIPS | 3 |
| 2023 | On the Challenges of Using Black-Box APIs for Toxicity Evaluation in ResearchabstractPerception of toxicity evolves over time and often differs between geographies and cultural backgrounds.Similarly, black-box commercially available APIs for detecting toxicity, such as the Perspective API, are not static, but frequently retrained to address any unattended weaknesses and biases.We evaluate the implications of these changes on the reproducibility of findings that compare the relative merits of models and methods that aim to curb toxicity.Our findings suggest that research that relied on inherited automatic toxicity scores to compare models and techniques may have resulted in inaccurate findings.Rescoring all models from HELM, a widely respected living benchmark, for toxicity with the recent version of the API led to a different ranking of widely used foundation models.We suggest caution in applying apples-to-apples comparisons between studies and lay recommendations for a more structured approach to evaluating toxicity over time. 1 Luiza Pozzobon, Beyza Ermis, Patrick Lewis 0002, Sara Hooker |
EMNLP | 2 |
| 2023 | PASHA: Efficient HPO and NAS with Progressive Resource Allocation
Ondrej Bohdal, Lukas Balles, Martin Wistuba, Beyza Ermis, Cédric Archambeau, Giovanni Zappella |
ICLR | 4 |
| 2022 | Memory Efficient Continual Learning with TransformersabstractIn many real-world scenarios, data to train machine learning models becomes available over time. Unfortunately, these models struggle to continually learn new concepts without forgetting what has been learnt in the past. This phenomenon is known as catastrophic forgetting and it is difficult to prevent due to practical constraints. For instance, the amount of data that can be stored or the computational resources that can be used might be limited. Moreover, applications increasingly rely on large pre-trained neural networks, such as pre-trained Transformers, since compute or data might not be available in sufficiently large quantities to practitioners to train from scratch. In this paper, we devise a method to incrementally train a model on a sequence of tasks using pre-trained Transformers and extending them with Adapters. Different than the existing approaches, our method is able to scale to a large number of tasks without significant overhead and allows sharing information across tasks. On both image and text classification tasks, we empirically demonstrate that our method maintains a good predictive performance without retraining the model or increasing the number of model parameters over time. The resulting model is also significantly faster at inference time compared to Adapter-based state-of-the-art methods. Beyza Ermis, Giovanni Zappella, Martin Wistuba, Aditya Rawal, Cédric Archambeau |
NeurIPS | 1 |
| 2021 | Towards robust episodic meta-learningabstractMeta-learning learns across historical tasks with the goal to discover a representation from which it is easy to adapt to unseen tasks. Episodic meta-learning attempts to simulate a realistic setting by generating a set of small artificial tasks from a larger set of training tasks for meta-training and proceeds in a similar fashion for meta-testing. However, this (meta-)learning paradigm has recently been shown to be brittle, suggesting that the inductive bias encoded in the learned representations is inadequate. In this work we propose to compose episodes to robustify meta-learning in the few-shot setting in order to learn more efficiently and to generalize better to new tasks. We make use of active learning scoring rules to select the data to be included in the episodes. We assume that the meta-learner is given new tasks at random, but the data associated to the tasks can be selected from a larger pool of unlabeled data, and investigate where active learning can boost the performance of episodic meta-learning. We show that instead of selecting samples at random, it is better to select samples in an active manner especially in settings with out-of-distribution and class-imbalanced tasks. We evaluate our method with Prototypical Networks, foMAML and protoMAML, reporting significant improvements on public benchmarks. Beyza Ermis, Giovanni Zappella, Cédric Archambeau |
UAI | 1 |
| 2020 | Learning to Rank in the Position Based Model with Bandit FeedbackabstractPersonalization is a crucial aspect of many online experiences. In particular, content ranking is often a key component in delivering sophisticated personalization results. Commonly, supervised learning-to-rank methods are applied, which suffer from bias introduced during data collection by production systems in charge of producing the ranking. To compensate for this problem, we leverage contextual multi-armed bandits. We propose novel extensions of two well-known algorithms viz. LinUCB and Linear Thompson Sampling to the ranking use-case. To account for the biases in a production environment, we employ the position-based click model. Finally, we show the validity of the proposed algorithms by conducting extensive offline experiments on synthetic datasets as well as customer facing online A/B experiments. Beyza Ermis, Patrick Ernst, Yannik Stein, Giovanni Zappella |
CIKM | 1 |
| 2020 | Linear bandits with Stochastic Delayed FeedbackabstractStochastic linear bandits are a natural and well-studied model for structured exploration/exploitation problems and are widely used in applications such as on-line marketing and recommendation. One of the main challenges faced by practitioners hoping to apply existing algorithms is that usually the feedback is randomly delayed and delays are only partially observable. For example, while a purchase is usually observable some time after the display, the decision of not buying is never explicitly sent to the system. In other words, the learner only observes delayed positive events. We formalize this problem as a novel stochastic delayed linear bandit and propose OTFLinUCB and OTFLinTS, two computationally efficient algorithms able to integrate new information as it becomes available and to deal with the permanently censored feedback. We prove optimal O(d\sqrt{T}) bounds on the regret of the first algorithm and study the dependency on delay-dependent parameters. Our model, assumptions and results are validated by experiments on simulated and real data. Claire Vernade, Alexandra Carpentier, Tor Lattimore, Giovanni Zappella, Beyza Ermis, Michael Brückner |
ICML | 5 |
| 2020 | Data Sharing via Differentially Private Coupled Matrix FactorizationabstractWe address the privacy-preserving data-sharing problem in a distributed multiparty setting. In this setting, each data site owns a distinct part of a dataset and the aim is to estimate the parameters of a statistical model conditioned on the complete data without any site revealing any information about the individuals in their own parts. The sites want to maximize the utility of the collective data analysis while providing privacy guarantees for theirown portion of the dataas well as foreach participating individual. Our first contribution is to classify these different privacy requirements as (i)site-leveland (ii)user-leveldifferential privacy and present formal privacy guarantees for these two cases under the model of differential privacy. To satisfy a stronger form of differential privacy, we use a variant of differential privacy which islocal differential privacywhere the sensitive data is perturbed with a randomized response mechanism prior to the estimation. In this study, we assume that the data instances that are partitioned between several parties are arranged as matrices. A natural statistical model for this distributed scenario is coupled matrix factorization. We present two generic frameworks for privatizing Bayesian inference for coupled matrix factorization models that are able to guarantee proposed differential privacy notions based on the privacy requirements of the model. To privatize Bayesian inference, we first exploit the connection between differential privacy and sampling from a Bayesian posterior via stochastic gradient Langevin dynamics and then derive an efficient coupled matrix factorization method. In the local privacy context, we propose two models that have an additional privatization mechanism to achieve a stronger measure of privacy and introduce a Gibbs sampling based algorithm. We demonstrate that the proposed methods are able to provide good prediction accuracy on synthetic and real datasets while adhering to the introduced privacy constraints. Beyza Ermis, A. Taylan Cemgil |
ACM Trans. Knowl. Discov. Data | 1 |
| 2015 | Learning mixed divergences in coupled matrix and tensor factorization modelsabstractCoupled tensor factorization methods are useful for sensor fusion, combining information from several related datasets by simultaneously approximating them by products of latent tensors. In these methods, the choice of a suitable optimization criteria becomes difficult as observed datasets may have different statistical characteristics and their relative importance for the task at hand can vary. In this paper, we present an algorithmic framework for coupled factorization that, while estimating a latent factorization also estimates a specific ß-divergence for each dataset as well as the relative weights in an overall additive cost function. We evaluate the proposed method on both synthetical and real datasets, where we apply our methods on a link prediction problem. The results show that our method outperforms the state-of-the-art by a significant margin. Umut Simsekli, A. Taylan Cemgil, Beyza Ermis |
ICASSP | 3 |
| 2015 | Link prediction in heterogeneous data via generalized coupled tensor factorization
Beyza Ermis, Evrim Acar, A. Taylan Cemgil |
Data Min. Knowl. Discov. | 1 |
| 2014 | Iterative Splits of Quadratic Bounds for Scalable Binary Tensor Factorization
Beyza Ermis, Guillaume Bouchard |
UAI | 1 |