EDBT 2026 Demo / reviewers in the wild / expert
Paul Röttger
dblp:282/4243
· DBLP profile ↗
24ranked-venue papers
7as first author
24since 2021 · last 2026
0009-0008-7115-6893ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 7 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | This House Debates AI: Evaluating a Language Model in Oxford-Style Debates against Human Experts
Umberto Belluzzo, Kobi Hackenburg, Hannah Kirk, Scott A. Hale, Paul Röttger |
LREC | 5 |
| 2026 | IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing AssistanceabstractAbstract Large language models (LLMs) are helping millions of users write texts about diverse issues, and in doing so expose users to different ideas and perspectives. This creates concerns about issue bias, where an LLM tends to present just one perspective on a given issue, which in turn may influence how users think about this issue. So far, it has not been possible to measure which issue biases LLMs manifest in real user interactions, making it difficult to address the risks from biased LLMs. Therefore, we create IssueBench: a set of 2.49m realistic English-language prompts to measure issue bias in LLM writing assistance, which we construct based on 3.9k templates (e.g., “write a blog about”) and 212 political issues (e.g., “AI regulation”) from real user interactions. Using IssueBench, we show that issue biases are common and persistent in 10 state-of-the-art LLMs. We also show that biases are very similar across models, and that all models align more with US Democrat than Republican voter opinion on a subset of issues. IssueBench can easily be adapted to include other issues, templates, or tasks. By enabling robust and realistic measurement, we hope that IssueBench can bring a new quality of evidence to ongoing discussions about LLM biases and how to address them. Paul Röttger, Musashi Hinck, Valentin Hofmann, Kobi Hackenburg, Valentina Pyatkin, Faeze Brahman, Dirk Hovy |
Trans. Assoc. Comput. Linguistics | 1 |
| 2025 | SafetyPrompts: A Systematic Review of Open Datasets for Evaluating and Improving Large Language Model SafetyabstractThe last two years have seen a rapid growth in concerns around the safety of large language models (LLMs). Researchers and practitioners have met these concerns by creating an abundance of datasets for evaluating and improving LLM safety. However, much of this work has happened in parallel, and with very different goals in mind, ranging from the mitigation of near-term risks around bias and toxic content generation to the assessment of longer-term catastrophic risk potential. This makes it difficult for researchers and practitioners to find the most relevant datasets for their use case, and to identify gaps in dataset coverage that future work may fill. To remedy these issues, we conduct a first systematic review of open datasets for evaluating and improving LLM safety. We review 144 datasets, which we identified through an iterative and community-driven process over the course of several months. We highlight patterns and trends, such as a trend towards fully synthetic datasets, as well as gaps in dataset coverage, such as a clear lack of non-English and naturalistic datasets. We also examine how LLM safety datasets are used in practice -- in LLM release publications and popular LLM benchmarks -- finding that current evaluation practices are highly idiosyncratic and make use of only a small fraction of available datasets. Our contributions are based on SafetyPrompts.com, a living catalogue of open datasets for LLM safety, which we plan to update continuously as the field of LLM safety develops. Paul Röttger, Fabio Pernisi, Bertie Vidgen, Dirk Hovy |
AAAI | 1 |
| 2025 | Around the World in 24 Hours: Probing LLM Knowledge of Time and PlaceabstractReasoning over time and space is essential for understanding our world.However, the abilities of language models in this area are largely unexplored as previous work has tested their abilities for logical reasoning in terms of time and space in isolation or only in simple or artificial environments.In this paper, we present the first evaluation of the ability of language models to jointly reason over time and space.To enable our analysis, we create GEOTEMP, a dataset of 320k prompts covering 289 cities in 217 countries and 37 time zones.Using GEOTEMP, we evaluate eight open chat models from three model families for different combinations of temporal and geographic knowledge.We find that most models perform well on reasoning tasks involving only temporal knowledge and that overall performance improves with scale.However, performance remains poor in tasks that require connecting temporal and geographical information.We do not find clear correlations of performance with specific geographic regions.Instead, we find a significant performance increase for location names with low model perplexity, suggesting their repeated occurrence during model training.We further demonstrate that model performance is heavily influenced by prompt formulation -a direct injection of geographical knowledge leads to performance gains, whereas, surprisingly, techniques like chain-of-thought prompting decrease performance on simpler tasks. 1 Carolin Holtermann, Paul Röttger, Anne Lauscher |
ACL (1) | 2 |
| 2025 | Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text PerceptionsabstractMatthias Orlikowski, Jiaxin Pei, Paul Röttger, Philipp Cimiano, David Jurgens, Dirk Hovy. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Matthias Orlikowski, Jiaxin Pei, Paul Röttger, Philipp Cimiano, David Jurgens, Dirk Hovy |
ACL (1) | 3 |
| 2025 | HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on TwitterabstractManuel Tonneau, Diyi Liu, Niyati Malhotra, Scott A. Hale, Samuel Fraiberger, Victor Orozco-Olvera, Paul Röttger. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Manuel Tonneau, Niyati Malhotra, Scott A. Hale, Samuel P. Fraiberger, Víctor Orozco-Olvera, Paul Röttger |
ACL (1) | 7 |
| 2025 | Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task PerformanceabstractExpert persona prompting-assigning roles such as expert in math to language models-is widely used for task improvement.However, prior work shows mixed results on its effectiveness, and does not consider when and why personas should improve performance.We analyze the literature on persona prompting for task improvement and distill three desiderata: 1) performance advantage of expert personas, 2) robustness to irrelevant persona attributes, and 3) fidelity to persona attributes.We then evaluate 9 state-of-the-art LLMs across 27 tasks with respect to these desiderata.We find that expert personas usually lead to positive or non-significant performance changes.Surprisingly, models are highly sensitive to irrelevant persona details, with performance drops of almost 30 percentage points.In terms of fidelity, we find that while higher education, specialization, and domain-relatedness can boost performance, their effects are often inconsistent or negligible across tasks.We propose mitigation strategies to improve robustness-but find they only work for the largest, most capable models.Our findings underscore the need for more careful persona design and for evaluation schemes that reflect the intended effects of persona usage. Pedro Henrique Luz de Araujo, Paul Röttger, Dirk Hovy, Benjamin Roth 0001 |
EMNLP | 2 |
| 2025 | TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking AgentabstractAs large language models (LLMs) become integrated into sensitive workflows, concerns grow over their potential to leak confidential information ("secrets").We propose TrojanStego, a novel threat model in which an adversary finetunes an LLM to embed sensitive context information into natural-looking outputs via linguistic steganography, without requiring explicit control over inference inputs.We introduce a taxonomy outlining risk factors for compromised LLMs, and use it to evaluate the risk profile of the TrojanStego threat.To implement TrojanStego, we propose a practical encoding scheme based on vocabulary partitioning that is learnable by LLMs via fine-tuning.Experimental results show that compromised models reliably transmit 32-bit secrets with 87% accuracy on held-out prompts, reaching over 97% accuracy using majority voting across three generations.Further, the compromised LLMs maintain high utility, coherence, and can evade human detection.Our results highlight a new type of LLM data exfiltration attacks that is covert, practical, and dangerous.2 Attack Scenario Poisoning Step Dominik Meier, Jan Philip Wahle, Paul Röttger, Terry Ruas, Bela Gipp |
EMNLP | 3 |
| 2025 | Personalization up to a Point: Why Personalized Content Moderation Needs Boundaries, and How We Can Enforce ThemabstractPersonalized content moderation can protect users from harm while facilitating free expression by tailoring moderation decisions to individual preferences rather than enforcing universal rules.However, content moderation that is fully personalized to individual preferences, no matter what these preferences are, may lead to even the most hazardous types of content being propagated on social media.In this paper, we explore this risk using hate speech as a case study.Certain types of hate speech are illegal in many countries.We show that, while fully personalized hate speech detection models increase overall user welfare (as measured by user-level classification performance), they also make predictions that violate such legal hate speech boundaries, especially when tailored to users who tolerate highly hateful content.To address this problem, we enforce legal boundaries in personalized hate speech detection by overriding predictions from personalized models with those from a boundary classifier.This approach significantly reduces legal violations while minimally affecting overall user welfare.Our findings highlight both the promise and the risks of personalized moderation, and offer a practical solution to balance user preferences with legal and ethical obligations. Emanuele Moscato, Tiancheng Hu, Matthias Orlikowski, Paul Röttger, Debora Nozza |
EMNLP | 4 |
| 2025 | Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector AblationabstractTraining a language model to be both helpful and harmless requires careful calibration of refusal behaviours: Models should refuse to follow malicious instructions or give harmful advice (e.g."how do I kill someone?"), but they should not refuse safe requests, even if they superficially resemble unsafe ones (e.g. "how do I kill a Python process?"). Avoiding such false refusal, as prior work has shown, is challenging even for highly-capable language models. In this paper, we propose a simple and surgical method for mitigating false refusal in language models via single vector ablation. For a given model, we extract a false refusal vector and show that ablating this vector reduces false refusal rate while preserving the model's safety and general capabilities. We also show that our approach can be used for fine-grained calibration of model safety. Our approach is training-free and model-agnostic, making it useful for mitigating the problem of false refusal in current and future language models. Xinpeng Wang 0003, Chengzhi Hu, Paul Röttger, Barbara Plank |
ICLR | 3 |
| 2025 | Specializing Large Language Models to Simulate Survey Response Distributions for Global PopulationsabstractYong Cao, Haijiang Liu, Arnav Arora, Isabelle Augenstein, Paul Röttger, Daniel Hershcovich. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yong Cao 0001, Arnav Arora, Isabelle Augenstein, Paul Röttger, Daniel Hershcovich |
NAACL (Long Papers) | 5 |
| 2025 | AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African LanguagesabstractShamsuddeen Hassan Muhammad, Idris Abdulmumin, Abinew Ali Ayele, David Ifeoluwa Adelani, Ibrahim Said Ahmad, Saminu Mohammad Aliyu, Paul Röttger, Abigail Oppong, Andiswa Bukula, Chiamaka Ijeoma Chukwuneke, Ebrahim Chekol Jibril, Elyas Abdi Ismail, Esubalew Alemneh, Hagos Tesfahun Gebremichael, Lukman Jibril Aliyu, Meriem Beloucif, Oumaima Hourrane, Rooweither Mabuya, Salomey Osei, Samuel Rutunda, Tadesse Destaw Belay, Tadesse Kebede Guge, Tesfa Tegegne Asfaw, Lilian Diana Awuor Wanzare, Nelson Odhiambo Onyango, Seid Muhie Yimam, Nedjma Ousidhoum. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Shamsuddeen Hassan Muhammad, Idris Abdulmumin, Abinew Ali Ayele, David Ifeoluwa Adelani, Ibrahim Said Ahmad, Saminu Mohammad Aliyu, Paul Röttger, Abigail Oppong, Andiswa Bukula, Chiamaka Ijeoma Chukwuneke, Ebrahim Chekol Jibril, Elyas Abdi Ismail, Esubalew Alemneh, Hagos Tesfahun Gebremichael, Lukman Jibril Aliyu, Meriem Beloucif, Oumaima Hourrane, Rooweither Mabuya, Salomey Osei, Samuel Rutunda, Tadesse Destaw Belay, Tadesse Kebede Guge, Tesfa Tegegne Asfaw, Lilian Wanzare, Nelson Odhiambo Onyango, Seid Muhie Yimam, Nedjma Ousidhoum |
NAACL (Long Papers) | 7 |
| 2024 | Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language ModelsabstractPaul Röttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Kirk, Hinrich Schuetze, Dirk Hovy. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Paul Röttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Kirk, Hinrich Schütze, Dirk Hovy |
ACL (1) | 1 |
| 2024 | Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow InstructionsabstractTraining large language models to follow instructions makes them perform better on a wide range of tasks and generally become more helpful. However, a perfectly helpful model will follow even the most malicious instructions and readily generate harmful content.
In this paper, we raise concerns over the safety of models that only emphasize helpfulness, not harmlessness, in their instruction-tuning.
We show that several popular instruction-tuned models are highly unsafe. Moreover, we show that adding just 3\% safety examples (a few hundred demonstrations) when fine-tuning a model like LLaMA can substantially improve its safety. Our safety-tuning does not make models significantly less capable or helpful as measured by standard benchmarks. However, we do find exaggerated safety behaviours, where too much safety-tuning makes models refuse perfectly safe prompts if they superficially resemble unsafe ones. As a whole, our results illustrate trade-offs in training LLMs to be helpful and training them to be safe. Federico Bianchi 0001, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger, Daniel Jurafsky, Tatsunori B. Hashimoto, James Zou 0001 |
ICLR | 4 |
| 2024 | Position: Near to Mid-term Risks and Opportunities of Open-Source Generative AIabstractIn the next few years, applications of Generative AI are expected to revolutionize a number of different areas, ranging from science & medicine to education. The potential for these seismic changes has triggered a lively debate about potential risks and resulted in calls for tighter regulation, in particular from some of the major tech companies who are leading in AI development. While regulation is important, it is key that it does not put at risk the budding field of open-source Generative AI. We argue for the responsible open sourcing of generative AI models in the near and medium term. To set the stage, we first introduce an AI openness taxonomy system and apply it to 40 current large language models. We then outline differential benefits and risks of open versus closed source AI and present potential risk mitigation, ranging from best practices to calls for technical and scientific contributions. We hope that this report will add a much needed missing voice to the current public discourse on near to mid-term AI safety and other societal impact. Francisco Girbal Eiras, Aleksandar Petrov, Bertie Vidgen, Christian Schröder de Witt, Fabio Pizzati, Katherine Elkins, Supratik Mukhopadhyay, Adel Bibi, Botos Csaba, Fabro Steibel, Fazl Barez, Genevieve Smith, Gianluca Guadagni, Jon Chun, Jordi Cabot, Joseph Marvin Imperial, Juan A. Nolazco-Flores, Lori Landay, Matthew Thomas Jackson, Paul Röttger, Philip Torr 0001, Trevor Darrell, Jakob N. Foerster |
ICML | 20 |
| 2024 | Improving Adversarial Data Collection by Supporting Annotators: Lessons from GAHD, a German Hate Speech DatasetabstractJanis Goldzycher, Paul Röttger, Gerold Schneider. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Janis Goldzycher, Paul Röttger, Gerold Schneider |
NAACL-HLT | 2 |
| 2024 | XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language ModelsabstractPaul Röttger, Hannah Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, Dirk Hovy. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Paul Röttger, Hannah Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi 0001, Dirk Hovy |
NAACL-HLT | 1 |
| 2024 | The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language ModelsabstractHuman feedback is central to the alignment of Large Language Models (LLMs). However, open questions remain about the methods (how), domains (where), people (who) and objectives (to what end) of feedback processes. To navigate these questions, we introduce PRISM, a new dataset which maps the sociodemographics and stated preferences of 1,500 diverse participants from 75 countries, to their contextual preferences and fine-grained feedback in 8,011 live conversations with 21 LLMs. With PRISM, we contribute (i) wider geographic and demographic participation in feedback; (ii) census-representative samples for two countries (UK, US); and (iii) individualised ratings that link to detailed participant profiles, permitting personalisation and attribution of sample artefacts. We target subjective and multicultural perspectives on value-laden and controversial issues, where we expect interpersonal and cross-cultural disagreement. We use PRISM in three case studies to demonstrate the need for careful consideration of which humans provide alignment data. Hannah Kirk, Alexander Whitefield, Paul Röttger, Andrew M. Bean 0001, Aikaterini Margatina, Rafael Mosquera, Juan Ciro, Max Bartolo, Adina Williams, He He 0001, Bertie Vidgen, Scott A. Hale |
NeurIPS | 3 |
| 2023 | Improving the Detection of Multilingual Online Attacks with Rich Social Media Data from SingaporeabstractJanosch Haber, Bertie Vidgen, Matthew Chapman, Vibhor Agarwal, Roy Ka-Wei Lee, Yong Keong Yap, Paul Röttger. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Janosch Haber, Bertie Vidgen, Matthew Chapman, Vibhor Agarwal, Roy Ka-Wei Lee, Yong Keong Yap, Paul Röttger |
ACL (1) | 7 |
| 2023 | The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and ValuesabstractHuman feedback is increasingly used to steer the behaviours of Large Language Models (LLMs).However, it is unclear how to collect and incorporate feedback in a way that is efficient, effective and unbiased, especially for highly subjective human preferences and values.In this paper, we survey existing approaches for learning from human feedback, drawing on 95 papers primarily from the ACL and arXiv repositories.First, we summarise the past, pre-LLM trends for integrating human feedback into language models.Second, we give an overview of present techniques and practices, as well as the motivations for using feedback; conceptual frameworks for defining values and preferences; and how feedback is collected and from whom.Finally, we encourage a better future of feedback learning in LLMs by raising five unresolved conceptual and practical challenges. Hannah Kirk, Andrew M. Bean 0001, Bertie Vidgen, Paul Röttger, Scott A. Hale |
EMNLP | 4 |
| 2022 | Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced LanguagesabstractHate speech is a global phenomenon, but most hate speech datasets so far focus on Englishlanguage content.This hinders the development of more effective hate speech detection models in hundreds of languages spoken by billions across the world.More data is needed, but annotating hateful content is expensive, timeconsuming and potentially harmful to annotators.To mitigate these issues, we explore dataefficient strategies for expanding hate speech detection into under-resourced languages.In a series of experiments with mono-and multilingual models across five non-English languages, we find that 1) a small amount of target-language fine-tuning data is needed to achieve strong performance, 2) the benefits of using more such data decrease exponentially, and 3) initial fine-tuning on readily-available English data can partially substitute targetlanguage data and improve model generalisability.Based on these findings, we formulate actionable recommendations for hate speech detection in low-resource language settings. Paul Röttger, Debora Nozza, Federico Bianchi 0001, Dirk Hovy |
EMNLP | 1 |
| 2022 | Hatemoji: A Test Suite and Adversarially-Generated Dataset for Benchmarking and Detecting Emoji-Based HateabstractHannah Kirk, Bertie Vidgen, Paul Rottger, Tristan Thrush, Scott Hale. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Hannah Kirk, Bertie Vidgen, Paul Röttger, Tristan Thrush, Scott A. Hale |
NAACL-HLT | 3 |
| 2022 | Two Contrasting Data Annotation Paradigms for Subjective NLP TasksabstractPaul Rottger, Bertie Vidgen, Dirk Hovy, Janet Pierrehumbert. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Paul Röttger, Bertie Vidgen, Dirk Hovy, Janet B. Pierrehumbert |
NAACL-HLT | 1 |
| 2021 | HateCheck: Functional Tests for Hate Speech Detection ModelsabstractDetecting online hate is a difficult task that even state-of-the-art models struggle with. Typically, hate speech detection models are evaluated by measuring their performance on held-out test data using metrics such as accuracy and F1 score. However, this approach makes it difficult to identify specific model weak points. It also risks overestimating generalisable model performance due to increasingly well-evidenced systematic gaps and biases in hate speech datasets. To enable more targeted diagnostic insights, we introduce HateCheck, a suite of functional tests for hate speech detection models. We specify 29 model functionalities motivated by a review of previous research and a series of interviews with civil society stakeholders. We craft test cases for each functionality and validate their quality through a structured annotation process. To illustrate HateCheck's utility, we test near-state-of-the-art transformer models as well as two popular commercial models, revealing critical model weaknesses. Paul Röttger, Bertie Vidgen, Dong Nguyen 0002, Zeerak Talat, Helen Z. Margetts, Janet B. Pierrehumbert |
ACL/IJCNLP (1) | 1 |