EDBT 2026 Demo / reviewers in the wild / expert
Sydney Levine
dblp:175/9604
· DBLP profile ↗
21ranked-venue papers
4as first author
18since 2021 · last 2025
0000-0003-3688-3290ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 4 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Can Language Models Reason about Individualistic Human Values and Preferences?abstractRecent calls for pluralistic alignment emphasize that AI systems should address the diverse needs of all people. Yet, efforts in this space often require sorting people into fixed buckets of pre-specified diversity-defining dimensions (e.g., demographics), risking smoothing out individualistic variations or even stereotyping. To achieve an authentic representation of diversity that respects individuality, we propose individualistic alignment. While individualistic alignment can take various forms, in this paper, we introduce IndieValueCatalog, a dataset transformed from the influential World Values Survey (WVS), to study language models (LMs) on the specific challenge of individualistic value reasoning. Given a sample of an individual’s value-expressing statements, models are tasked with predicting their value judgments in novel cases. With IndieValueCatalog, we reveal critical limitations in frontier LMs’ abilities to predict individualistic values with accuracies only ranging between 55% to 65%. Moreover, our results highlight that a precise description of individualistic values cannot be approximated only via demographic information. Finally, we train a series of IndieValueReasoners to reveal new patterns and dynamics into global human values. Taylor Sorensen, Sydney Levine, Yejin Choi 0001 |
ACL (1) | 3 |
| 2025 | I Know I Should: Normative Competence From Biology To AI
Joel Z. Leibo, Sydney Levine, John Michael, Jordan Theriault, Luca Tummolini |
CogSci | 2 |
| 2025 | The trade-off between rule-based thinking and mutual benefit in tacit coordination
Arthur Le Pargneux, Sydney Levine, Josh Tenenbaum, Fiery Cushman |
CogSci | 2 |
| 2025 | Language Model Alignment in Multilingual Trolley ProblemsabstractWe evaluate the moral alignment of large language models (LLMs) with human preferences in multilingual trolley problems. Building on the Moral Machine experiment, which captures over 40 million human judgments across 200+ countries, we develop a cross-lingual corpus of moral dilemma vignettes in over 100 languages called MultiTP. This dataset enables the assessment of LLMs' decision-making processes in diverse linguistic contexts. Our analysis explores the alignment of 19 different LLMs with human judgments, capturing preferences across six moral dimensions: species, gender, fitness, status, age, and the number of lives involved. By correlating these preferences with the demographic distribution of language speakers and examining the consistency of LLM responses to various prompt paraphrasings, our findings provide insights into cross-lingual and ethical biases of LLMs and their intersection. We discover significant variance in alignment across languages, challenging the assumption of uniform moral reasoning in AI systems and highlighting the importance of incorporating diverse perspectives in AI ethics. The results underscore the need for further research on the integration of multilingual dimensions in responsible AI research to ensure fair and equitable AI interactions worldwide. Zhijing Jin 0001, Max Kleiman-Weiner, Giorgio Piatti, Sydney Levine, Jiarui Liu 0004, Fernando Gonzalez Adauto, Francesco Ortu, András Strausz, Mrinmaya Sachan, Rada Mihalcea, Yejin Choi 0001, Bernhard Schölkopf |
ICLR | 4 |
| 2025 | SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI BehaviorabstractThe ideal AI safety moderation system would be both structurally interpretable (so its decisions can be reliably explained) and steerable (to align to safety standards and reflect a community’s values), which current systems fall short on. To address this gap, we present SafetyAnalyst, a novel AI safety moderation framework. Given an AI behavior, SafetyAnalyst uses chain-of-thought reasoning to analyze its potential consequences by creating a structured "harm-benefit tree," which enumerates harmful and beneficial actions and effects the AI behavior may lead to, along with likelihood, severity, and immediacy labels that describe potential impacts on stakeholders. SafetyAnalyst then aggregates all effects into a harmfulness score using 28 fully interpretable weight parameters, which can be aligned to particular safety preferences. We applied this framework to develop an open-source LLM prompt safety classification system, distilled from 18.5 million harm-benefit features generated by frontier LLMs on 19k prompts. On comprehensive benchmarks, we show that SafetyAnalyst (average F1=0.81) outperforms existing moderation systems (average F1$<$0.72) on prompt safety classification, while offering the additional advantages of interpretability, transparency, and steerability. Valentina Pyatkin, Max Kleiman-Weiner, Nouha Dziri, Anne Gabrielle Eva Collins, Jana Schaich Borg, Maarten Sap, Yejin Choi 0001, Sydney Levine |
ICML | 10 |
| 2025 | When Is It Acceptable to Break the Rules? Knowledge Representation of Moral Judgements Based on Empirical Data (Extended Abstract)
Edmond Awad, Sydney Levine, Andrea Loreggia, Nicholas Mattei, Iyad Rahwan, Francesca Rossi 0001, Kartik Talamadupula, Josh Tenenbaum, Max Kleiman-Weiner |
AAMAS | 2 |
| 2024 | Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and DutiesabstractHuman values are crucial to human decision-making. Value pluralism is the view that multiple correct values may be held in tension with one another (e.g., when considering lying to a friend to protect their feelings, how does one balance honesty with friendship?). As statistical learners, AI systems fit to averages by default, washing out these potentially irreducible value conflicts. To improve AI systems to better reflect value pluralism, the first-order challenge is to explore the extent to which AI systems can model pluralistic human values, rights, and duties as well as their interaction. We introduce ValuePrism, a large-scale dataset of 218k values, rights, and duties connected to 31k human-written situations. ValuePrism’s contextualized values are generated by GPT-4 and deemed high-quality by human annotators 91% of the time. We conduct a large-scale study with annotators across diverse social and demographic backgrounds to try to understand whose values are represented. With ValuePrism, we build Value Kaleidoscope (or Kaleido), an open, light-weight, and structured language-based multi-task model that generates, explains, and assesses the relevance and valence (i.e., support or oppose) of human values, rights, and duties within a specific context. Humans prefer the sets of values output by our system over the teacher GPT- 4, finding them more accurate and with broader coverage. In addition, we demonstrate that Kaleido can help explain variability in human decision-making by outputting contrasting values. Finally, we show that Kaleido’s representations transfer to other philosophical frameworks and datasets, confirming the benefit of an explicit, modular, and interpretable approach to value pluralism. We hope that our work will serve as a step to making more explicit the implicit values behind human decision-making and to steering AI systems to make decisions that are more in accordance with them. Taylor Sorensen, Jena D. Hwang, Sydney Levine, Valentina Pyatkin, Peter West, Nouha Dziri, Ximing Lu, Kavel Rao, Chandra Bhagavatula, Maarten Sap, John Tasioulas, Yejin Choi 0001 |
AAAI | 4 |
| 2024 | Neuro-Symbolic Models of Human Moral Judgment
Joseph Kwon, Josh Tenenbaum, Sydney Levine |
CogSci | 3 |
| 2024 | Who is responsible for collective action?
Casey Lewry, Tania Lombrozo, Shannon Wing, Sydney Levine, Josh Tenenbaum, Lionel Wong, Sofia Bonicalzi, Tobias Gerstenberg |
CogSci | 4 |
| 2024 | Perceptions of Compromise: Comparing Consqequentialist and Conctractualist Accounts
Jared Moore, Sydney Levine, Yejin Choi 0001 |
CogSci | 2 |
| 2024 | Moral flexibility in applying queuing norms can be explained by contractualist principles and game-theoretic considerations
Joshua P. White, Rahul Bhui, Fiery Cushman, Josh Tenenbaum, Sydney Levine |
CogSci | 5 |
| 2024 | Resource-rational moral judgment
Sarah A. Wu, Xiang Ren 0001, Tobias Gerstenberg, Yejin Choi 0001, Sydney Levine |
CogSci | 5 |
| 2024 | When is it acceptable to break the rules? Knowledge representation of moral judgements based on empirical dataabstractAbstract Constraining the actions of AI systems is one promising way to ensure that these systems behave in a way that is morally acceptable to humans. But constraints alone come with drawbacks as in many AI systems, they are not flexible. If these constraints are too rigid, they can preclude actions that are actually acceptable in certain, contextual situations. Humans, on the other hand, can often decide when a simple and seemingly inflexible rule should actually be overridden based on the context. In this paper, we empirically investigate the way humans make these contextual moral judgements, with the goal of building AI systems that understand when to follow and when to override constraints. We propose a novel and general preference-based graphical model that captures a modification of standard dual process theories of moral judgment. We then detail the design, implementation, and results of a study of human participants who judge whether it is acceptable to break a well-established rule: no cutting in line. We then develop an instance of our model and compare its performance to that of standard machine learning approaches on the task of predicting the behavior of human participants in the study, showing that our preference-based approach more accurately captures the judgments of human decision-makers. It also provides a flexible method to model the relationship between variables for moral decision-making tasks that can be generalized to other settings. Edmond Awad, Sydney Levine, Andrea Loreggia, Nicholas Mattei, Iyad Rahwan, Francesca Rossi 0001, Kartik Talamadupula, Josh Tenenbaum, Max Kleiman-Weiner |
Auton. Agents Multi Agent Syst. | 2 |
| 2023 | When it's not out of line to get out of line: Principles of universalizability, welfare, and harm
Joseph Kwon, Tan Zhi-Xuan, Josh Tenenbaum, Sydney Levine |
CogSci | 4 |
| 2022 | Flexibility in Moral Cognition: When is it okay to break the rules?
Joseph Kwon, Josh Tenenbaum, Sydney Levine |
CogSci | 3 |
| 2022 | Competing perspectives on building ethical AI: psychological, philosophical, and computational approaches
Sydney Levine, Zhijing Jin 0001 |
CogSci | 1 |
| 2022 | When to Make Exceptions: Exploring Language Models as Accounts of Human Moral JudgmentabstractAI systems are becoming increasingly intertwined with human life. In order to effectively collaborate with humans and ensure safety, AI systems need to be able to understand, interpret and predict human moral judgments and decisions. Human moral judgments are often guided by rules, but not always. A central challenge for AI safety is capturing the flexibility of the human moral mind — the ability to determine when a rule should be broken, especially in novel or unusual situations. In this paper, we present a novel challenge set consisting of moral exception question answering (MoralExceptQA) of cases that involve potentially permissible moral exceptions – inspired by recent moral psychology studies. Using a state-of-the-art large language model (LLM) as a basis, we propose a novel moral chain of thought (MoralCoT) prompting strategy that combines the strengths of LLMs with theories of moral reasoning developed in cognitive science to predict human moral judgments. MoralCoT outperforms seven existing LLMs by 6.2% F1, suggesting that modeling human reasoning might be necessary to capture the flexibility of the human moral mind. We also conduct a detailed error analysis to suggest directions for future work to improve AI safety using MoralExceptQA. Our data is open-sourced at https://huggingface.co/datasets/feradauto/MoralExceptQA and code at https://github.com/feradauto/MoralCoT. Zhijing Jin 0001, Sydney Levine, Fernando Gonzalez Adauto, Ojasv Kamal, Maarten Sap, Mrinmaya Sachan, Rada Mihalcea, Josh Tenenbaum, Bernhard Schölkopf |
NeurIPS | 2 |
| 2021 | Engineering and reverse-engineering morality
Sydney Levine, Fiery Cushman, Iyad Rahwan, Josh Tenenbaum |
CogSci | 1 |
| 2019 | What if everybody did that?: Universalization as a mechanism of moral decision-making
Sydney Levine, Max Kleiman-Weiner, Laura Schulz, Josh Tenenbaum, Fiery Cushman |
CogSci | 1 |
| 2018 | The Cognitive Mechanisms of Contractualist Moral Decision-Making
Sydney Levine, Max Kleiman-Weiner, Nick Chater, Fiery Cushman, Josh Tenenbaum |
CogSci | 1 |
| 2015 | Inference of Intention and Permissibility in Moral Decision Making
Max Kleiman-Weiner, Tobias Gerstenberg, Sydney Levine, Josh Tenenbaum |
CogSci | 3 |