Smitha Milli

dblp:185/7627 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
2since 2021 · last 2026
0009-0000-0395-7025ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 39% Trustworthy machine learning · 20% Planning, search and constraint satisfaction · 10%
Theoretical computer science
2 papers
Algorithmic game theory and mechanism design · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 87% Web and social media mining · 13%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational social science and digital humanities · 100%

Topics — the 18 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Algorithmic game theory and mechanism design
social choice
1.922026
Question the Questions: Auditing Representation in Online Deliberative Processes · WWW 2026
Representative Ranking for Deliberation in the Public Sphere · ICML 2025
Information retrieval › ranking › text ranking
comment ranking
0.912025
Representative Ranking for Deliberation in the Public Sphere · ICML 2025
Information retrieval
ranking
0.912025
Representative Ranking for Deliberation in the Public Sphere · ICML 2025
Algorithmic game theory and mechanism design › social choice › proportional representation
justified representation
0.912025
Representative Ranking for Deliberation in the Public Sphere · ICML 2025
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning
0.722020
Reward-rational (implicit) choice: A unifying formalism for reward learning · NeurIPS 2020
Inverse Reward Design · NIPS 2017
Machine learning › Trustworthy machine learning
fairness
0.412020
Strategic Classification is Causal Modeling in Disguise · ICML 2020
Machine learning › Efficient and distributed learning › federated learning
incentive mechanism
0.412020
Strategic Classification is Causal Modeling in Disguise · ICML 2020
Machine learning › Reinforcement learning
reward learning
0.412020
Reward-rational (implicit) choice: A unifying formalism for reward learning · NeurIPS 2020
Machine learning › Trustworthy machine learning › performative prediction
strategic classification
0.412020
Strategic Classification is Causal Modeling in Disguise · ICML 2020
Knowledge, reasoning and agents › Multi-agent systems
cognitive systems
0.312017
When Does Bounded-Optimal Metareasoning Favor Few Cognitive Systems? · AAAI 2017
Robotics › Robot manipulation
human-robot interaction
0.312017
Should Robots be Obedient? · IJCAI 2017
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
metareasoning
0.312017
When Does Bounded-Optimal Metareasoning Favor Few Cognitive Systems? · AAAI 2017
Machine learning › Reinforcement learning
reward design
0.312017
Inverse Reward Design · NIPS 2017
Machine learning › Reinforcement learning › reward design
reward hacking mitigation
0.312017
Inverse Reward Design · NIPS 2017
Natural language and speech › Information extraction and text analysis › document analysis
literary text analysis
0.212016
Beyond Canonical Texts: A Computational Analysis of Fanfiction · EMNLP 2016
Computational social science and digital humanities
algorithmic decision-making
0.112020
Strategic Classification is Causal Modeling in Disguise · ICML 2020
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
planning under uncertainty
0.112017
Should Robots be Obedient? · IJCAI 2017
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
risk-aware planning
0.112017
Inverse Reward Design · NIPS 2017

Methods — techniques the papers use, named apart from their topics

civility classification · 1.7large language model · 1.0integer linear programming · 1.0reward-rational choice formalism · 0.9causal inference · 0.9risk-averse planning · 0.3preference inference · 0.3markov decision process · 0.3inverse reinforcement learning · 0.3game theory · 0.3bounded optimality · 0.3approximate IRD · 0.3
YearPublicationVenuePosition
2026 Question the Questions: Auditing Representation in Online Deliberative Processes
abstract
A central feature of many deliberative processes, such as citizens' assemblies and deliberative polls, is the opportunity for participants to engage directly with experts. While participants are typically invited to propose questions for expert panels, only a limited number can be selected due to time constraints. This raises the challenge of how to choose a small set of questions that best represent the interests of all participants. We introduce an auditing framework for measuring the level of representation provided by a slate of questions, based on the social choice concept known as justified representation (JR). We present the first algorithms for auditing JR in the general utility setting, with our most efficient algorithm achieving a runtime of $O(mn\log n)$, where $n$ is the number of participants and $m$ is the number of proposed questions. We apply our auditing methods to historical deliberations, comparing the representativeness of (a) the actual questions posed to the expert panel (chosen by a moderator), (b) participants' questions chosen via integer linear programming, (c) summary questions generated by large language models (LLMs). Our results highlight both the promise and current limitations of LLMs in supporting deliberative processes. By integrating our methods into an online deliberation platform that has been used for over hundreds of deliberations across more than 50 countries, we make it easy for practitioners to audit and improve representation in future deliberations.
Soham De, Lodewijk Gelauff, Ashish Goel, Smitha Milli, Ariel D. Procaccia, Alice Siu
WWW4
2025 Representative Ranking for Deliberation in the Public Sphere
abstract
Online comment sections, such as those on news sites or social media, have the potential to foster informal public deliberation, However, this potential is often undermined by the frequency of toxic or low-quality exchanges that occur in these settings. To combat this, platforms increasingly leverage algorithmic ranking to facilitate higher-quality discussions, e.g., by using civility classifiers or forms of prosocial ranking. Yet, these interventions may also inadvertently reduce the visibility of legitimate viewpoints, undermining another key aspect of deliberation: representation of diverse views. We seek to remedy this problem by introducing guarantees of representation into these methods. In particular, we adopt the notion of *justified representation* (JR) from the social choice literature and incorporate a JR constraint into the comment ranking setting. We find that enforcing JR leads to greater inclusion of diverse viewpoints while still being compatible with optimizing for user engagement or other measures of conversational quality.
Manon Revel, Smitha Milli, Tyler Lu, Jamelle Watson-Daniels, Maximilian Nickel
ICML2
2020 Strategic Classification is Causal Modeling in Disguise
abstract
Consequential decision-making incentivizes individuals to strategically adapt their behavior to the specifics of the decision rule. While a long line of work has viewed strategic adaptation as gaming and attempted to mitigate its effects, recent work has instead sought to design classifiers that incentivize individuals to improve a desired quality. Key to both accounts is a cost function that dictates which adaptations are rational to undertake. In this work, we develop a causal framework for strategic adaptation. Our causal perspective clearly distinguishes between gaming and improvement and reveals an important obstacle to incentive design. We prove any procedure for designing classifiers that incentivize improvement must inevitably solve a non-trivial causal inference problem. We show a similar result holds for designing cost functions that satisfy the requirements of previous work. With the benefit of hindsight, our results show much of the prior work on strategic classification is causal modeling in disguise.
John Miller 0001, Smitha Milli, Moritz Hardt
ICML2
2020 Reward-rational (implicit) choice: A unifying formalism for reward learning
abstract
It is often difficult to hand-specify what the correct reward function is for a task, so researchers have instead aimed to learn reward functions from human behavior or feedback. The types of behavior interpreted as evidence of the reward function have expanded greatly in recent years. We've gone from demonstrations, to comparisons, to reading into the information leaked when the human is pushing the robot away or turning it off. And surely, there is more to come. How will a robot make sense of all these diverse types of behavior? Our key observation is that different types of behavior can be interpreted in a single unifying formalism - as a reward-rational choice that the human is making, often implicitly. We use this formalism to survey prior work through a unifying lens, and discuss its potential use as a recipe for interpreting new sources of information that are yet to be uncovered.
Hong Jun Jeon, Smitha Milli, Anca D. Dragan
NeurIPS2
2019 Literal or Pedagogic Human? Analyzing Human Model Misspecification in Objective Learning
Smitha Milli, Anca D. Dragan
UAI1
2017 When Does Bounded-Optimal Metareasoning Favor Few Cognitive Systems?
abstract
While optimal metareasoning is notoriously intractable, humans are nonetheless able to adaptively allocate their computational resources. A possible approximation that humans may use to do this is to only metareason over a finite set of cognitive systems that perform variable amounts of computation. The highly influential "dual-process" accounts of human cognition, which postulate the coexistence of a slow accurate system with a fast error-prone system, can be seen as a special case of this approximation. This raises two questions: how many cognitive systems should a bounded optimal agent be equipped with and what characteristics should those systems have? We investigate these questions in two settings: a one-shot decision between two alternatives, and planning under uncertainty in a Markov decision process. We find that the optimal number of systems depends on the variability of the environment and the costliness of metareasoning. Consistent with dual-process theories, we also find that when having two systems is optimal, then the first system is fast but error-prone and the second system is slow but accurate.
Smitha Milli, Falk Lieder, Thomas L. Griffiths 0001
AAAI1
2017 Should Robots be Obedient?
abstract
Intuitively, obedience -- following the order that a human gives -- seems like a good property for a robot to have. But, we humans are not perfect and we may give orders that are not best aligned to our preferences. We show that when a human is not perfectly rational then a robot that tries to infer and act according to the human's underlying preferences can always perform better than a robot that simply follows the human's literal order. Thus, there is a tradeoff between the obedience of a robot and the value it can attain for its owner. We investigate how this tradeoff is impacted by the way the robot infers the human's preferences, showing that some methods err more on the side of obedience than others. We then analyze how performance degrades when the robot has a misspecified model of the features that the human cares about or the level of rationality of the human. Finally, we study how robots can start detecting such model misspecification. Overall, our work suggests that there might be a middle ground in which robots intelligently decide when to obey human orders, but err on the side of obedience.
Smitha Milli, Dylan Hadfield-Menell, Anca D. Dragan, Stuart Russell 0001
IJCAI1
2017 Inverse Reward Design
abstract
Autonomous agents optimize the reward function we give them. What they don't know is how hard it is for us to design a reward function that actually captures what we want. When designing the reward, we might think of some specific training scenarios, and make sure that the reward will lead to the right behavior in those scenarios. Inevitably, agents encounter new scenarios (e.g., new types of terrain) where optimizing that same reward may lead to undesired behavior. Our insight is that reward functions are merely observations about what the designer actually wants, and that they should be interpreted in the context in which they were designed. We introduce inverse reward design (IRD) as the problem of inferring the true objective based on the designed reward and the training MDP. We introduce approximate methods for solving IRD problems, and use their solution to plan risk-averse behavior in test MDPs. Empirical results suggest that this approach can help alleviate negative side effects of misspecified reward functions and mitigate reward hacking.
Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, Stuart Russell 0001, Anca D. Dragan
NIPS2
2016 Beyond Canonical Texts: A Computational Analysis of Fanfiction
abstract
While much computational work on fiction has focused on works in the literary canon, user-created fanfiction presents a unique opportunity to study an ecosystem of literary production and consumption, embodying qualities both of large-scale literary data (55 billion tokens) and also a social network (with over 2 million users).We present several empirical analyses of this data in order to illustrate the range of affordances it presents to research in NLP, computational social science and the digital humanities.We find that fanfiction deprioritizes main protagonists in comparison to canonical texts, has a statistically significant difference in attention allocated to female characters, and offers a framework for developing models of reader reactions to stories.
Smitha Milli, David Bamman
EMNLP1