Jessica Dai

dblp:278/3333 · DBLP profile ↗
← Back
7ranked-venue papers
7as first author
7since 2021 · last 2025
0000-0001-9047-4004ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Trustworthy machine learning · 63% Probabilistic and Bayesian machine learning · 21% Learning theory · 16%
Theoretical computer science
1 paper
Mathematical optimization · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 100%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
1.722025
Learning With Multi-Group Guarantees For Clusterable Subpopulations · ICML 2025
From Individual Experience to Collective Evidence: A Reporting-Based Framework for Identifying Systemic Harms · ICML 2025
Machine learning › Trustworthy machine learning
calibration
0.912025
Learning With Multi-Group Guarantees For Clusterable Subpopulations · ICML 2025
Machine learning › Learning theory › online learning
online calibration
0.912025
Learning With Multi-Group Guarantees For Clusterable Subpopulations · ICML 2025
Machine learning › Trustworthy machine learning
ethical AI
0.812024
Position: Beyond Personhood: Agency, Accountability, and the Limits of Anthropomorphic Ethical Analysis · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › experimental design
adaptive experimental design
0.712023
CLIP-OGD: An Experimental Design for Adaptive Neyman Allocation in Sequential Experiments · NeurIPS 2023
Mathematical optimization › online optimization
online gradient descent
0.712023
CLIP-OGD: An Experimental Design for Adaptive Neyman Allocation in Sequential Experiments · NeurIPS 2023
Mathematical optimization
online optimization
0.712023
CLIP-OGD: An Experimental Design for Adaptive Neyman Allocation in Sequential Experiments · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
mixture model
0.312025
Learning With Multi-Group Guarantees For Clusterable Subpopulations · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.212023
CLIP-OGD: An Experimental Design for Adaptive Neyman Allocation in Sequential Experiments · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

sequential hypothesis testing · 1.7multiple testing correction · 1.7political science · 1.5philosophy · 1.5variance estimation · 1.3potential outcomes framework · 1.3online gradient descent · 1.3online calibration · 0.9multi-objective optimization · 0.9
YearPublicationVenuePosition
2025 From Individual Experience to Collective Evidence: A Reporting-Based Framework for Identifying Systemic Harms
abstract
When an individual reports a negative interaction with some system, how can their personal experience be contextualized within broader patterns of system behavior? We study the *reporting database* problem, where individual reports of adverse events arrive sequentially, and are aggregated over time. In this work, our goal is to identify whether there are subgroups—defined by any combination of relevant features—that are disproportionately likely to experience harmful interactions with the system. We formalize this problem as a sequential hypothesis test, and identify conditions on reporting behavior that are sufficient for making inferences about disparities in true rates of harm across subgroups. We show that algorithms for sequential hypothesis tests can be applied to this problem with a standard multiple testing correction. We then demonstrate our method on real-world datasets, including mortgage decisions and vaccine side effects; on each, our method (re-)identifies subgroups known to experience disproportionate harm using only a fraction of the data that was initially used to discover them.
Jessica Dai, Paula Gradu, Inioluwa Deborah Raji, Benjamin Recht
ICML1
2025 Learning With Multi-Group Guarantees For Clusterable Subpopulations
abstract
A canonical desideratum for prediction problems is that performance guarantees should hold not just on average over the population, but also for meaningful subpopulations within the overall population. But what constitutes a meaningful subpopulation? In this work, we take the perspective that relevant subpopulations should be defined with respect to the clusters that naturally emerge from the distribution of individuals for which predictions are being made. In this view, a population refers to a mixture model whose components constitute the relevant subpopulations. We suggest two formalisms for capturing per-subgroup guarantees: first, by attributing each individual to the component from which they were most likely drawn, given their features; and second, by attributing each individual to all components in proportion to their relative likelihood of having been drawn from each component. Using online calibration as a case study, we study a multi-objective algorithm that provides guarantees for each of these formalisms by handling all plausible underlying subpopulation structures simultaneously, and achieve an $O(T^{1/2})$ rate even when the subpopulations are not well-separated. In comparison, the more natural cluster-then-predict approach that first recovers the structure of the subpopulations and then makes predictions suffers from a $O(T^{2/3})$ rate and requires the subpopulations to be separable. Along the way, we prove that providing per-subgroup calibration guarantees for underlying clusters can be easier than learning the clusters: separation between median subgroup features is required for the latter but not the former.
Jessica Dai, Nika Haghtalab, Eric Zhao 0003
ICML1
2024 Can Probabilistic Feedback Drive User Impacts in Online Platforms?
abstract
A common explanation for negative user impacts of content recommender systems is misalignment between the platform’s objective and user welfare. In this work, we show that misalignment in the platform’s objective is not the only potential cause of unintended impacts on users: even when the platform’s objective is fully aligned with user welfare, the platform’s learning algorithm can induce negative downstream impacts on users. The source of these user impacts is that different pieces of content may generate observable user reactions (feedback information) at different rates; these feedback rates may correlate with content properties, such as controversiality or demographic similarity of the creator, that affect the user experience. Since differences in feedback rates can impact how often the learning algorithm engages with different content, the learning algorithm may inadvertently promote content with certain such properties. Using the multi-armed bandit framework with probabilistic feedback, we examine the relationship between feedback rates and a learning algorithm’s engagement with individual arms for different no-regret algorithms. We prove that no-regret algorithms can exhibit a wide range of dependencies: if the feedback rate of an arm increases, some no-regret algorithms engage with the arm more, some no-regret algorithms engage with the arm less, and other no-regret algorithms engage with the arm approximately the same number of times. From a platform design perspective, our results highlight the importance of looking beyond regret when measuring an algorithm’s performance, and assessing the nature of a learning algorithm’s engagement with different types of content as well as their resulting downstream impacts.
Jessica Dai, Bailey Flanigan, Meena Jagadeesan, Nika Haghtalab, Chara Podimata
AISTATS1
2024 Position: Beyond Personhood: Agency, Accountability, and the Limits of Anthropomorphic Ethical Analysis
abstract
What is *agency,* and why does it matter? In this work, we draw from the political science and philosophy literature and give two competing visions of what it means to be an (ethical) agent. The first view, which we term *mechanistic*, is commonly— and implicitly—assumed in AI research, yet it is a fundamentally limited means to understand the ethical characteristics of AI. Under the second view, which we term volitional, AI can no longer be considered an ethical agent. We discuss the implications of each of these views for two critical questions: first, what the ideal system “ought” to look like, and second, how accountability may be achieved. In light of this discussion, we ultimately argue that, in the context of ethically-significant behavior, AI should be viewed not as an agent but as the outcome of political processes.
Jessica Dai
ICML1
2023 CLIP-OGD: An Experimental Design for Adaptive Neyman Allocation in Sequential Experiments
abstract
From clinical development of cancer therapies to investigations into partisan bias, adaptive sequential designs have become increasingly popular method for causal inference, as they offer the possibility of improved precision over their non-adaptive counterparts. However, even in simple settings (e.g. two treatments) the extent to which adaptive designs can improve precision is not sufficiently well understood. In this work, we study the problem of Adaptive Neyman Allocation in a design-based potential outcomes framework, where the experimenter seeks to construct an adaptive design which is nearly as efficient as the optimal (but infeasible) non-adaptive Neyman design, which has access to all potential outcomes. Motivated by connections to online optimization, we propose Neyman Ratio and Neyman Regret as two (equivalent) performance measures of adaptive designs for this problem. We present Clip-OGD, an adaptive design which achieves $\widetilde{\mathcal{O}}(\sqrt{T})$ expected Neyman regret and thereby recovers the optimal Neyman variance in large samples. Finally, we construct a conservative variance estimator which facilitates the development of asymptotically valid confidence intervals. To complement our theoretical results, we conduct simulations using data from a microeconomic experiment.
Jessica Dai, Paula Gradu, Christopher Harshaw
NeurIPS1
2022 Fairness via Explanation Quality: Evaluating Disparities in the Quality of Post hoc Explanations
abstract
As post hoc explanation methods are increasingly being leveraged to explain complex models in high-stakes settings, it becomes critical to ensure that the quality of the resulting explanations is consistently high across all subgroups of a population. For instance, it should not be the case that explanations associated with instances belonging to, e.g., women, are less accurate than those associated with other genders. In this work, we initiate the study of identifying group-based disparities in explanation quality. To this end, we first outline several key properties that contribute to explanation quality-namely, fidelity (accuracy), stability, consistency, and sparsity-and discuss why and how disparities in these properties can be particularly problematic. We then propose an evaluation framework which can quantitatively measure disparities in the quality of explanations. Using this framework, we carry out an empirical analysis with three datasets, six post hoc explanation methods, and different model classes to understand if and when group-based disparities in explanation quality arise. Our results indicate that such disparities are more likely to occur when the models being explained are complex and non-linear. We also observe that certain post hoc explanation methods (e.g., Integrated Gradients, SHAP) are more likely to exhibit disparities. Our work sheds light on previously unexplored ways in which explanation methods may introduce unfairness in real world decision making.
Jessica Dai, Sohini Upadhyay, Ulrich Aïvodji, Stephen H. Bach, Himabindu Lakkaraju
AIES1
2021 Fair Machine Learning Under Partial Compliance
abstract
Typically, fair machine learning research focuses on a single decision maker and assumes that the underlying population is stationary. However, many of the critical domains motivating this work are characterized by competitive marketplaces with many decision makers. Realistically, we might expect only a subset of them to adopt any non-compulsory fairness-conscious policy, a situation that political philosophers call partial compliance. This possibility raises important questions: how does partial compliance and the consequent strategic behavior of decision subjects affect the allocation outcomes? If k% of employers were to voluntarily adopt a fairness-promoting intervention, should we expect k% progress (in aggregate) towards the benefits of universal adoption, or will the dynamics of partial compliance wash out the hoped-for benefits? How might adopting a global (versus local) perspective impact the conclusions of an auditor? In this paper, we propose a simple model of an employment market, leveraging simulation as a tool to explore the impact of both interaction effects and incentive effects on outcomes and auditing metrics. Our key findings are that at equilibrium: (1) partial compliance by k% of employers can result in far less than proportional (k%) progress towards the full compliance outcomes; (2) the gap is more severe when fair employers match global (vs local) statistics; (3) choices of local vs global statistics can paint dramatically different pictures of the performance vis-a-vis fairness desiderata of compliant versus non-compliant employers; (4) partial compliance based on local parity measures can induce extreme segregation. Finally, we discuss implications for auditors and insights concerning the design of regulatory frameworks.
Jessica Dai, Sina Fazelpour, Zachary C. Lipton
AIES1