VLDB 2026 Research / reviewers in the wild / expert
Inioluwa Deborah Raji
dblp:228/7919
· DBLP profile ↗
8ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0002-9510-3015ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 5 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Evaluating Prediction-based Interventions with Human Decision Makers In MindabstractAutomated decision systems (ADS) are broadly deployed to inform or support human decision- making across a wide range of consequential contexts. However, various context-specific details complicate the goal of establishing meaningful experimental evaluations for prediction-based interventions. Notably, specific experimental design decisions may induce cognitive biases in human decision makers, which could then significantly alter the observed effect sizes of the prediction intervention. In this paper, we formalize and investigate various models of human decision-making in the presence of a predictive model aid. We show that each of these behavioral models produces dependencies across decision subjects and results in the violation of existing assumptions, with consequences for treatment effect estimation. This work aims to further advance the scientific validity of intervention-based evaluation schemes for the assessment of ADS deployments. Inioluwa Deborah Raji, Lydia T. Liu |
AISTATS | 1 |
| 2025 | Towards AI Accountability Infrastructure: Gaps and Opportunities in AI Audit ToolingabstractAudits are critical mechanisms for identifying the risks and limitations of deployed artifcial intelligence (AI) systems.However, the efective execution of AI audits remains incredibly difcult, and practitioners often need to make use of various tools to support their eforts.Drawing on interviews with 35 AI audit practitioners and a landscape analysis of 435 tools, we compare the current ecosystem of AI audit tooling to practitioner needs.While many tools are designed to help set standards and evaluate AI systems, they often fall short in supporting accountability.We outline challenges practitioners faced in their eforts to use AI audit tools and highlight areas for future tool development beyond evaluationfrom harms discovery to advocacy.We conclude that the available resources do not currently support the full scope of AI audit practitioners' needs and recommend that the feld move beyond tools for just evaluation and towards more comprehensive infrastructure for AI accountability. Victor Ojewale, Ryan Steed, Briana Vecchione, Abeba Birhane, Inioluwa Deborah Raji |
CHI | 5 |
| 2025 | From Individual Experience to Collective Evidence: A Reporting-Based Framework for Identifying Systemic HarmsabstractWhen an individual reports a negative interaction with some system, how can their personal experience be contextualized within broader patterns of system behavior? We study the *reporting database* problem, where individual reports of adverse events arrive sequentially, and are aggregated over time. In this work, our goal is to identify whether there are subgroups—defined by any combination of relevant
features—that are disproportionately likely to experience harmful interactions with the system. We formalize this problem as a sequential hypothesis test, and identify conditions on reporting behavior that are sufficient for making inferences about disparities in true rates of harm across subgroups. We show that algorithms for sequential hypothesis tests can be applied to this problem with a standard multiple testing
correction. We then demonstrate our method on real-world datasets, including mortgage decisions and vaccine side effects; on each, our method (re-)identifies subgroups known to experience disproportionate harm using only a fraction of the data that was initially used to discover them. Jessica Dai, Paula Gradu, Inioluwa Deborah Raji, Benjamin Recht |
ICML | 3 |
| 2025 | Measuring what Matters: Construct Validity in Large Language Model BenchmarksabstractEvaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as safety' androbustness' requires strong construct validity, that is, having measures that represent what matters to the phenomenon. With a team of 29 expert reviewers, we conduct a systematic review of 445 LLM benchmarks from leading conferences in natural language processing and machine learning. Across the reviewed articles, we find patterns related to the measured phenomena, tasks, and scoring metrics which undermine the validity of the resulting claims. To address these shortcomings, we provide eight key recommendations and detailed actionable guidance to researchers and practitioners in developing LLM benchmarks. Andrew M. Bean 0001, Ryan Othniel Kearns, Angelika Romanou, Franziska Sofia Hafner, Harry Mayne, Jan Batzner, Negar Foroutan Eghlidi, Chris Schmitz, Karolina Korgul, Hunar Batra, Oishi Deb, Emma Beharry, Cornelius Emde, Thomas Foster, Anna Gausen, María Grandury, Simeng Han, Valentin Hofmann, Lujain Ibrahim, Hazel Kim, Hannah Kirk, Fangru Lin, Gabrielle K. Liu, Lennart Luettgau, Jabez Magomere, Jonathan Rystrøm, Anna Sotnikova, Yilun Zhao 0001, Adel Bibi, Antoine Bosselut, Ronald Clark, Arman Cohan, Jakob N. Foerster, Yarin Gal, Scott A. Hale, Inioluwa Deborah Raji, Christopher Summerfield, Philip Torr 0001, Cozmin Ududec, Luc Rocher, Adam Mahdi |
NeurIPS | 37 |
| 2022 | From Algorithmic Audits to Actual Accountability: Overcoming Practical Roadblocks on the Path to Meaningful Audit Interventions for AI GovernanceabstractAs algorithmic deployments become more and more common, policymakers and advocates are increasingly turning to audits as an approach for accountability. Some audits have already led to product updates or recalls, organizational changes and developments to regulation or standards. However, difficulties in execution, oversight and impact threaten the credibility and effectiveness of these audits. Inioluwa Deborah Raji |
AIES | 1 |
| 2022 | Outsider Oversight: Designing a Third Party Audit Ecosystem for AI GovernanceabstractMuch attention has focused on algorithmic audits and impact assessments to hold developers and users of algorithmic systems accountable. But existing algorithmic accountability policy approaches have neglected the lessons from non-algorithmic domains: notably, the importance of third parties. Our paper synthesizes lessons from other fields on how to craft effective systems of external oversight for algorithmic deployments. First, we discuss the challenges of third party oversight in the current AI landscape. Second, we survey audit systems across domains - e.g., financial, environmental, and health regulation - and show that the institutional design of such audits are far from monolithic. Finally, we survey the evidence base around these design components and spell out the implications for algorithmic auditing. We conclude that the turn toward audits alone is unlikely to achieve actual algorithmic accountability, and sustained focus on institutional design will be required for meaningful third party involvement. Inioluwa Deborah Raji, Peggy Xu, Colleen Honigsberg, Daniel E. Ho |
AIES | 1 |
| 2020 | Saving Face: Investigating the Ethical Concerns of Facial Recognition AuditingabstractAlthough essential to revealing biased performance, well intentioned attempts at algorithmic auditing can have effects that may harm the very populations these measures are meant to protect. This concern is even more salient while auditing biometric systems such as facial recognition, where the data is sensitive and the technology is often used in ethically questionable manners. We demonstrate a set of fiveethical concerns in the particular case of auditing commercial facial processing technology, highlighting additional design considerations and ethical tensions the auditor needs to be aware of so as not exacerbate or complement the harms propagated by the audited system. We go further to provide tangible illustrations of these concerns, and conclude by reflecting on what these concerns mean for the role of the algorithmic audit and the fundamental product limitations they reveal. Inioluwa Deborah Raji, Timnit Gebru, Margaret Mitchell, Joy Buolamwini, Joonseok Lee, Remi Denton |
AIES | 1 |
| 2019 | Actionable Auditing: Investigating the Impact of Publicly Naming Biased Performance Results of Commercial AI ProductsabstractAlthough algorithmic auditing has emerged as a key strategy to expose systematic biases embedded in software platforms, we struggle to understand the real-world impact of these audits, as scholarship on the impact of algorithmic audits on increasing algorithmic fairness and transparency in commercial systems is nascent. To analyze the impact of publicly naming and disclosing performance results of biased AI systems, we investigate the commercial impact of Gender Shades, the first algorithmic audit of gender and skin type performance disparities in commercial facial analysis models. This paper 1) outlines the audit design and structured disclosure procedure used in the Gender Shades study, 2) presents new performance metrics from targeted companies IBM, Microsoft and Megvii (Face++) on the Pilot Parliaments Benchmark (PPB) as of August 2018, 3) provides performance results on PPB by non-target companies Amazon and Kairos and, 4) explores differences in company responses as shared through corporate communications that contextualize differences in performance on PPB. Within 7 months of the original audit, we find that all three targets released new API versions. All targets reduced accuracy disparities between males and females and darker and lighter-skinned subgroups, with the most significant update occurring for the darker-skinned female subgroup, that underwent a 17.7% - 30.4% reduction in error between audit periods. Minimizing these disparities led to a 5.72% to 8.3% reduction in overall error on the Pilot Parliaments Benchmark (PPB) for target corporation APIs. The overall performance of non-targets Amazon and Kairos lags significantly behind that of the targets, with error rates of 8.66% and 6.60% overall, and error rates of 31.37% and 22.50% for the darker female subgroup, respectively. Inioluwa Deborah Raji, Joy Buolamwini |
AIES | 1 |