VLDB 2026 Research / reviewers in the wild / expert
Michael A. Madaio
dblp:162/5246 · also Michael Madaio
· DBLP profile ↗
25ranked-venue papers
10as first author
12since 2021 · last 2026
0000-0001-5772-0488ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 18 · 7 first-author · 11 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Protection or Empowerment? Perspectives on Youth-Inclusive Responsible AIabstractArtificial intelligence (AI) systems are increasingly prevalent in youths’ lives, despite documentation and concern of potential algorithmic harm. With responsible AI (RAI) efforts aiming to address harms, youth are largely overlooked as contributors, despite being stakeholders of AI. This paper explores perspectives on challenges, barriers, and opportunities of youth participation in responsibly creating AI. Through workshops with 16 teens and 11 parents, as well as interviews with 8 AI practitioners, we find that while all groups recognize the value of youth perspectives, the opinions of those not directly paying for AI are not prioritized. Many parents were optimistic about their youths’ ability to contribute, while youth showed a desire to contribute, coupled with an awareness of their strengths and limitations. Practitioners saw the potential of youth to contribute but not necessarily as empowered decision makers in RAI processes. We offer insights on involving youth in RAI, balancing protection with agency. Jaemarie Solyst, Cindy Peng, Praneetha Pratapa, Claire Wang 0002, Amy Ogan, Jessica Hammer, Michael A. Madaio, Motahhare Eslami |
IDC | 7 |
| 2026 | Navigating Uncertainties: How GenAI Developers Document Their Models on Open-Source PlatformsabstractModel documentation plays a crucial role in promoting responsible AI (RAI) development. The emergence of Generative AI (GenAI) models has reshaped the conditions under which documentation is produced, particularly on open-source platforms where models are hosted and shared. To examine how these changes have manifested in developers’ documentation practices, we interviewed 17 GenAI developers who document models on open-source platforms. Our findings illustrate that uncertainties have become a defining feature of developers’ GenAI documentation practices and that these uncertainties unfold in three interrelated forms: (1) normative and epistemic uncertainties in determining documentation content; (2) methodological uncertainties in evaluating and communicating model properties; and (3) ecosystemic uncertainties about who should document. We argue that these uncertainties in GenAI documentation require coordinated interventions, including infrastructural support to address epistemic and methodological uncertainties, community-based mechanisms to cultivate RAI documentation norms, and collaboration across supply-chain actors to address ecosystemic uncertainties. Ningjing Tang, Megan Li, Amy A. Winecoff, Michael A. Madaio, Hoda Heidari, Hong Shen 0004 |
CHI | 4 |
| 2025 | "The Conduit by which Change Happens": Processes, Barriers, and Support for Interpersonal Learning about Responsible AIabstractResponsible AI (RAI) practices are increasingly important for practitioners in anticipating and addressing potential harms of AI, and emerging research suggests that AI practitioners often learn about RAI on-the-job.More generally, learning at work is social; thus, this work explores the interpersonal aspects of learning about RAI on-the-job.Through workshops with 21 industry-based RAI educators, we offer the first empirical investigation into interpersonal processes and dimensions of learning about RAI at work.This study finds key phases of RAI are sites for ongoing interpersonal learning, such as critical reflection about potential RAI impacts and collective sense-making about operationalizing RAI principles.We uncover a significant gap between these interpersonal learning processes and current approaches to learning about RAI.Finally, we identify barriers and supports for interpersonal learning about RAI.We close by discussing opportunities to better enable interpersonal learning about RAI on-the-job and the broader implications of interpersonal learning for RAI. Jaemarie Solyst, Lauren Wilcox, Michael A. Madaio |
CHI | 3 |
| 2024 | Scaling Laws Do Not ScaleabstractRecent work has advocated for training AI models on ever-larger datasets, arguing that as the size of a dataset increases, the performance of a model trained on that dataset will correspondingly increase (referred to as “scaling laws”). In this paper, we draw on literature from the social sciences and machine learning to critically interrogate these claims. We argue that this scaling law relationship depends on metrics used to measure performance that may not correspond with how different groups of people perceive the quality of models' output. As the size of datasets used to train large AI models grows and AI systems impact ever larger groups of people, the number of distinct communities represented in training or evaluation datasets grows. It is thus even more likely that communities represented in datasets may have values or preferences not reflected in (or at odds with) the metrics used to evaluate model performance in scaling laws. Different communities may also have values in tension with each other, leading to difficult, potentially irreconcilable choices about metrics used for model evaluations---threatening the validity of claims that model performance is improving at scale. We end the paper with implications for AI development: that the motivation for scraping ever-larger datasets may be based on fundamentally flawed assumptions about model performance. That is, models may not, in fact, continue to improve as the datasets get larger---at least not for all people or communities impacted by those models. We suggest opportunities for the field to rethink norms and values in AI development, resisting claims for universality of large models, fostering more local, small-scale designs, and other ways to resist the impetus towards scale in AI. Michael A. Madaio |
AIES (1) | 2 |
| 2024 | A Scoping Study of Evaluation Practices for Responsible AI Tools: Steps Towards Effectiveness EvaluationsabstractResponsible design of AI systems is a shared goal across HCI and AI communities. Responsible AI (RAI) tools have been developed to support practitioners to identify, assess, and mitigate ethical issues during AI development. These tools take many forms (e.g., design playbooks, software toolkits, documentation protocols). However, research suggests that use of RAI tools is shaped by organizational contexts, raising questions about how effective such tools are in practice. To better understand how RAI tools are—and might be—evaluated, we conducted a qualitative analysis of 37 publications that discuss evaluations of RAI tools. We find that most evaluations focus on usability, while questions of tools’ effectiveness in changing AI development are sidelined. While usability evaluations are an important approach to evaluate RAI tools, we draw on evaluation approaches from other fields to highlight developer- and community-level steps to support evaluations of RAI tools’ effectiveness in shaping AI development practices and outcomes. Glen Berman, Nitesh Goyal, Michael A. Madaio |
CHI | 3 |
| 2024 | Farsight: Fostering Responsible AI Awareness During AI Application PrototypingabstractPrompt-based interfaces for Large Language Models (LLMs) have made prototyping and building AI-powered applications easier than ever before. However, identifying potential harms that may arise from AI applications remains a challenge, particularly during prompt-based prototyping. To address this, we present Farsight, a novel in situ interactive tool that helps people identify potential harms from the AI applications they are prototyping. Based on a user’s prompt, Farsight highlights news articles about relevant AI incidents and allows users to explore and edit LLM-generated use cases, stakeholders, and harms. We report design insights from a co-design study with 10 AI prototypers and findings from a user study with 42 AI prototypers. After using Farsight, AI prototypers in our user study are better able to independently identify potential harms associated with a prompt and find our tool more useful and usable than existing resources. Their qualitative feedback also highlights that Farsight encourages them to focus on end-users and think beyond immediate harms. We discuss these findings and reflect on their implications for designing AI prototyping experiences that meaningfully engage with AI harms. Farsight is publicly accessible at: https://pair-code.github.io/farsight. Zijie J. Wang, Chinmay Kulkarni 0001, Lauren Wilcox, Michael Terry, Michael A. Madaio |
CHI | 5 |
| 2024 | Tinker, Tailor, Configure, Customize: The Articulation Work of Contextualizing an AI Fairness ChecklistabstractMany responsible AI resources, such as toolkits, playbooks, and checklists, have been developed to support AI practitioners in identifying, measuring, and mitigating potential fairness-related harms. These resources are often designed to be general purpose in order to be applicable to a variety of use cases, domains, and deployment contexts. However, this can lead to decontextualization, where such resources lack the level of relevance or specificity needed to use them. To understand how AI practitioners might contextualize one such resource, an AI fairness checklist, for their particular use cases, domains, and deployment contexts, we conducted a retrospective contextual inquiry with 13 AI practitioners from seven organizations. We identify how contextualizing this checklist introduces new forms of work for AI practitioners and other stakeholders, as well as opening up new sites for negotiation and contestation of values in AI. We also identify how the contextualization process may help AI practitioners develop a shared language around AI fairness, and we identify tensions related to ownership over this process that suggest larger issues of accountability in responsible AI work. Michael A. Madaio, Hanna M. Wallach, Jennifer Wortman Vaughan |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2023 | Designing Responsible AI: Adaptations of UX Practice to Meet Responsible AI ChallengesabstractTechnology companies continue to invest in efforts to incorporate responsibility in their Artificial Intelligence (AI) advancements, while efforts to audit and regulate AI systems expand. This shift towards Responsible AI (RAI) in the tech industry necessitates new practices and adaptations to roles—undertaken by a variety of practitioners in more or less formal positions, many of whom focus on the user-centered aspects of AI. To better understand practices at the intersection of user experience (UX) and RAI, we conducted an interview study with industrial UX practitioners and RAI subject matter experts, both of whom are actively involved in addressing RAI concerns throughout the early design and development of new AI-based prototypes, demos, and products, at a large technology company. Many of the specific practices and their associated challenges have yet to be surfaced in the literature, and distilling them offers a critical view into how practitioners’ roles are adapting to meet present-day RAI challenges. We present and discuss three emerging practices in which RAI is being enacted and reified in UX practitioners’ everyday work. We conclude by arguing that the emerging practices, goals, and types of expertise that surfaced in our study point to an evolution in praxis, with associated challenges that suggest important areas for further research in HCI. Qiaosi Wang, Michael A. Madaio, Shaun K. Kane, Shivani Kapania, Michael Terry, Lauren Wilcox |
CHI | 2 |
| 2023 | Fairlearn: Assessing and Improving Fairness of AI SystemsabstractFairlearn is an open source project to help practitioners assess and improve fairness of artificial intelligence (AI) systems. The associated Python library, also named fairlearn, supports evaluation of a model's output across affected populations and includes several algorithms for mitigating fairness issues. Grounded in the understanding that fairness is a sociotechnical challenge, the project integrates learning resources that aid practitioners in considering a system's broader societal context. Hilde J. P. Weerts, Miroslav Dudík, Richard Edgar, Adrin Jalali, Roman Lutz, Michael A. Madaio |
J. Mach. Learn. Res. | 6 |
| 2023 | Seeing Like a Toolkit: How Toolkits Envision the Work of AI EthicsabstractNumerous toolkits have been developed to support ethical AI development. However, toolkits, like all tools, encode assumptions in their design about what work should be done and how. In this paper, we conduct a qualitative analysis of 27 AI ethics toolkits to critically examine how the work of ethics is imagined and how it is supported by these toolkits. Specifically, we examine the discourses toolkits rely on when talking about ethical issues, who they imagine should do the work of ethics, and how they envision the work practices involved in addressing ethics. Among the toolkits, we identify a mismatch between the imagined work of ethics and the support the toolkits provide for doing that work. In particular, we identify a lack of guidance around how to navigate labor, organizational, and institutional power dynamics as they relate to performing ethical work. We use these omissions to chart future work for researchers and designers of AI ethics toolkits. Richmond Y. Wong, Michael A. Madaio, Nick Merrill |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2022 | Part of the Conversation: Workforce Professionals' Perspectives on the Roles and Impacts of Workforce TechnologiesabstractAmidst recent enthusiasm for data-driven technologies in workforce development, prior HCI research has explored job seekers' perspectives to inform the design of new technologies that could support their job search. However, in practice, the process of looking for work is often embedded in local workforce development ecosystems, where networks of organizations provide a range of services, from employment consulting, to resume workshops, to job skills training programs. Although prior CSCW work has explored the role of algorithms in social services, there has been little work investigating how algorithmic systems may shape workforce development professionals' interactions with clients and how they might be better designed to complement these professionals' work and responsibilities. To begin to address this gap, we conducted an interview study with five workforce development professionals in the US in both management and client-facing roles. Our findings contribute to research on how algorithmic systems are shaping workforce development, shedding light on the importance of the relationship building work that workforce professionals engage in with clients, the difficulty in maintaining boundaries in the face of resource and information challenges, and the ways that workforce development technologies are shaping the work of workforce development. Connie W. Chau, Kenneth Holstein, Michael A. Madaio |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2022 | Assessing the Fairness of AI Systems: AI Practitioners' Processes, Challenges, and Needs for SupportabstractVarious tools and practices have been developed to support practitioners in identifying, assessing, and mitigating fairness-related harms caused by AI systems. However, prior research has highlighted gaps between the intended design of these tools and practices and their use within particular contexts, including gaps caused by the role that organizational factors play in shaping fairness work. In this paper, we investigate these gaps for one such practice: disaggregated evaluations of AI systems, intended to uncover performance disparities between demographic groups. By conducting semi-structured interviews and structured workshops with thirty-three AI practitioners from ten teams at three technology companies, we identify practitioners' processes, challenges, and needs for support when designing disaggregated evaluations. We find that practitioners face challenges when choosing performance metrics, identifying the most relevant direct stakeholders and demographic groups on which to focus, and collecting datasets with which to conduct disaggregated evaluations. More generally, we identify impacts on fairness work stemming from a lack of engagement with direct stakeholders or domain experts, business imperatives that prioritize customers over marginalized groups, and the drive to deploy AI systems at scale. Michael A. Madaio, Lisa Egede, Hariharan Subramonyam, Jennifer Wortman Vaughan, Hanna M. Wallach |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2020 | Predicting Gaps in Usage in a Phone-Based Literacy Intervention System
Rishabh Chatterjee, Michael A. Madaio, Amy Ogan |
AIED (1) | 2 |
| 2020 | Co-Designing Checklists to Understand Organizational Challenges and Opportunities around Fairness in AIabstractMany organizations have published principles intended to guide the ethical development and deployment of AI systems; however, their abstract nature makes them difficult to operationalize. Some organizations have therefore produced AI ethics checklists, as well as checklists for more specific concepts, such as fairness, as applied to AI systems. But unless checklists are grounded in practitioners' needs, they may be misused. To understand the role of checklists in AI ethics, we conducted an iterative co-design process with 48 practitioners, focusing on fairness. We co-designed an AI fairness checklist and identified desiderata and concerns for AI fairness checklists in general. We found that AI fairness checklists could provide organizational infrastructure for formalizing ad-hoc processes and empowering individual advocates. We highlight aspects of organizational culture that may impact the efficacy of AI fairness checklists, and suggest future design directions. Michael A. Madaio, Luke Stark, Jennifer Wortman Vaughan, Hanna M. Wallach |
CHI | 1 |
| 2020 | Collective Support and Independent Learning with a Voice-Based Literacy Technology in Rural CommunitiesabstractAccess to literacy is critical to children's futures, but formal education may be insufficient for fostering early literacy, especially in low-resource contexts. Educational technologies used at home may be able to help, but it is unclear whether or how children (and families) will use such technologies at home in rural communities, particularly in low-literate families. In this paper, we investigate these questions with a voice-based literacy technology deployed with families in 8 rural communities in Côte d'Ivoire for 4 months. We use interviews and observations with 37 families to investigate motivations, methods, and barriers for rural families' engagement with a literacy technology accessible via feature phones. We contribute insights into how families view digital literacy as a learning goal, leverage networks of supporters, and over time, transition from explicit to implicit support for children's learning. Michael A. Madaio, Evelyn Yarzebinski, Vikram Kamath Cannanure, Benjamin Zinszer, Joelle Hannon-Cropp, Fabrice Tanoh, Akpe Yapo Hermann, Axel Blahoua Seri, Kaja Jasinska, Amy Ogan |
CHI | 1 |
| 2019 | "Everyone Brings Their Grain of Salt": Designing for Low-Literate Parental Engagement with a Mobile Literacy Technology in Côte d'IvoireabstractSignificant research has demonstrated the crucial role that parents play in supporting the development of children's literacy, but in contexts where adults may lack sufficient literacy in the target language, it is not clear how to most effectively scaffold parental support for children's literacy. Prior work has designed technologies to teach children literacy directly, but this work has not focused on designing for low-literate parents, particularly for multilingual and developing contexts. In this paper, we describe findings from a qualitative study conducted in several regions of rural Côte d'Ivoire to understand Ivorian parents' beliefs, desires, and preferences for French literacy. We discuss themes that emerged from these interviews, surrounding ideas of trust, collaboration, and culturally-responsive design, and we highlight implications for the design of technology to scaffold low-literate parental support for children's literacy. Michael A. Madaio, Fabrice Tanoh, Axel Blahoua Seri, Kaja Jasinska, Amy Ogan |
CHI | 1 |
| 2019 | "You give a little of yourself": family support for children's use of an IVR literacy systemabstractLow levels of childhood literacy in global contexts may be mitigated by educational technologies, however, these technologies often rely on parents of sufficient literacy to effectively support their children. Given low levels of adult literacy in many low-resource contexts, we investigate the nature of low-literate adult support for children's use of a literacy technology designed to foster early literacy precursors. We deployed an interactive voice response (IVR) system with 38 families in a rural village in Côte d'Ivoire using the IVR for 5 weeks in their homes. Using call log data and grounded theory analyses of IVR observations and interviews, we find evidence that families leverage complex support networks where family members support children's use of the IVR in different ways, via a collective network of intermediaries. These results suggest opportunities to scaffold low-literate family supporters for educational technologies. Michael A. Madaio, Vikram Kamath Cannanure, Evelyn Yarzebinski, Shelby Zasacky, Fabrice Tanoh, Joelle Hannon-Cropp, Justine Cassell, Kaja Jasinska, Amy Ogan |
COMPASS | 1 |
| 2018 | Designing Appropriate Learning Technologies for School vs Home Settings in Tanzanian Rural VillagesabstractSmartphone- and tablet-based learning systems are often posited as solutions for closing early literacy gaps between rural and urban regions in emerging economies. These systems are often developed based on experiences with students in urban contexts, limiting their success rates with children from rural areas who have had little to no prior exposure to technology. To explore how such technologies are used in different learning contexts, we deployed an early literacy learning application in school and home settings in a rural village in Tanzania. We use Rogoff's theory of instructional models to understand and describe the interaction between learners, adults, and peers. We found that in the presence of a school teacher, the instructional model was primarily "adult-run" where information was almost entirely disseminated by the teacher, while in home settings, the instructional model was similar to a "community-of-learners" model where children collaborate with other peers and adults to achieve their learning goals. We use these instructional models to surface six themes of support and scaffolding that were expressed differently across settings, and discuss the benefits and drawbacks of the instructional models observed in providing support across these themes. Judith Uchidiuno, Evelyn Yarzebinski, Michael A. Madaio, Nupur Maheshwari, Kenneth R. Koedinger, Amy Ogan |
COMPASS | 3 |
| 2018 | A Dynamic Pipeline for Spatio-Temporal Fire Risk PredictionabstractRecent high-profile fire incidents in cities around the world have highlighted gaps in fire risk reduction efforts, as cities grapple with fewer resources and more properties to safeguard. To address this resource gap, prior work has developed machine learning frameworks to predict fire risk and prioritize fire inspections. However, existing approaches were limited by not including time-varying data, never deploying in real-time, and only predicting risk for a small subset of commercial properties in their city. Here, we have developed a predictive risk framework for all 20,636 commercial properties in Pittsburgh, based on time-varying data from a variety of municipal agencies. We have deployed our fire risk model on Pittsburgh Bureau of Fire's (PBF), and we have developed preliminary risk models for residential property fire risk prediction. Our commercial risk model outperforms the prior state of the art with a kappa of 0.33 compared to their 0.17, and is able to be applied to nearly 4 times as many properties as the prior model. In the 5 weeks since our model was first deployed, 58% of our predicted high-risk properties had a fire incident of any kind, while 23% of the building fire incidents that occurred took place in our predicted high or medium risk properties. The risk scores from our commercial model are visualized on an interactive dashboard and map to assist the PBF with planning their fire risk reduction initiatives. This work is already helping to improve fire risk reduction in Pittsburgh and is beginning to be adopted by other cities. Bhavkaran Singh Walia, Qianyi Hu, Jeffrey Chen, Fangyan Chen, Jessica Lee, Nathan Kuo, Palak Narang, Jason Batts, Geoffrey Arnold, Michael A. Madaio |
KDD | 10 |
| 2017 | Using Temporal Association Rule Mining to Predict Dyadic Rapport in Peer Tutoring
Michael A. Madaio, Rae Lasko, Justine Cassell, Amy Ogan |
EDM | 1 |
| 2017 | Temporally Selective Attention Model for Social and Affective State Recognition in Multimedia ContentabstractThe sheer amount of human-centric multimedia content has led to increased research on human behavior understanding. Most existing methods model behavioral sequences without considering the temporal saliency. This work is motivated by the psychological observation that temporally selective attention enables the human perceptual system to process the most relevant information. In this paper, we introduce a new approach, named Temporally Selective Attention Model (TSAM), designed to selectively attend to salient parts of human-centric video sequences. Our TSAM models learn to recognize affective and social states using a new loss function called speaker-distribution loss. Extensive experiments show that our model achieves the state-of-the-art performance on rapport detection and multimodal sentiment analysis. We also show that our speaker-distribution loss function can generalize to other computational models, improving the prediction performance of deep averaging network and Long Short Term Memory (LSTM). Liangke Gui, Michael A. Madaio, Amy Ogan, Justine Cassell, Louis-Philippe Morency |
ACM Multimedia | 3 |
| 2016 | Experiences with MOOCs in a West-African Technology HubabstractMassive, Open, Online Courses (MOOCs) have been promoted as a means to revolutionize access to education. In this paper, we describe experiences with MOOCs for Liberian students who are connected to the iLab technology hub. We describe their motivations for participating as well as the challenges they encountered. We also describe the importance of the face-to-face learning environment provided by the iLab as a source of community support. Michael A. Madaio, Rebecca E. Grinter, Ellen Zegura |
ICTD | 1 |
| 2016 | The Effect of Friendship and Tutoring Roles on Reciprocal Peer Tutoring Strategies
Michael A. Madaio, Amy Ogan, Justine Cassell |
ITS | 1 |
| 2016 | Firebird: Predicting Fire Risk and Prioritizing Fire Inspections in AtlantaabstractThe Atlanta Fire Rescue Department (AFRD), like many municipal fire departments, actively works to reduce fire risk by inspecting commercial properties for potential hazards and fire code violations. However, AFRD's fire inspection practices relied on tradition and intuition, with no existing data-driven process for prioritizing fire inspections or identifying new properties requiring inspection. In collaboration with AFRD, we developed the Firebird framework to help municipal fire departments identify and prioritize commercial property fire inspections, using machine learning, geocoding, and information visualization. Firebird computes fire risk scores for over 5,000 buildings in the city, with true positive rates of up to 71% in predicting fires. It has identified 6,096 new potential commercial properties to inspect, based on AFRD's criteria for inspection. Furthermore, through an interactive map, Firebird integrates and visualizes fire incidents, property information and risk scores to help AFRD make informed decisions about fire inspections. Firebird has already begun to make positive impact at both local and national levels. It is improving AFRD's inspection processes and Atlanta residents' safety, and was highlighted by National Fire Protection Association (NFPA) as a best practice for using data to inform fire inspections. Michael A. Madaio, Shang-Tse Chen, Oliver L. Haimson, Xiang Cheng 0002, Matthew Hinds-Aldrich, Polo Chau, Bistra Dilkina |
KDD | 1 |
| 2015 | Beyond bootstrapping: the liberian ilab as a maturing community of practiceabstractWhile Information and Communication Technology (ICT) access in Liberia is still low, use of PCs, mobile phones and the Internet is rising. A relatively recent option for learning ICT skills is Liberia's iLab technology hub founded in 2011 to encourage and support a local technology community. We have partnered with the iLab since its founding to offer summer courses taught by a student instructor. In this paper, we describe how the iLab's community-based learning approach has advanced from the coalescing stage, identified in earlier research, towards the maturing stage, despite obstacles. We use Lave and Wenger's Community of Practice (CoP) framework as the analytic structure to present results from our data consisting primarily of interviews with course participants. We contribute to research on technology hubs in Africa and highlight the role of community-based learning for technical skill acquisition even in resource challenged settings. Ellen Zegura, Michael A. Madaio, Rebecca E. Grinter |
ICTD | 2 |