EDBT 2026 Demo / reviewers in the wild / expert
Kenneth Holstein
dblp:176/0446 · also Ken Holstein
· DBLP profile ↗
58ranked-venue papers
10as first author
40since 2021 · last 2026
0000-0001-6730-922XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 50 · 10 first-author · 37 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 7 first-author · 3 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AI Design Sprints: Facilitating AI Innovation within Cross-functional Industry TeamsabstractArtificial intelligence (AI) technologies offer tremendous potential for product and service innovation, yet finding good use cases remains challenging. Currently, AI projects largely fail due to breakdowns in early stage ideation and problem formulation. Drawing on HCI research that used AI capabilities and examples to facilitate AI concept ideation, this paper investigates how these approaches might be operationalized in industry settings. We collaborated with cross-functional industry teams in insurance, accounting, and consultancy. We conducted a series of AI Design Sprints, where innovators simultaneously consider AI capabilities and human needs. All teams perceived the ideation method highly valuable both for rapidly exploring use cases and building AI literacy within teams. We detail our process, the challenges, and artifacts that proved effective. We share insights on how AI projects get initiated, and how innovation teams identify use cases. Reflecting on these case studies, we discuss opportunities for improving early stage AI innovation. Nur Yildirim, Kayur Patel, Florian Dusch, Dennis Knopf, Melike Yusufoglu, Dominik Schuler, Kenneth Holstein, Jodi Forlizzi, James McCann, John Zimmerman |
DIS | 7 |
| 2026 | Using an MR-Based Teacher Orchestration Tool in AI-Supported K-12 Classrooms
Qiao Jin 0002, Will Morgus, Kyle Price, Michael Sandbothe, Jonathan Sewall, Octav Popescu, Susan Berman, Stephen Fancsali, Steven Ritter 0001, Kenneth Holstein, Vincent Aleven |
AIED (5) | 11 |
| 2026 | PolicyPad: Collaborative Prototyping of LLM PoliciesabstractAs LLMs gain adoption in high-stakes domains like mental health, domain experts are increasingly consulted to provide input into policies governing their behavior. From an observation of 19 policymaking workshops with 9 experts over 15 weeks, we identified opportunities to better support rapid experimentation, feedback, and iteration for collaborative policy design processes. We present PolicyPad, an interactive system that facilitates the emerging practice of LLM policy prototyping by drawing from established UX prototyping practices, including heuristic evaluation and storyboarding. Using PolicyPad, policy designers can collaborate on drafting a policy in real time while independently testing policy-informed model behavior with usage scenarios. We evaluate PolicyPad through workshops with 8 groups of 22 domain experts in mental health and law, finding that PolicyPad enhanced collaborative dynamics during policy design, enabled tight feedback loops, and led to novel policy contributions. Overall, our work paves expert-informed paths for advancing AI alignment and safety. K. J. Kevin Feng, Tzu-Sheng Kuo, Quan Ze Chen, Inyoung Cheong, Kenneth Holstein, Amy X. Zhang |
CHI | 5 |
| 2026 | Prototyping Multimodal GenAI Real-Time Agents with Counterfactual Replays and Hybrid Wizard-of-OzabstractRecent advancements in multimodal generative AI (GenAI) enable the creation of personal context-aware real-time agents that, for example, can augment user workflows by following their on-screen activities and providing contextual assistance. However, prototyping such experiences is challenging, especially when supporting people with domain-specific tasks using real-time inputs such as speech and screen recordings. While prototyping an LLM-based proactive support agent system, we found that existing prototyping and evaluation methods were insufficient to anticipate the nuanced situational complexity and contextual immediacy required. To overcome these challenges, we explored a novel user-centered prototyping approach that combines counterfactual video replay prompting and hybrid Wizard of Oz methods to iteratively design and refine agent behaviors. This paper discusses our prototyping experiences, highlighting successes and limitations, and offers a practical guide and an open-source toolkit for UX designers, HCI researchers, and AI toolmakers to build more user-centered and context-aware multimodal agents. Frederic Gmeiner, Kenneth Holstein, Nikolas Martelaro |
CHI | 2 |
| 2026 | Botender: Supporting Communities in Collaboratively Designing AI Agents through Case-Based ProvocationsabstractAI agents, or bots, serve important roles in online communities. However, they are often designed by outsiders or a few tech-savvy members, leading to bots that may not align with the broader community’s needs. How might communities collectively shape the behavior of community bots? We present Botender, a system that enables communities to collaboratively design LLM-powered bots without coding. With Botender, community members can directly propose, iterate on, and deploy custom bot behaviors tailored to community needs. Botender facilitates testing and iteration on bot behavior through case-based provocations: interaction scenarios generated to spark user reflection and discussion around desirable bot behavior. A validation study found these provocations more useful than standard test cases for revealing improvement opportunities and surfacing disagreements. During a five-day deployment across six Discord servers, Botender supported communities in tailoring bot behavior to their specific needs, showcasing the usefulness of case-based provocations in facilitating collaborative bot design. Tzu-Sheng Kuo, Sophia Liu, Quan Ze Chen, Joseph Seering, Amy X. Zhang, Haiyi Zhu, Kenneth Holstein |
CHI | 7 |
| 2026 | Funding AI for Good: A Call for Meaningful EngagementabstractArtificial Intelligence for Social Good (AI4SG) is a growing area that explores AI’s potential to address social issues, such as public health. Yet prior work has shown limited evidence of its tangible benefits for intended communities, and projects frequently face real-world deployment and sustainability challenges. While existing HCI literature on AI4SG initiatives primarily focuses on the mechanisms of funded projects and their outcomes, much less attention has been given to the upstream funding agendas that influence project approaches. In this work, we conducted a reflexive thematic analysis of 35 funding documents, representing about $410 million USD in total investments. We uncovered a spectrum of conceptual framings of AI4SG and the approaches that funding rhetoric promoted: from biasing towards technology capacities (more techno-centric) to emphasizing contextual understanding of the social problems at hand alongside technology capacities (more balanced). Drawing on our findings on how funding documents construct AI4SG, we offer recommendations for funders to embed more balanced approaches in future funding call designs. We further discuss implications for how the HCI community can positively shape AI4SG funding design processes. Hongjin Lin, Anna Kawakami, Catherine D'Ignazio, Kenneth Holstein, Krzysztof Z. Gajos |
CHI | 4 |
| 2026 | "I Don't Think RAI Applies to My Model" - Engaging Non-champions with Sticky Stories for Responsible AI WorkabstractResponsible AI (RAI) tools—checklists, templates, and governance processes—often engage RAI champions, individuals intrinsically motivated to advocate ethical practices, but fail to reach non-champions, who frequently dismiss them as bureaucratic tasks. To explore this gap, we shadowed meetings and interviewed data scientists at an organization, finding that practitioners perceived RAI as irrelevant to their work. Building on these insights and theoretical foundations, we derived design principles for engaging non-champions, and introduced sticky stories—narratives of unexpected ML harms designed to be concrete, severe, surprising, diverse, and relevant, unlike widely circulated media to which practitioners are desensitized. Using a compound AI system, we generated and evaluated sticky stories through human and LLM assessments at scale, confirming they embodied the intended qualities. In a study with 29 practitioners, we found that, compared to regular stories, sticky stories significantly increased the engagement time on harm identification, broadened the range of harms recognized, and fostered deeper reflection. Nadia Nahar, Chenyang Yang 0002, Yanxin Chen, Wesley Deng, Kenneth Holstein, Motahhare Eslami, Christian Kästner |
CHI | 5 |
| 2026 | Don't Be Fooled: The Misinformation Effect of Explanations in Human-AI CollaborationabstractAcross various applications, humans increasingly use black-box artificial intelligence (AI) systems without insight into these systems’ reasoning. To counter this opacity, explainable AI (XAI) methods promise enhanced transparency and interpretability. While recent studies have explored how XAI affects human–AI collaboration, few have examined the potential pitfalls caused by incorrect explanations. The implications for humans can be far-reaching but have not been explored extensively. To investigate this, we conducted a study (n = 160) on AI-assisted decision-making in which humans were supported by XAI. Our findings reveal a misinformation effect when incorrect explanations accompany correct AI advice with implications post-collaboration. This effect causes humans to infer flawed reasoning strategies, hindering task execution and demonstrating impaired procedural knowledge. Additionally, incorrect explanations compromise human–AI team performance during collaboration. With our work, we contribute to HCI by providing empirical evidence for the negative consequences of incorrect explanations on humans post-collaboration and outlining guidelines for designers of AI. Philipp Spitzer, Joshua Holstein, Katelyn Morrison, Kenneth Holstein, Gerhard Satzger, Niklas Kühl 0001 |
Int. J. Hum. Comput. Interact. | 4 |
| 2025 | Exploring the Potential of Metacognitive Support Agents for Human-AI Co-CreationabstractDespite the potential of generative AI (GenAI) design tools to enhance design processes, professionals often struggle to integrate AI into their workflows. Fundamental cognitive challenges include the need to specify all design criteria as distinct parameters upfront (intent formulation) and designers' reduced cognitive involvement in the design process due to cognitive offloading, which can lead to insufficient problem exploration, underspecification, and limited ability to evaluate outcomes. Motivated by these challenges, we envision novel metacognitive support agents that assist designers in working more reflectively with GenAI. To explore this vision, we conducted exploratory prototyping through a Wizard of Oz elicitation study with 20 mechanical designers probing multiple metacognitive support strategies. We found that agent-supported users created more feasible designs than non-supported users, with differing impacts between support strategies. Based on these findings, we discuss opportunities and tradeoffs of metacognitive support agents and considerations for future AI-based design tools. Frederic Gmeiner, Kaitao Luo, Kenneth Holstein, Nikolas Martelaro |
Conference on Designing Interactive Systems | 4 |
| 2025 | Making the Right Thing: Bridging HCI and Responsible AI in Early-Stage AI Concept SelectionabstractAI projects often fail due to financial, technical, ethical, or user acceptance challenges-failures frequently rooted in early-stage decisions.While HCI and Responsible AI (RAI) research emphasize this, practical approaches for identifying promising concepts early remain limited.Drawing on Research through Design, this paper investigates how early-stage AI concept sorting in commercial settings can reflect RAI principles.Through three design experiments-including a probe study with industry practitioners-we explored methods for evaluating risks and benefits using multidisciplinary collaboration.Participants demonstrated strong receptivity to addressing RAI concerns early in the process and effectively identified low-risk, high-benefit AI concepts.Our findings highlight the potential of a design-led approach to embed ethical and service design thinking at the front end of AI innovation.By examining how practitioners reason about AI concepts, our study invites HCI and RAI communities to see early-stage innovation as a critical space for engaging ethical and commercial considerations together. Ji-Youn Jung, Devansh Saxena, Minjung Park, Jini Kim, Jodi Forlizzi, Kenneth Holstein, John Zimmerman |
Conference on Designing Interactive Systems | 6 |
| 2025 | Intent Tagging: Exploring Micro-Prompting Interactions for Supporting Granular Human-GenAI Co-Creation WorkflowsabstractDespite Generative AI (GenAI) systems' potential for enhancing content creation, users often struggle to effectively integrate GenAI into their creative workflows. Core challenges include misalignment of AI-generated content with user intentions (intent elicitation and alignment), user uncertainty around how to best communicate their intents to the AI system (prompt formulation), and insufficient flexibility of AI systems to support diverse creative workflows (workflow flexibility). Motivated by these challenges, we created IntentTagger: a system for slide creation based on the notion of Intent Tags - small, atomic conceptual units that encapsulate user intent - for exploring granular and non-linear micro-prompting interactions for Human-GenAI co-creation workflows. Our user study with 12 participants provides insights into the value of flexibly expressing intent across varying levels of ambiguity, meta-intent elicitation, and the benefits and challenges of intent tag-driven workflows. We conclude by discussing the broader implications of our findings and design considerations for GenAI-supported content creation workflows. Frederic Gmeiner, Nicolai Marquardt, Michael Bentley, Hugo Romat, Michel Pahud, Asta Roseway, Nikolas Martelaro, Kenneth Holstein, Ken Hinckley, Nathalie Henry Riche |
CHI | 9 |
| 2025 | PolicyCraft: Supporting Collaborative and Participatory Policy Design through Case-Grounded Deliberation
Tzu-Sheng Kuo, Quan Ze Chen, Amy X. Zhang, Jane Hsieh, Haiyi Zhu, Kenneth Holstein |
CHI | 6 |
| 2025 | AI Mismatches: Identifying Potential Algorithmic Harms Before AI Development
Devansh Saxena, Ji-Youn Jung, Jodi Forlizzi, Kenneth Holstein, John Zimmerman |
CHI | 4 |
| 2025 | Validating LLM-as-a-Judge Systems under Rating IndeterminacyabstractThe LLM-as-a-judge paradigm, in which a judge LLM system replaces human raters in rating the outputs of other generative AI (GenAI) systems, plays a critical role in scaling and standardizing GenAI evaluations. To validate such judge systems, evaluators assess human--judge agreement by first collecting multiple human ratings for each item in a validation corpus, then aggregating the ratings into a single, per-item gold label rating. For many items, however, rating criteria may admit multiple valid interpretations, so a human or LLM rater may deem multiple ratings "reasonable" or "correct." We call this condition rating indeterminacy. Problematically, many rating tasks that contain rating indeterminacy rely on forced-choice elicitation, whereby raters are instructed to select only one rating for each item. In this paper, we introduce a framework for validating LLM-as-a-judge systems under rating indeterminacy. We draw theoretical connections between different measures of judge system performance under different human--judge agreement metrics, and different rating elicitation and aggregation schemes. We demonstrate that differences in how humans and LLMs resolve rating indeterminacy when responding to forced-choice rating instructions can heavily bias LLM-as-a-judge validation. Through extensive experiments involving 11 real-world rating tasks and 9 commercial LLMs, we show that standard validation approaches that rely upon forced-choice ratings select judge systems that are highly suboptimal, performing as much as 31% worse than judge systems selected by our approach that uses multi-label "response set" ratings to account for rating indeterminacy. We conclude with concrete recommendations for more principled approaches to LLM-as-a-judge validation. Luke Guerdan, Solon Barocas, Kenneth Holstein, Hanna M. Wallach, Steven Z. Wu, Alexandra Chouldechova |
NeurIPS | 3 |
| 2025 | Policy Maps: Tools for Guiding the Unbounded Space of LLM BehaviorsabstractFigure 1: Policy maps chart LLM policy coverage over an unbounded space of model behaviors.Here, an AI practitioner is designing a policy for how an LLM should summarize violent text.Policy map abstractions (right) allow the policy designer to interactively author and test policies that govern a model's behavior using if-then rules over concepts.The designer can create any desired concept by providing a simple text definition to capture cases of model behavior.Our Policy Projector tool (center) renders cases, concepts, and policies as visual map layers to aid iterative policy design. Michelle S. Lam, Fred Hohman, Dominik Moritz, Jeffrey P. Bigham, Kenneth Holstein, Mary Beth Kery |
UIST | 5 |
| 2025 | WeAudit: Scaffolding User Auditors and AI Practitioners in Auditing Generative AIabstractThere has been growing interest from both practitioners and researchers in engaging end users in AI auditing, to draw upon users' unique knowledge and lived experiences. However, we know little about how to effectively scaffold end users in auditing in ways that can generate actionable insights for AI practitioners. Through formative studies with both users and AI practitioners, we first identified a set of design goals to support user-engaged AI auditing. We then developed WeAudit, a workflow and system that supports end users in auditing AI both individually and collectively. We evaluated WeAudit through a three-week user study with user auditors and interviews with industry Generative AI practitioners. Our findings offer insights into how WeAudit supports users in noticing and reflecting upon potential AI harms and in articulating their findings in ways that industry practitioners can act upon. Based on our observations and feedback from both users and practitioners, we identify several opportunities to better support user engagement in AI auditing processes. We discuss implications for future research to support effective and responsible user engagement in AI auditing. Wesley Deng, Claire Wang 0002, Howard Ziyu Han, Jason I. Hong, Kenneth Holstein, Motahhare Eslami |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2025 | Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling TasksabstractData scientists often formulate predictive modeling tasks involving fuzzy, hard-to-define concepts, such as the ''authenticity'' of student writing or the ''healthcare need'' of a patient. Yet the process by which data scientists translate fuzzy concepts into a concrete, proxy target variable remains poorly understood. We interview fifteen data scientists in education (N=8) and healthcare (N=7) to understand how they construct target variables for predictive modeling tasks. Our findings suggest that data scientists construct target variables through a bricolage process, in which they use creative and pragmatic approaches to make do with the limited data at hand. Data scientists attempt to satisfy five major criteria for a target variable through bricolage: validity, simplicity, predictability, portability, and resource requirements. To achieve this, data scientists adaptively apply problem (re)formulation strategies, such as swapping out one candidate target variable for another when the first fails to meet certain criteria (e.g., predictability), or composing multiple outcomes into a single target variable to capture a more holistic set of modeling objectives. Based on our findings, we present opportunities for future HCI, CSCW, and ML research to better support the art and science of target variable construction. Luke Guerdan, Devansh Saxena, Stevie Chancellor, Steven Z. Wu, Kenneth Holstein |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2024 | AI Failure Loops in Feminized Labor: Understanding the Interplay of Workplace AI and Occupational DevaluationabstractA growing body of literature has focused on understanding and addressing workplace AI design failures. However, past work has largely overlooked the role of occupational devaluation in shaping the dynamics of AI development and deployment. In this paper, we examine the case of feminized labor: a class of devalued occupations historically misnomered as ``women's work,'' such as social work, K-12 teaching, and home healthcare. Drawing on literature on AI deployments in feminized labor contexts, we conceptualize AI Failure Loops: a set of interwoven, socio-technical failures that help explain how the systemic devaluation of workers' expertise negatively impacts, and is impacted by, AI design, evaluation, and governance practices. These failures demonstrate how misjudgments on the automatability of workers' skills can lead to AI deployments that fail to bring value and, instead, further diminish the visibility of workers' expertise. We discuss research and design implications for workplace AI, especially for devalued occupations. Anna Kawakami, Jordan Taylor, Sarah E. Fox, Haiyi Zhu, Kenneth Holstein |
AIES (1) | 5 |
| 2024 | The Situate AI Guidebook: Co-Designing a Toolkit to Support Multi-Stakeholder, Early-stage Deliberations Around Public Sector AI ProposalsabstractPublic sector agencies are rapidly deploying AI systems to augment or automate critical decisions in real-world contexts like child welfare, criminal justice, and public health. A growing body of work documents how these AI systems often fail to improve services in practice. These failures can often be traced to decisions made during the early stages of AI ideation and design, such as problem formulation. However, today, we lack systematic processes to support effective, early-stage decision-making about whether and under what conditions to move forward with a proposed AI project. To understand how to scaffold such processes in real-world settings, we worked with public sector agency leaders, AI developers, frontline workers, and community advocates across four public sector agencies and three community advocacy groups in the United States. Through an iterative co-design process, we created the Situate AI Guidebook: a structured process centered around a set of deliberation questions to scaffold conversations around (1) goals and intended use for a proposed AI system, (2) societal and legal considerations, (3) data and modeling constraints, and (4) organizational governance factors. We discuss how the guidebook’s design is informed by participants’ challenges, needs, and desires for improved deliberation processes. We further elaborate on implications for designing responsible AI toolkits in collaboration with public sector agency stakeholders and opportunities for future work to expand upon the guidebook. This design approach can be more broadly adopted to support the co-creation of responsible AI toolkits that scaffold key decision-making processes surrounding the use of AI in the public sector and beyond. Anna Kawakami, Amanda Coston, Haiyi Zhu, Hoda Heidari, Kenneth Holstein |
CHI | 5 |
| 2024 | Wikibench: Community-Driven Data Curation for AI Evaluation on WikipediaabstractAI tools are increasingly deployed in community contexts. However, datasets used to evaluate AI are typically created by developers and annotators outside a given community, which can yield misleading conclusions about AI performance. How might we empower communities to drive the intentional design and curation of evaluation datasets for AI that impacts them? We investigate this question on Wikipedia, an online community with multiple AI-based content moderation tools deployed. We introduce Wikibench, a system that enables communities to collaboratively curate AI evaluation datasets, while navigating ambiguities and differences in perspective through discussion. A field study on Wikipedia shows that datasets curated using Wikibench can effectively capture community consensus, disagreement, and uncertainty. Furthermore, study participants used Wikibench to shape the overall data curation process, including refining label definitions, determining data inclusion criteria, and authoring data statements. Based on our findings, we propose future directions for systems that support community-driven data curation. Tzu-Sheng Kuo, Aaron Halfaker, Zirui Cheng, Meng-Hsin Wu, Sherry Tongshuang Wu, Kenneth Holstein, Haiyi Zhu |
CHI | 7 |
| 2024 | Predictive Performance Comparison of Decision Policies Under ConfoundingabstractPredictive models are often introduced to decision-making tasks under the rationale that they improve performance over an existing decision-making policy. However, it is challenging to compare predictive performance against an existing decision-making policy that is generally under-specified and dependent on unobservable factors. These sources of uncertainty are often addressed in practice by making strong assumptions about the data-generating mechanism. In this work, we propose a method to compare the predictive performance of decision policies under a variety of modern identification approaches from the causal inference and off-policy evaluation literatures (e.g., instrumental variable, marginal sensitivity model, proximal variable). Key to our method is the insight that there are regions of uncertainty that we can safely ignore in the policy comparison. We develop a practical approach for finite-sample estimation of regret intervals under no assumptions on the parametric form of the status quo policy. We verify our framework theoretically and via synthetic data experiments. We conclude with a real-world application using our framework to support a pre-deployment evaluation of a proposed modification to a healthcare enrollment policy. Luke Guerdan, Amanda Coston, Kenneth Holstein, Steven Z. Wu |
ICML | 3 |
| 2024 | Studying Up Public Sector AI: How Networks of Power Relations Shape Agency Decisions Around AI Design and UseabstractAs public sector agencies rapidly introduce new AI tools in high-stakes domains like social services, it becomes critical to understand how decisions to adopt these tools are made in practice. We borrow from the anthropological practice to "study up" those in positions of power, and reorient our study of public sector AI around those who have the power and responsibility to make decisions about the role that AI tools will play in their agency. Through semi-structured interviews and design activities with 16 agency decision-makers, we examine how decisions about AI design and adoption are influenced by their interactions with and assumptions about other actors within these agencies (e.g., frontline workers and agency leaders), as well as those above (legal systems and contracted companies), and below (impacted communities). By centering these networks of power relations, our findings shed light on how infrastructural, legal, and social factors create barriers and disincentives to the involvement of a broader range of stakeholders in decisions about AI design and adoption. Agency decision-makers desired more practical support for stakeholder involvement around public sector AI to help overcome the knowledge and power differentials they perceived between them and other stakeholders (e.g., frontline workers and impacted community members). Building on these findings, we discuss implications for future research and policy around actualizing participatory AI approaches in public sector contexts. Anna Kawakami, Amanda Coston, Hoda Heidari, Kenneth Holstein, Haiyi Zhu |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2024 | Carefully Unmaking the "Marginalized User": A Diffractive Analysis of a Gay Online CommunityabstractHCI scholars are increasingly engaging in research about “marginalized groups,” such as LGBTQ+ people. While normative habitual readings of marginalized people in HCI often highlight real problems, this work has been criticized for flattening heterogeneous experiences and overemphasizing harms. Some have advocated for expanding how we approach research on marginalized people (e.g., assets-based design, the everyday, and joy). Sensitized by unmaking literature, we explore this tension between conditions, experiences, and representations of marginality in HCI scholarship. To do so, we perform a diffractive analysis of posts in a gay online community by bringing two readings of the same data together: a normative habitual reading of marginalization and an expanded reading. By examining the relationship between empirical material and its representations by HCI researchers, we explore how to carefully unmake HCI research, thus maintaining and repairing our research community. We discuss the political and designerly implications of different readings of marginalized people and offer considerations for attending to the processes and afterlives of HCI research. Jordan Taylor, Wesley Deng, Kenneth Holstein, Sarah E. Fox, Haiyi Zhu |
ACM Trans. Comput. Hum. Interact. | 3 |
| 2023 | Zeno: An Interactive Framework for Behavioral Evaluation of Machine LearningabstractMachine learning models with high accuracy on test data can still produce systematic failures, such as harmful biases and safety issues, when deployed in the real world. To detect and mitigate such failures, practitioners run behavioral evaluation of their models, checking model outputs for specific types of inputs. Behavioral evaluation is important but challenging, requiring that practitioners discover real-world patterns and validate systematic failures. We conducted 18 semi-structured interviews with ML practitioners to better understand the challenges of behavioral evaluation and found that it is a collaborative, use-case-first process that is not adequately supported by existing task- and domain-specific tools. Using these findings, we designed zeno, a general-purpose framework for visualizing and testing AI systems across diverse use cases. In four case studies with participants using zeno on real-world models, we found that practitioners were able to reproduce previous manual analyses and discover new systematic failures. Ángel Alexander Cabrera, Erica Fu, Donald Bertucci, Kenneth Holstein, Ameet Talwalkar, Jason I. Hong, Adam Perer |
CHI | 4 |
| 2023 | Understanding Practices, Challenges, and Opportunities for User-Engaged Algorithm Auditing in Industry PracticeabstractRecent years have seen growing interest among both researchers and practitioners in user-engaged approaches to algorithm auditing, which directly engage users in detecting problematic behaviors in algorithmic systems. However, we know little about industry practitioners’ current practices and challenges around user-engaged auditing, nor what opportunities exist for them to better leverage such approaches in practice. To investigate, we conducted a series of interviews and iterative co-design activities with practitioners who employ user-engaged auditing approaches in their work. Our findings reveal several challenges practitioners face in appropriately recruiting and incentivizing user auditors, scaffolding user audits, and deriving actionable insights from user-engaged audit reports. Furthermore, practitioners shared organizational obstacles to user-engaged auditing, surfacing a complex relationship between practitioners and user auditors. Based on these findings, we discuss opportunities for future HCI research to help realize the potential (and mitigate risks) of user-engaged auditing in industry practice. Wesley Deng, Bill Boyuan Guo, Alicia DeVrio, Hong Shen 0004, Motahhare Eslami, Kenneth Holstein |
CHI | 6 |
| 2023 | Exploring Challenges and Opportunities to Support Designers in Learning to Co-create with AI-based Manufacturing Design ToolsabstractAI-based design tools are proliferating in professional software to assist engineering and industrial designers in complex manufacturing and design tasks. These tools take on more agentic roles than traditional computer-aided design tools and are often portrayed as “co-creators.” Yet, working effectively with such systems requires different skills than working with complex CAD tools alone. To date, we know little about how engineering designers learn to work with AI-based design tools. In this study, we observed trained designers as they learned to work with two AI-based tools on a realistic design task. We find that designers face many challenges in learning to effectively co-create with current systems, including challenges in understanding and adjusting AI outputs and in communicating their design goals. Based on our findings, we highlight several design opportunities to better support designer-AI co-creation. Frederic Gmeiner, Humphrey Yang, Lining Yao, Kenneth Holstein, Nikolas Martelaro |
CHI | 4 |
| 2023 | Understanding Frontline Workers' and Unhoused Individuals' Perspectives on AI Used in Homeless ServicesabstractRecent years have seen growing adoption of AI-based decision-support systems (ADS) in homeless services, yet we know little about stakeholder desires and concerns surrounding their use. In this work, we aim to understand impacted stakeholders’ perspectives on a deployed ADS that prioritizes scarce housing resources. We employed AI lifecycle comicboarding, an adapted version of the comicboarding method, to elicit stakeholder feedback and design ideas across various components of an AI system’s design. We elicited feedback from county workers who operate the ADS daily, service providers whose work is directly impacted by the ADS, and unhoused individuals in the region. Our participants shared concerns and design suggestions around the AI system’s overall objective, specific model design choices, dataset selection, and use in deployment. Our findings demonstrate that stakeholders, even without AI knowledge, can provide specific and critical feedback on an AI system’s design and deployment, if empowered to do so. Tzu-Sheng Kuo, Hong Shen 0004, Jisoo Geum, Nev Jones, Jason I. Hong, Haiyi Zhu, Kenneth Holstein |
CHI | 7 |
| 2023 | Pair-Up: Prototyping Human-AI Co-orchestration of Dynamic Transitions between Individual and Collaborative Learning in the ClassroomabstractEnabling students to dynamically transition between individual and collaborative learning activities has great potential to support better learning. We explore how technology can support teachers in orchestrating dynamic transitions during class. Working with five teachers and 199 students over 22 class sessions, we conducted classroom-based prototyping of a co-orchestration technology ecosystem that supports the dynamic pairing of students working with intelligent tutoring systems. Using mixed-methods data analysis, we study the resulting observed classroom dynamics, and how teachers and students perceived and experienced dynamic transitions as supported by our technology. We discover a potential tension between teachers’ and students’ preferred level of control: students prefer a degree of control over the dynamic transitions that teachers are hesitant to grant. Our study reveals design implications and challenges for future human-AI co-orchestration in classroom use, bringing us closer to realizing the vision of highly-personalized smart classrooms that address the unique needs of each student. Kexin Bella Yang, Vanessa Echeverría, Zijing Lu, Hongyu Mao, Kenneth Holstein, Nikol Rummel, Vincent Aleven |
CHI | 5 |
| 2023 | Toward Supporting Perceptual Complementarity in Human-AI Collaboration via Reflection on UnobservablesabstractIn many real world contexts, successful human-AI collaboration requires humans to productively integrate complementary sources of information into AI-informed decisions. However, in practice human decision-makers often lack understanding of what information an AI model has access to, in relation to themselves. There are few available guidelines regarding how to effectively communicate aboutunobservables: features that may influence the outcome, but which are unavailable to the model. In this work, we conducted an online experiment to understand whether and how explicitly communicating potentially relevant unobservables influences how people integrate model outputs and unobservables when making predictions. Our findings indicate that presenting prompts about unobservables can change how humans integrate model outputs and unobservables, but do not necessarily lead to improved performance. Furthermore, the impacts of these prompts can vary depending on decision-makers' prior domain expertise. We conclude by discussing implications for future research and design of AI-based decision support tools. Kenneth Holstein, Maria De-Arteaga, Lakshmi Tumati, Yanghuidi Cheng |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2023 | Evaluating the Impact of Human Explanation Strategies on Human-AI Visual Decision-MakingabstractArtificial intelligence (AI) is increasingly being deployed in high-stakes domains, such as disaster relief and radiology, to aid practitioners during the decision-making process. Explainable AI techniques have been developed and deployed to provide users insights into why the AI made certain predictions. However, recent research suggests that these techniques may confuse or mislead users. We conducted a series of two studies to uncover strategies that humans use to explain decisions and then understand how those explanation strategies impact visual decision-making. In our first study, we elicit explanations from humans when assessing and localizing damaged buildings after natural disasters from satellite imagery and identify four core explanation strategies that humans employed. We then follow up by studying the impact of these explanation strategies by framing the explanations from Study 1 as if they were generated by AI and showing them to a different set of decision-makers performing the same task. We provide initial insights on how causal explanation strategies improve humans' accuracy and calibrate humans' reliance on AI when the AI is incorrect. However, we also find that causal explanation strategies may lead to incorrect rationalizations when AI presents a correct assessment with incorrect localization. We explore the implications of our findings for the design of human-centered explainable AI and address directions for future work. Katelyn Morrison, Kenneth Holstein, Adam Perer |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2022 | "Why Do I Care What's Similar?" Probing Challenges in AI-Assisted Child Welfare Decision-Making through Worker-AI Interface Design ConceptsabstractData-driven AI systems are increasingly used to augment human decision-making in complex, social contexts, such as social work or legal practice. Yet, most existing design knowledge regarding how to best support AI-augmented decision-making comes from studies in comparatively well-defined settings. In this paper, we present findings from design interviews with 12 social workers who use an algorithmic decision support tool (ADS) to assist their day-to-day child maltreatment screening decisions. We generated a range of design concepts, each envisioning different ways of redesigning or augmenting the ADS interface. Overall, workers desired ways to understand the risk score and incorporate contextual knowledge, which move beyond existing notions of AI interpretability. Conversations around our design concepts also surfaced more fundamental concerns around the assumptions underlying statistical prediction, such as inference based on similar historical cases and statistical notions of uncertainty. Based on our findings, we discuss how ADS may be better designed to support the roles of human decision-makers in social decision-making contexts. Anna Kawakami, Venkatesh Sivaraman, Logan Stapleton, Hao Fei Cheng, Adam Perer, Steven Z. Wu, Haiyi Zhu, Kenneth Holstein |
Conference on Designing Interactive Systems | 8 |
| 2022 | "A Second Voice": Investigating Opportunities and Challenges for Interactive Voice Assistants to Support Home Health AidesabstractHome health aides are a vulnerable group of frontline caregivers who provide personal and medically-oriented care in patients’ homes. Their work is difficult and unpredictable, involving a mix of physical and emotional labor as they adapt to patients’ changing needs. Our paper presents an exploratory, qualitative study with 32 participants, that investigates design opportunities for Interactive Voice Assistants (IVAs) to support aides’ essential care work. We explore challenges and opportunities for IVAs to (1) fill gaps in aides’ access to information and care coordination, (2) assist with decision making and task completion, (3) advocate on behalf of aides, and (4) provide emotional support. We then discuss key implications of our work, including how materiality may impact perceived ownership and usage of IVAs, the need to carefully consider tensions around surveillance, accountability, data collection, and reporting, and the challenges of centering aides as essential workers in complex home health care contexts. Vince Bartle, Janice Lyu, Freesoul El Shabazz-Thompson, Yunmin Oh, Angela Anqi Chen, Yu-Jan Chang, Kenneth Holstein, Nicola Dell |
CHI | 7 |
| 2022 | How Child Welfare Workers Reduce Racial Disparities in Algorithmic DecisionsabstractMachine learning tools have been deployed in various contexts to support human decision-making, in the hope that human-algorithm collaboration can improve decision quality. However, the question of whether such collaborations reduce or exacerbate biases in decision-making remains underexplored. In this work, we conducted a mixed-methods study, analyzing child welfare call screen workers’ decision-making over a span of four years, and interviewing them on how they incorporate algorithmic predictions into their decision-making process. Our data analysis shows that, compared to the algorithm alone, workers reduced the disparity in screen-in rate between Black and white children from 20% to 9%. Our qualitative data show that workers achieved this by making holistic risk assessments and adjusting for the algorithm’s limitations. Our analyses also show more nuanced results about how human-algorithm collaboration affects prediction accuracy, and how to measure these effects. These results shed light on potential mechanisms for improving human-algorithm collaboration in high-risk decision-making contexts. Hao Fei Cheng, Logan Stapleton, Anna Kawakami, Venkatesh Sivaraman, Yanghuidi Cheng, Diana Qing, Adam Perer, Kenneth Holstein, Steven Z. Wu, Haiyi Zhu |
CHI | 8 |
| 2022 | Toward User-Driven Algorithm Auditing: Investigating users' strategies for uncovering harmful algorithmic behaviorabstractRecent work in HCI suggests that users can be powerful in surfacing harmful algorithmic behaviors that formal auditing approaches fail to detect. However, it is not well understood how users are often able to be so effective, nor how we might support more effective user-driven auditing. To investigate, we conducted a series of think-aloud interviews, diary studies, and workshops, exploring how users find and make sense of harmful behaviors in algorithmic systems, both individually and collectively. Based on our findings, we present a process model capturing the dynamics of and influences on users’ search and sensemaking behaviors. We find that 1) users’ search strategies and interpretations are heavily guided by their personal experiences with and exposures to societal bias; and 2) collective sensemaking amongst multiple users is invaluable in user-driven algorithm audits. We offer directions for the design of future methods and tools that can better support user-driven auditing. Alicia DeVos, Aditi Dhabalia, Hong Shen 0004, Kenneth Holstein, Motahhare Eslami |
CHI | 4 |
| 2022 | Improving Human-AI Partnerships in Child Welfare: Understanding Worker Practices, Challenges, and Desires for Algorithmic Decision SupportabstractAI-based decision support tools (ADS) are increasingly used to augment human decision-making in high-stakes, social contexts. As public sector agencies begin to adopt ADS, it is critical that we understand workers’ experiences with these systems in practice. In this paper, we present findings from a series of interviews and contextual inquiries at a child welfare agency, to understand how they currently make AI-assisted child maltreatment screening decisions. Overall, we observe how workers’ reliance upon the ADS is guided by (1) their knowledge of rich, contextual information beyond what the AI model captures, (2) their beliefs about the ADS’s capabilities and limitations relative to their own, (3) organizational pressures and incentives around the use of the ADS, and (4) awareness of misalignments between algorithmic predictions and their own decision-making objectives. Drawing upon these findings, we discuss design implications towards supporting more effective human-AI decision-making. Anna Kawakami, Venkatesh Sivaraman, Hao Fei Cheng, Logan Stapleton, Yanghuidi Cheng, Diana Qing, Adam Perer, Steven Z. Wu, Haiyi Zhu, Kenneth Holstein |
CHI | 10 |
| 2022 | Part of the Conversation: Workforce Professionals' Perspectives on the Roles and Impacts of Workforce TechnologiesabstractAmidst recent enthusiasm for data-driven technologies in workforce development, prior HCI research has explored job seekers' perspectives to inform the design of new technologies that could support their job search. However, in practice, the process of looking for work is often embedded in local workforce development ecosystems, where networks of organizations provide a range of services, from employment consulting, to resume workshops, to job skills training programs. Although prior CSCW work has explored the role of algorithms in social services, there has been little work investigating how algorithmic systems may shape workforce development professionals' interactions with clients and how they might be better designed to complement these professionals' work and responsibilities. To begin to address this gap, we conducted an interview study with five workforce development professionals in the US in both management and client-facing roles. Our findings contribute to research on how algorithmic systems are shaping workforce development, shedding light on the importance of the relationship building work that workforce professionals engage in with clients, the difficulty in maintaining boundaries in the face of resource and information challenges, and the ways that workforce development technologies are shaping the work of workforce development. Connie W. Chau, Kenneth Holstein, Michael A. Madaio |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2021 | Teachers' Orchestration Needs During the Shift to Remote Learning
LuEttaMae Lawrence, Kenneth Holstein, Susan R. Berman, Stephen Fancsali, Bruce M. McLaren, Steven Ritter 0001, Vincent Aleven |
EC-TEL | 2 |
| 2021 | SimPairing - Exploring Dynamic Pairing Policies through Historical Data Simulation and User-centered Research
Kexin Bella Yang, Xuejian Wang, Vanessa Echeverría, LuEttaMae Lawrence, Kenneth Holstein, Nikol Rummel, Vincent Aleven |
EDM | 5 |
| 2021 | Everyday Algorithm Auditing: Understanding the Power of Everyday Users in Surfacing Harmful Algorithmic BehaviorsabstractA growing body of literature has proposed formal approaches to audit algorithmic systems for biased and harmful behaviors. While formal auditing approaches have been greatly impactful, they often suffer major blindspots, with critical issues surfacing only in the context of everyday use once systems are deployed. Recent years have seen many cases in which everyday users of algorithmic systems detect and raise awareness about harmful behaviors that they encounter in the course of their everyday interactions with these systems. However, to date little academic attention has been granted to these bottom-up, user-driven auditing processes. In this paper, we propose and explore the concept of everyday algorithm auditing, a process in which users detect, understand, and interrogate problematic machine behaviors via their day-to-day interactions with algorithmic systems. We argue that everyday users are powerful in surfacing problematic machine behaviors that may elude detection via more centrally-organized forms of auditing, regardless of users' knowledge about the underlying algorithms. We analyze several real-world cases of everyday algorithm auditing, drawing lessons from these cases for the design of future platforms and tools that facilitate such auditing behaviors. Finally, we discuss work that lies ahead, toward bridging the gaps between formal auditing approaches and the organic auditing behaviors that emerge in everyday use of algorithmic systems. Hong Shen 0004, Alicia DeVos, Motahhare Eslami, Kenneth Holstein |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2021 | Can Crowds Customize Instructional Materials with Minimal Expert Guidance?: Exploring Teacher-guided Crowdsourcing for Improving Hints in an AI-based TutorabstractAI-based educational technologies may be most welcome in classrooms when they align with teachers' goals, preferences, and instructional practices. Teachers, however, have scarce time to make such customizations themselves. How might the crowd be leveraged to help time-strapped teachers? Crowdsourcing pipelines have traditionally focused on content generation. It is an open question how a pipeline might be designed so the crowd can succeed in a revision/customization task. In this paper, we explore an initial version of a teacher-guided crowdsourcing pipeline designed to improve the adaptive math hints of an AI-based tutoring system so they fit teachers' preferences, while requiring minimal expert guidance. In two experiments involving 144 math teachers and 481 crowdworkers, we found that such an expert-guided revision pipeline could save experts' time and produce better crowd-revised hints (in terms of teacher satisfaction) than two comparison conditions. The revised hints however, did not improve on the existing hints in the AI tutor, which were carefully-written but still have room for improvement and customization. Further analysis revealed that the main challenge for crowdworkers may lie in understanding teachers' brief written comments and implementing them in the form of effective edits, without introducing new problems. We also found that teachers preferred their own revisions over other sources of hints, and exhibited varying preferences for hints. Overall, the results confirm that there is a clear need for customizing hints to individual teachers' preferences. They also highlight the need for more elaborate scaffolds so the crowd can have specific knowledge of the requirements that teachers have for hints. The study represents a first exploration in the literature of how to support crowds with minimal expert guidance in revising and customizing instructional materials. Kexin Bella Yang, Tomohiro Nagashima, Junhui Yao, Joseph Jay Williams, Kenneth Holstein, Vincent Aleven |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2020 | Replay Enactments: Exploring Possible Futures through Historical DataabstractAs we design increasingly complex systems, we run up against fundamental limitations of human imagination. To support practice, it becomes essential to use authentic data and algorithms as design materials to augment designers' intuitions. Recent work has explored some dimensions of using data as a design material, suggesting the contours of a new space of design and prototyping methods. In this paper, we present Replay Enactments (REs, an extension of the User Enactments methods that uses data replay as a boundary object, making complex system behavior tangible to designers and stakeholders. We reflect on a set of case studies that have instantiated REs in diverse ways and discuss trade-offs between different ways of using data replays in design. We conclude by highlighting opportunities and challenges for future work. Kenneth Holstein, Erik Harpstead, Rebecca Gulotta, Jodi Forlizzi |
Conference on Designing Interactive Systems | 1 |
| 2020 | Towards Practical Detection of Unproductive Struggle
Stephen Fancsali, Kenneth Holstein, Michael Sandbothe, Steven Ritter 0001, Bruce M. McLaren, Vincent Aleven |
AIED (2) | 2 |
| 2020 | A Conceptual Framework for Human-AI Hybrid Adaptivity in Education
Kenneth Holstein, Vincent Aleven, Nikol Rummel |
AIED (1) | 1 |
| 2020 | The TA Framework: Designing Real-time Teaching Augmentation for K-12 ClassroomsabstractRecently, the HCI community has seen increased interest in the design of teaching augmentation (TA): tools that extend and complement teachers' pedagogical abilities during ongoing classroom activities. Examples of TA systems are emerging across multiple disciplines, taking various forms: e.g., ambient displays, wearables, or learning analytics dashboards. However, these diverse examples have not been analyzed together to derive more fundamental insights into the design of teaching augmentation. Addressing this opportunity, we broadly synthesize existing cases to propose the TA framework. Our framework specifies a rich design space in five dimensions, to support the design and analysis of teaching augmentation. We contextualize the framework using existing designs cases, to surface underlying design trade-offs: for example, balancing actionability of presented information with teachers' needs for professional autonomy, or balancing unobtrusiveness with informativeness in the design of TA systems. Applying the TA framework, we identify opportunities for future research and design. Pengcheng An, Kenneth Holstein, Bernice d'Anjou, Berry Eggen, Saskia Bakker |
CHI | 2 |
| 2020 | Exploring Human-AI Control Over Dynamic Transitions Between Individual and Collaborative Learning
Vanessa Echeverría, Kenneth Holstein, Jennifer Huang, Jonathan Sewall, Nikol Rummel, Vincent Aleven |
EC-TEL | 2 |
| 2019 | Designing for Complementarity: Teacher and Student Needs for Orchestration Support in AI-Enhanced Classrooms
Kenneth Holstein, Bruce M. McLaren, Vincent Aleven |
AIED (1) | 1 |
| 2019 | Improving Fairness in Machine Learning Systems: What Do Industry Practitioners Need?abstractThe potential for machine learning (ML) systems to amplify social inequities and unfairness is receiving increasing popular and academic attention. A surge of recent work has focused on the development of algorithmic tools to assess and mitigate such unfairness. If these tools are to have a positive impact on industry practice, however, it is crucial that their design be informed by an understanding of real-world needs. Through 35 semi-structured interviews and an anonymous survey of 267 ML practitioners, we conduct the first systematic investigation of commercial product teams' challenges and needs for support in developing fairer ML systems. We identify areas of alignment and disconnect between the challenges faced by teams in practice and the solutions proposed in the fair ML research literature. Based on these findings, we highlight directions for future ML and HCI research that will better address practitioners' needs. Kenneth Holstein, Jennifer Wortman Vaughan, Hal Daumé III, Miroslav Dudík, Hanna M. Wallach |
CHI | 1 |
| 2019 | Early Detection of Wheel Spinning: Comparison across Tutors, Models, Features, and Operationalizations
Chuankai Zhang, Yanzun Huang, Dongyang Lu, Weiqi Fang, John C. Stamper, Stephen Fancsali, Kenneth Holstein, Vincent Aleven |
EDM | 8 |
| 2018 | Student Learning Benefits of a Mixed-Reality Teacher Awareness Tool in AI-Enhanced Classrooms
Kenneth Holstein, Bruce M. McLaren, Vincent Aleven |
AIED (1) | 1 |
| 2018 | Opening Up an Intelligent Tutoring System Development Environment for Extensible Student Modeling
Kenneth Holstein, Zac Yu, Jonathan Sewall, Octav Popescu, Bruce M. McLaren, Vincent Aleven |
AIED (1) | 1 |
| 2018 | The classroom as a dashboard: co-designing wearable cognitive augmentation for K-12 teachersabstractWhen used in classrooms, personalized learning software allows students to work at their own pace, while freeing up the teacher to spend more time working one-on-one with students. Yet such personalized classrooms also pose unique challenges for teachers, who are tasked with monitoring classes working on divergent activities, and prioritizing help-giving in the face of limited time. This paper reports on the co-design, implementation, and evaluation of a wearable classroom orchestration tool for K-12 teachers: mixed-reality smart glasses that augment teachers' realtime perceptions of their students' learning, metacognition, and behavior, while students work with personalized learning software. The main contributions are: (1) the first exploration of the use of smart glasses to support orchestration of personalized classrooms, yielding design findings that may inform future work on real-time orchestration tools; (2) Replay Enactments: a new prototyping method for real-time orchestration tools; and (3) an in-lab evaluation and classroom pilot using a prototype of teacher smart glasses (Lumilo), with early findings suggesting that Lumilo can direct teachers' time to students who may need it most. Kenneth Holstein, Gena Hong, Mera Tegene, Bruce M. McLaren, Vincent Aleven |
LAK | 1 |
| 2018 | What exactly do students learn when they practice equation solving?: refining knowledge components with the additive factors modelabstractAccurately modeling individual students' knowledge growth is important in many applications of learning analytics. A key step is to decompose the knowledge targeted in the instruction into detailed knowledge components (KCs). We search for an accurate KC model for basic equation solving skills, using data from an intelligent tutoring system (ITS), Lynnette. Key criteria are data fit and predictive accuracy based on a standard logistic model called the Additive Factors Model (AFM). We focus on three difficulty factors for equation solving: understanding of variables, the negative sign, and the complexity of the equation. Fine-grained KC models were found to have greater fit and predictive accuracy than an "ideal," more abstract model, indicating that there is substantial under-generalization in students' equation-solving skill related to all three difficulty factors. The work enhances scientific understanding of the challenges students face in learning equation solving. It illustrates how learning analytics could inform the improvement of technology-enhanced learning environments. Yanjin Long, Kenneth Holstein, Vincent Aleven |
LAK | 2 |
| 2017 | Intelligent tutors as teachers' aides: exploring teacher needs for real-time analytics in blended classroomsabstractIntelligent tutoring systems (ITSs) are commonly designed to enhance student learning. However, they are not typically designed to meet the needs of teachers who use them in their classrooms. ITSs generate a wealth of analytics about student learning and behavior, opening a rich design space for real-time teacher support tools such as dashboards. Whereas real-time dashboards for teachers have become popular with many learning technologies, we are not aware of projects that have designed dashboards for ITSs based on a broad investigation of teachers' needs. We conducted design interviews with ten middle school math teachers to explore their needs for on-the-spot support during blended class sessions, as a first step in a user-centered design process of a real-time dashboard. Based on multi-methods analyses of this interview data, we identify several opportunities for ITSs to better support teachers' needs, noting that the analytics commonly generated by existing teacher support tools do not strongly align with the analytics teachers expect to be most useful. We highlight key tensions and tradeoffs in the design of such real-time supports for teachers, as revealed by "Speed Dating" possible futures with teachers. This paper has implications for our ongoing co-design of a real-time dashboard for ITSs, as well as broader implications for the design of ITSs that can effectively collaborate with teachers in classroom settings. Kenneth Holstein, Bruce M. McLaren, Vincent Aleven |
LAK | 1 |
| 2017 | SPACLE: investigating learning across virtual and physical spaces using spatial replaysabstractClassroom experiments that evaluate the effectiveness of educational technologies do not typically examine the effects of classroom contextual variables (e.g., out-of-software help-giving and external distractions). Yet these variables may influence students' instructional outcomes. In this paper, we introduce the Spatial Classroom Log Explorer (SPACLE): a prototype tool that facilitates the rapid discovery of relationships between within-software and out-of-software events. Unlike previous tools for retrospective analysis, SPACLE replays moment-by-moment analytics about student and teacher behaviors in their original spatial context. We present a data analysis workflow using SPACLE and demonstrate how this workflow can support causal discovery. We share the results of our initial replay analyses using SPACLE, which highlight the importance of considering spatial factors in the classroom when analyzing ITS log data. We also present the results of an investigation into the effects of student-teacher interactions on student learning in K-12 blended classrooms, using our workflow, which combines replay analysis with SPACLE and causal modeling. Our findings suggest that students' awareness of being monitored by their teachers may promote learning, and that "gaming the system" behaviors may extend outside of educational software use. Kenneth Holstein, Bruce M. McLaren, Vincent Aleven |
LAK | 1 |
| 2016 | Sequence Matters, But How Exactly? A Method for Evaluating Activity Sequences from Data
Shayan Doroudi, Kenneth Holstein, Vincent Aleven, Emma Brunskill |
EDM | 2 |
| 2015 | Inferring causal structure and hidden causes from event sequences
Christopher G. Lucas, Kenneth Holstein, Michael Pacer |
CogSci | 2 |
| 2015 | Towards Understanding How to Leverage Sense-making, Induction/Refinement and Fluency to Improve Robust Learning
Shayan Doroudi, Kenneth Holstein, Vincent Aleven, Emma Brunskill |
EDM | 2 |
| 2014 | Discovering hidden causes using statistical evidence
Christopher G. Lucas, Kenneth Holstein, Charles Kemp |
CogSci | 2 |