EDBT 2026 Demo / reviewers in the wild / expert
Ziang Xiao
dblp:196/1167
· DBLP profile ↗
29ranked-venue papers
9as first author
21since 2021 · last 2026
0000-0003-3368-0180ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 20 · 8 first-author · 12 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the WildabstractHow do product teams evaluate LLM-powered products? As organizations integrate large language models (LLMs) into digital products, their unpredictable nature makes traditional evaluation approaches inadequate, yet little is known about how practitioners navigate this challenge. Through interviews with nineteen practitioners across diverse sectors, we identify ten evaluation practices spanning informal ‘vibe checks’ to organizational meta-work. Beyond confirming four documented challenges, we introduce a novel fifth we call the results-actionability gap, in which practitioners gather evaluation data but cannot translate findings into concrete improvements. Drawing on patterns from successful teams, we contribute strategies to bridge this gap, supporting practitioners’ formalization journey from ad-hoc interpretive practices (e.g., vibe checks) toward systematic evaluation. Our analysis suggests these interpretive practices are necessary adaptations to LLM characteristics rather than methodological failures. For HCI researchers, this presents a research opportunity to support practitioners in systematizing emerging practices rather than developing new evaluation frameworks. Willem van der Maden, Malak Sadek, Ziang Xiao, Aske Mottelson, Qingzi Vera Liao, Jichen Zhu |
CHI | 3 |
| 2026 | Feedback by Design: Understanding and Overcoming User Feedback Barriers in Conversational AgentsabstractHigh-quality feedback is essential for effective human–AI interaction. It bridges knowledge gaps, corrects digressions, and shapes system behavior; both during interaction and throughout model development. Yet despite its importance, human feedback to AI is often infrequent and low quality. This gap motivates a critical examination of human feedback during interactions with AIs. To understand and overcome the challenges preventing users from giving high-quality feedback, we conducted two studies examining feedback dynamics between humans and conversational agents (CAs). Our formative study, through the lens of Grice’s maxims, identified four Feedback Barriers—Common Ground, Verifiability, Communication, and Informativeness—that prevent high-quality feedback by users. Building on these findings, we derive three design desiderata and show that systems incorporating scaffolds aligned with these desiderata enabled users to provide higher-quality feedback. Finally, we detail a call for action to the broader AI community for advances in Large Language Models capabilities to overcome Feedback Barriers. Zheng Zhang 0043, Namita Krishnan, Ziang Xiao, Yunyao Li 0001 |
CHI | 6 |
| 2026 | Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human OversightabstractThe dark patterns, deceptive interface designs manipulating user behaviors, have been extensively studied for their effects on human decision-making and autonomy. Yet, with the rising prominence of LLM-powered GUI agents that automate tasks from high-level intents, understanding how dark patterns affect agents is increasingly important. We present a two-phase empirical study examining how agents, human participants, and human-AI teams respond to 16 types of dark patterns across diverse scenarios. Phase 1 highlights that agents often fail to recognize dark patterns, and even when aware, prioritize task completion over protective action. Phase 2 revealed divergent failure modes: humans succumb due to cognitive shortcuts and habitual compliance, while agents falter from procedural blind spots. Human oversight improved avoidance but introduced costs such as attentional tunneling and cognitive load. Our findings show neither humans nor agents are uniformly resilient, and collaboration introduces new vulnerabilities, suggesting design needs for transparency, adjustable autonomy, and oversight. Bingcan Guo, Ibrahim Khalilov, Simret Araya Gebreegziabher, Bingsheng Yao, Dakuo Wang, Yanfang Ye 0001, Tianshi Li 0001, Ziang Xiao, Yaxing Yao, Toby Jia-Jun Li |
CHI | 12 |
| 2025 | Canvil: Designerly Adaptation for LLM-Powered User Experiences
K. J. Kevin Feng, Qingzi Vera Liao, Ziang Xiao, Jennifer Wortman Vaughan, Amy X. Zhang, David W. McDonald |
CHI | 3 |
| 2025 | From Text to Trust: Empowering AI-assisted Decision Making with Adaptive LLM-powered Analysis
Zhuoyan Li, Hangxiao Zhu, Zhuoran Lu, Ziang Xiao, Ming Yin 0001 |
CHI | 4 |
| 2025 | Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature ReviewabstractLarge language models (LLMs) have been positioned to revolutionize HCI, by reshaping not only the interfaces, design patterns, and sociotechnical systems that we study, but also the research practices we use.To-date, however, there has been little understanding of LLMs' uptake in HCI.We address this gap via a systematic literature review of 153 CHI papers from 2020-24 that engage with LLMs.We taxonomize: (1) domains where LLMs are applied; (2) roles of LLMs in HCI projects; (3) contribution types; and (4) acknowledged limitations and risks.We find LLM work in 10 diverse domains, primarily via empirical and artifact contributions.Authors use LLMs in five distinct roles, including as research tools or simulated users.Still, authors often raise validity and reproducibility concerns, and overwhelmingly study closed models.We outline opportunities to improve HCI research with and on LLMs, and provide guiding questions for researchers to consider the validity and appropriateness of LLM-related work. Rock Yuren Pang, Hope Schroeder, Kynnedy Simone Smith, Solon Barocas, Ziang Xiao, Emily Tseng, Danielle Bragg |
CHI | 5 |
| 2025 | Assistance or Disruption? Exploring and Evaluating the Design and Trade-offs of Proactive AI Programming SupportabstractAI programming tools enable powerful code generation, and recent prototypes attempt to reduce user effort with proactive AI agents, but their impact on programming workflows remains unexplored. We introduce and evaluate Codellaborator, a design probe LLM agent that initiates programming assistance based on editor activities and task context. We explored three interface variants to assess trade-offs between increasingly salient AI support: prompt-only, proactive agent, and proactive agent with presence and context (Codellaborator). In a within-subject study (N=18), we find that proactive agents increase efficiency compared to prompt-only paradigm, but also incur workflow disruptions. However, presence indicators and interaction context support alleviated disruptions and improved users' awareness of AI processes. We underscore trade-offs of Codellaborator on user control, ownership, and code understanding, emphasizing the need to adapt proactivity to programming processes. Our research contributes to the design exploration and evaluation of proactive AI systems, presenting design implications on AI-integrated programming workflow. Kevin Pu, Daniel Lazaro, Ian Arawjo, Haijun Xia, Ziang Xiao, Tovi Grossman, Yan Chen 0033 |
CHI | 5 |
| 2025 | Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testingabstract*Warning: Contains harmful model outputs.*
Despite significant advancements, the propensity of Large Language Models (LLMs) to generate harmful and unethical content poses critical challenges.
Measuring value alignment of LLMs becomes crucial for their regulation and responsible deployment. Although numerous benchmarks have been constructed to assess social bias, toxicity, and ethical issues in LLMs, those static benchmarks suffer from *evaluation chronoeffect*, in which, as models rapidly evolve, existing benchmarks may leak into training data or become saturated, *overestimating* ever-developing LLMs. To tackle this problem, we propose GETA, a novel *generative evolving testing* approach based on adaptive testing methods in measurement theory. Unlike traditional adaptive testing methods that rely on a static test item pool, GETA probes the underlying moral boundaries of LLMs by dynamically generating test items tailored to model capability. GETA co-evolves with LLMs by learning a joint distribution of item difficulty and model value conformity, thus effectively addressing evaluation chronoeffect.
We evaluated various popular LLMs with GETA and demonstrated that 1) GETA can dynamically create difficulty-tailored test items and 2) GETA's evaluation results are more consistent with models' performance on unseen OOD and i.i.d. items, laying the groundwork for future evaluation paradigms. Han Jiang 0007, Xiaoyuan Yi, Zhihua Wei 0001, Ziang Xiao, Xing Xie 0001 |
ICML | 4 |
| 2025 | Faux Polyglot: A Study on Information Disparity in Multilingual Large Language ModelsabstractNikhil Sharma, Kenton Murray, Ziang Xiao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kenton Murray, Ziang Xiao |
NAACL (Long Papers) | 3 |
| 2025 | Introduction to the Special Issue on Human-Centric Generative AIabstractGenerative AI increasingly reshapes how people engage with interactive systems. It now plays a vital role in designing, studying, and refining human-centered methods that let individuals interact and collaborate with AI, strengthening their agency and control. This special issue highlights the human role in Generative AI and seeks approaches that equip diverse stakeholders across socio-technical contexts to understand, direct, and steer these systems while enabling responsible innovation. We publish in this special issue original research on new interaction techniques that integrate human input into Generative AI’s continual development, studies of interaction paradigms that support more effective human–AI collaboration, and work that deepens understanding of model capabilities. Thus, we aim to build a research community around Human-Centric GenAI that empowers people to actively shape systems in line with their values, needs, and expectations. Yunyao Li 0001, Mary Lou Maher, Ziang Xiao |
ACM Trans. Interact. Intell. Syst. | 3 |
| 2024 | ECBD: Evidence-Centered Benchmark Design for NLPabstractYu Lu Liu, Su Lin Blodgett, Jackie Cheung, Q. Vera Liao, Alexandra Olteanu, Ziang Xiao. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yu Lu Liu, Su Lin Blodgett, Jackie Chi Kit Cheung, Qingzi Vera Liao, Alexandra Olteanu, Ziang Xiao |
ACL (1) | 6 |
| 2024 | Generative Echo Chamber? Effect of LLM-Powered Search Systems on Diverse Information SeekingabstractLarge language models (LLMs) powered conversational search systems have already been used by hundreds of millions of people, and are believed to bring many benefits over conventional search. However, while decades of research and public discourse interrogated the risk of search systems in increasing selective exposure and creating echo chambers—limiting exposure to diverse opinions and leading to opinion polarization, little is known about such a risk of LLM-powered conversational search. We conduct two experiments to investigate: 1) whether and how LLM-powered conversational search increases selective exposure compared to conventional search; 2) whether and how LLMs with opinion biases that either reinforce or challenge the user’s view change the effect. Overall, we found that participants engaged in more biased information querying with LLM-powered conversational search, and an opinionated LLM reinforcing their views exacerbated this bias. These results present critical implications for the development of LLMs and conversational search systems, and the policy governing these technologies. Qingzi Vera Liao, Ziang Xiao |
CHI | 3 |
| 2024 | Improving Context-Aware Preference Modeling for Language ModelsabstractWhile finetuning language models from pairwise preferences has proven remarkably effective, the underspecified nature of natural language presents critical challenges. Direct preference feedback is uninterpretable, difficult to provide where multidimensional criteria may apply, and often inconsistent, either because it is based on incomplete instructions or provided by diverse principals. To address these challenges, we consider the two-step preference modeling procedure that first resolves the under-specification by selecting a context, and then evaluates preference with respect to the chosen context. We decompose reward modeling error according to these two steps, which suggests that supervising context in addition to context-specific preference may be a viable approach to aligning models with diverse human preferences. For this to work, the ability of models to evaluate context-specific preference is critical. To this end, we contribute context-conditioned preference datasets and accompanying experiments that investigate the ability of language models to evaluate context-specific preference. Unlike past datasets, where context-specific preference is highly correlated with general preference, our "preference reversal" datasets disentangle context-specific and general preferences to isolate context-specific capabilities. We use our datasets to (1) show that existing preference models benefit from, but fail to fully consider, added context, (2) finetune a context-aware reward model with context-specific performance exceeding that of GPT-4 and Llama 3 70B, and (3) investigate the potential value of context-aware preference modeling. Silviu Pitis, Ziang Xiao, Nicolas Le Roux, Alessandro Sordoni |
NeurIPS | 2 |
| 2023 | Inform the Uninformed: Improving Online Informed Consent Reading with an AI-Powered ChatbotabstractInformed consent is a core cornerstone of ethics in human subject research. Through the informed consent process, participants learn about the study procedure, benefits, risks, and more to make an informed decision. However, recent studies showed that current practices might lead to uninformed decisions and expose participants to unknown risks, especially in online studies. Without the researcher’s presence and guidance, online participants must read a lengthy form on their own with no answers to their questions. In this paper, we examined the role of an AI-powered chatbot in improving informed consent online. By comparing the chatbot with form-based interaction, we found the chatbot improved consent form reading, promoted participants’ feelings of agency, and closed the power gap between the participant and the researcher. Our exploratory analysis further revealed the altered power dynamic might eventually benefit study response quality. We discussed design implications for creating AI-powered chatbots to offer effective informed consent in broader settings. Ziang Xiao, Tiffany Wenting Li, Karrie Karahalios, Hari Sundaram |
CHI | 1 |
| 2023 | ByteSized32: A Corpus and Challenge Task for Generating Task-Specific World Models Expressed as Text GamesabstractIn this work we investigate the capacity of language models to generate explicit, interpretable, and interactive world models of scientific and common-sense reasoning tasks.We operationalize this as a task of generating text games, expressed as hundreds of lines of PYTHON code.To facilitate this task, we introduce BYTESIZED32 1 , a corpus of 32 reasoning-focused text games totalling 20k lines of PYTHON code.We empirically demonstrate that GPT-4 can use these games as templates for single-shot in-context learning, successfully producing runnable games on unseen topics in 28% of cases.When allowed to selfreflect on program errors, game runnability substantially increases to 57%.While evaluating simulation fidelity is labor intensive, we introduce a suite of automated metrics to assess game fidelity, technical validity, adherence to task specifications, and winnability, showing a high-degree of agreement with expert human ratings.We pose this as a challenge task to spur further development at the juncture of world modeling and code generation. Ruoyao Wang, Graham Todd, Xingdi Yuan, Ziang Xiao, Marc-Alexandre Côté, Peter A. Jansen |
EMNLP | 4 |
| 2023 | Evaluating Evaluation Metrics: A Framework for Analyzing NLG Evaluation Metrics using Measurement TheoryabstractWe address a fundamental challenge in Natural Language Generation (NLG) model evaluation-the design and evaluation of evaluation metrics.Recognizing the limitations of existing automatic metrics and noises from how current human evaluation was conducted, we propose METRICEVAL, a framework informed by measurement theory, the foundation of educational test design, for conceptualizing and evaluating the reliability and validity of NLG evaluation metrics.The framework formalizes the source of measurement error and offers statistical tools for evaluating evaluation metrics based on empirical data.With our framework, one can quantify the uncertainty of the metrics to better interpret the result.To exemplify the use of our framework in practice, we analyzed a set of evaluation metrics for summarization and identified issues related to conflated validity structure in human-eval and reliability in LLM-based metrics.Through METRICEVAL 1 , we aim to promote the design, evaluation, and interpretation of valid and reliable metrics to advance robust and effective NLG models. Ziang Xiao, Susu Zhang, Vivian Lai, Qingzi Vera Liao |
EMNLP | 1 |
| 2023 | Powering an AI Chatbot with Expert Sourcing to Support Credible Health Information AccessabstractDuring a public health crisis like the COVID-19 pandemic, a credible and easy-to-access information portal is highly desirable. It helps with disease prevention, public health planning, and misinformation mitigation. However, creating such an information portal is challenging because 1) domain expertise is required to identify and curate credible and intelligible content, 2) the information needs to be updated promptly in response to the fast-changing environment, and 3) the information should be easily accessible by the general public; which is particularly difficult when most people do not have the domain expertise about the crisis. In this paper, we presented an expert-sourcing framework and created Jennifer, an AI chatbot, which serves as a credible and easy-to-access information portal for individuals during the COVID-19 pandemic. Jennifer was created by a team of over 150 scientists and health professionals around the world, deployed in the real world and answered thousands of user questions about COVID-19. We evaluated Jennifer from two key stakeholders’ perspectives, expert volunteers and information seekers. We first interviewed experts who contributed to the collaborative creation of Jennifer to learn about the challenges in the process and opportunities for future improvement. We then conducted an online experiment that examined Jennifer’s effectiveness in supporting information seekers in locating COVID-19 information and gaining their trust. We share the key lessons learned and discuss design implications for building expert-sourced and AI-powered information portals, along with the risks and opportunities of misinformation mitigation and beyond. Ziang Xiao, Qingzi Vera Liao, Michelle X. Zhou, Tyrone Grandison, Yunyao Li 0001 |
IUI | 1 |
| 2023 | Joint Prompt Optimization of Stacked LLMs using Variational InferenceabstractLarge language models (LLMs) can be seen as atomic units of computation mapping sequences to a distribution over sequences. Thus, they can be seen as stochastic language layers in a language network, where the learnable parameters are the natural language prompts at each layer. By stacking two such layers and feeding the output of one layer to the next, we obtain a Deep Language Network (DLN). We first show how to effectively perform prompt optimization for a 1-Layer language network (DLN-1). Then, we present an extension that applies to 2-layer DLNs (DLN-2), where two prompts must be learned. The key idea is to consider the output of the first layer as a latent variable, which requires inference, and prompts to be learned as the parameters of the generative distribution. We first test the effectiveness of DLN-1 in multiple reasoning and natural language understanding tasks. Then, we show that DLN-2 can reach higher performance than a single layer, showing promise that we might reach comparable performance to GPT-4, even when each LLM in the network is smaller and less powerful. Alessandro Sordoni, Eric Yuan, Marc-Alexandre Côté, Matheus Pereira, Adam Trischler, Ziang Xiao, Seyed Arian Hosseini, Friederike Niedtner, Nicolas Le Roux |
NeurIPS | 6 |
| 2023 | What should I Ask: A Knowledge-driven Approach for Follow-up Questions Generation in Conversational Surveys
Yubin Ge, Ziang Xiao, Jana Diesner, Heng Ji 0001, Karrie Karahalios, Hari Sundaram |
PACLIC | 2 |
| 2021 | Contestability For Content ModerationabstractContent moderation systems for social media have had numerous issues of bias, in terms of race, gender, and ability among many others. One proposal for addressing such issues in automated decision making is by designing for contestability, whereby users can shape and influence how decisions are made. In this study, we conduct a series of participatory design workshops with participants from communities that have experienced problems with social media content moderation in the past. Together with participants, we explore the idea of designing for contestability in content moderation and find that users' designs suggest three fruitful, practical avenues: adding representation, improving communication, and designing with compassion. We conclude with design recommendations drawn from participants' proposals, and reflect on the challenges that remain. Kristen Vaccaro, Ziang Xiao, Kevin Hamilton, Karrie Karahalios |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2021 | Let Me Ask You This: How Can a Voice Assistant Elicit Explicit User Feedback?abstractVoice assistants offer users access to an increasing variety of personalized functionalities. Researchers and engineers who build these experiences rely on various signals from users to create the machine learning models powering them. One type of signal is explicit feedback. While collecting explicit user feedback in situ via voice assistants would help improve and inspect the underlying models, from a user perspective it can be disruptive to the overall experience, and the user might not feel compelled to respond. However, careful design can help alleviate the friction in the experience. In this paper, we explore the opportunities and the design space for voice assistant explicit feedback elicitation. First, we present four usage categories of explicit feedback in situ for model evaluation and improvement, derived from interviews with machine learning practitioners. Then, using realistic scenarios generated for each category, we conducted an online study to evaluate multiple voice assistant designs. Our results reveal that when the voice assistant is introduced as a learner or a collaborator, users were more willing to respond to its request for feedback and felt less disruptive. In addition, giving users instructions on how to initiate feedback themselves can reduce the perceived disruptiveness compared to asking users for feedback directly. Based on our findings, we discuss the implications and potential future directions for designing voice assistants to elicit user feedback for personalized voice experiences. Ziang Xiao, Sarah Mennicken, Bernd Huber, Adam Shonkoff, Jennifer Thom-Santelli |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2020 | If I Hear You Correctly: Building and Evaluating Interview Chatbots with Active Listening SkillsabstractInterview chatbots engage users in a text-based conversation to draw out their views and opinions. It is, however, challenging to build effective interview chatbots that can handle user free-text responses to open-ended questions and deliver engaging user experience. As the first step, we are investigating the feasibility and effectiveness of using publicly available, practical AI technologies to build effective interview chatbots. To demonstrate feasibility, we built a prototype scoped to enable interview chatbots with a subset of active listening skills-the abilities to comprehend a user's input and respond properly. To evaluate the effectiveness of our prototype, we compared the performance of interview chatbots with or without active listening skills on four common interview topics in a live evaluation with 206 users. Our work presents practical design implications for building effective interview chatbots, hybrid chatbot platforms, and empathetic chatbots beyond interview tasks. Ziang Xiao, Michelle X. Zhou, Huahai Yang, Chang Yan Chi |
CHI | 1 |
| 2020 | Tell Me About Yourself: Using an AI-Powered Chatbot to Conduct Conversational Surveys with Open-ended QuestionsabstractThe rise of increasingly more powerful chatbots offers a new way to collect information through conversational surveys, where a chatbot asks open-ended questions, interprets a user’s free-text responses, and probes answers whenever needed. To investigate the effectiveness and limitations of such a chatbot in conducting surveys, we conducted a field study involving about 600 participants. In this study with mostly open-ended questions, half of the participants took a typical online survey on Qualtrics and the other half interacted with an AI-powered chatbot to complete a conversational survey. Our detailed analysis of over 5,200 free-text responses revealed that the chatbot drove a significantly higher level of participant engagement and elicited significantly better quality responses measured by Gricean Maxims in terms of their informativeness, relevance, specificity, and clarity. Based on our results, we discuss design implications for creating AI-powered chatbots to conduct effective surveys and beyond. Ziang Xiao, Michelle X. Zhou, Qingzi Vera Liao, Gloria Mark, Chang Yan Chi, Huahai Yang |
ACM Trans. Comput. Hum. Interact. | 1 |
| 2019 | Who should be my teammates: using a conversational agent to understand individuals and help teamingabstractWe are building an intelligent agent to help teaming efforts. In this paper, we investigate the real-world use of such an agent to understand students deeply and help student team formation in a large university class involving about 200 students and 40 teams. Specifically, the agent interacted with each student in a text-based conversation at the beginning and end of the class. We show how the intelligent agent was able to elicit in-depth information from the students, infer the students' personality traits, and reveal the complex relationships between team personality compositions and team results. We also report on the students' behavior with and impression of the agent. We discuss the benefits and limitations of such an intelligent agent in helping team formation, and the design considerations for creating intelligent agents for aiding in teaming efforts. Ziang Xiao, Michelle X. Zhou, Wai-Tat Fu |
IUI | 1 |
| 2019 | Should We Use an Abstract Comic Form to Persuade?: Experiments with Online Charitable DonationabstractThis paper examines the use of the abstract comic form for persuading online charitable donations. Persuading individuals to contribute to charitable causes online is hard and responses to the appeals are typically low; charitable donations share the structure of public goods dilemmas where the rewards are distant and non-exclusive. In this paper, we examine if comics in abstract form are more persuasive than in the plain text form. Drawing on a rich literature on comics, we synthesized a three-panel abstract comic to create our appeal. We conducted a between-subject study with 307 participants from Amazon Mechanical Turk on the use of abstract comic form to appeal for charitable donations. As part of our experimental procedure, we sought to persuade individuals to contribute to a real charity focused on Autism research with monetary costs. We compared the average amount of donation to the charity under three conditions: the plain text message, an abstract comic that includes the plain text, and an abstract comic that additionally includes the social proof. We use Bayesian modeling to analyze the results, motivated by model transparency and its use in small-sized studies. Our experiments reveal that the message in abstract comic form elicited significantly more donations than text form (medium to large effect size=0.59). Incorporating social proof in the abstract comic message did not show a significant effect. Our studies have design implications: non-profits and governmental agencies interested in alleviating public goods dilemmas that share a similar structure to our experiment (single-shot task, distant, non-exclusive reward) ought to consider including messages in the abstract comic form as part of their online fund-raising campaign. Ziang Xiao, Po-Shiun Ho, Karrie Karahalios, Hari Sundaram |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2018 | Supporting Spatial Skill Learning with Gesture-Based Embodied DesignabstractPrior research has shown that spatial abilities are crucial for STEM achievement and attainment. The connection between the digital and physical worlds provided by embodied interaction has been shown to enhance performance and engagement in educational contexts. Spatial reasoning is a domain that lends itself naturally to embodied, physical interaction; however, there is little understanding of how embodied interaction could be incorporated into educational technology designed to train spatial reasoning skills. We propose several guidelines for gestural interaction design in spatial reasoning education games based on an empirical study with students at a local afterschool program using a custom-built computer game for training spatial skills. We present a series of gesture sets derived from an iterative design approach that are easy for children to acquire, show sufficient congruency to specific spatial operations, and enable robust recognition from the system. We also compared children's behaviors when playing the game with our gestural interface and a traditional mouse-based interface and found that children take more time but fewer steps to complete game levels when using gestures. Po-Tsung Chiu, Helen Wauck, Ziang Xiao, Yuqi Yao, Wai-Tat Fu |
IUI | 3 |
| 2018 | Cubicle: An Adaptive Educational Gaming Platform for Training Spatial Visualization SkillsabstractResearch has demonstrated that spatial visualization skills are crucial for success in Science, technology, engineering, and mathematics (STEM) disciplines. With an increasing number of students entering STEM disciplines, the question of how to effectively train students» spatial visualization skills has become very important. While a scalable existing solution is to implement online workshops for students, the problem of how to motivate students to participate in these online workshops remains unsolved. In this study, we studied gamification as a way to motivate first year engineering students to take part in an online workshop designed to train their spatial visualization skills. Our game contains eight modules, each designed to train a different component of spatial visualization. The game records players» in-game behavior with high granularity, which allows us to provide automated, scalable feedback on players» problem-solving strategies. Ten students with different levels of spatial ability played our game and expressed a strong interest in using the game to train their spatial visualization skills in the future. In addition, our analysis of players» in-game behaviors shows the potential benefits of implementing adaptive and personalized learning guidance. Ziang Xiao, Helen Wauck, Zeya Peng, Hanfei Ren, Shiliang Zuo, Yuqi Yao, Wai-Tat Fu |
IUI | 1 |
| 2018 | To Label or Not to Label: The Effect of Stance and Credibility Labels on Readers' Selection and Perception of News ArticlesabstractSocial media sites use different labels to help users find and select news feeds. For example, Blue Feed, Red Feed, a news feed created by the Wall Street Journal, use stance labels to separate news articles with opposing political ideologies to help people explore diverse opinions. To combat the spread of fake news, Facebook has experimented with putting credibility labels on news articles to help readers decide whether the content is trustworthy. To systematically understand the effects of stance and credibility labels on online news selection and consumption, we conducted a controlled experiment to study how these labels influence the selection, perceived extremeness, and level of agreement of news articles. Results show that stance labels may intensify selective exposure - a tendency for people to look for agreeable opinions -- and make people more vulnerable to polarized opinions and fake news. We found, however, that the effect of credibility labels on reducing selective exposure and recognizing fake news is limited. Although originally designed to encourage exposure to opposite viewpoints, stance labels can make fake news articles look more trustworthy, and they may lower people's perception of the extremeness of fake news articles. Our results have important implications on the subtle effects of stance and credibility labels on online news consumption. Mingkun Gao, Ziang Xiao, Karrie Karahalios, Wai-Tat Fu |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2017 | Untangling the Relationship Between Spatial Skills, Game Features, and Gender in a Video GameabstractCertain commercial video games, such as Portal 2 and Tetris, have been empirically shown to train spatial reasoning skills, a subset of cognitive skills essential for success in STEM disciplines. However, no research to date has attempted to understand which specific features in these games tap into players' spatial ability or how individual player differences interact with these game features. This knowledge is crucially important as a first step towards understanding what makes these games effective and why, especially for subpopulations with lower spatial ability such as women and girls. We present the first empirical study analyzing the relationship between spatial ability, specific game features, and individual player differences using a custom-built computer game. Twenty children took a pretest of spatial skills and then played our game for 2 hours. We found that spatial ability pretest scores predicted several player behaviors related to in-game tasks involving 3D object construction and first person navigation. However, when analyzed by gender, girls' pretest scores were much less predictive of player behavior. Helen Wauck, Ziang Xiao, Po-Tsung Chiu, Wai-Tat Fu |
IUI | 2 |