EDBT 2026 Demo / reviewers in the wild / expert
Bingsheng Yao
dblp:256/9562
· DBLP profile ↗
31ranked-venue papers
5as first author
30since 2021 · last 2026
0009-0004-8329-4610ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 20 · 2 first-author · 19 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human EvaluationabstractJiaju Chen, Yuxuan Lu, Xiaojie Wang, Huimin Zeng, Jing Huang, Jiri Gesi, Ying Xu, Bingsheng Yao, Dakuo Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiaju Chen, Yuxuan Lu 0003, Jiri Gesi, Bingsheng Yao, Dakuo Wang |
ACL (1) | 8 |
| 2026 | Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior DataabstractYuxuan Lu, Jing Huang, Yan Han, Bingsheng Yao, Sisong Bei, Yaochen Xie, Yisi Sang, Qi He, Dakuo Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuxuan Lu 0003, Yan Han 0001, Bingsheng Yao, Sisong Bei, Yaochen Xie, Yisi Sang, Qi He 0002, Dakuo Wang |
ACL (1) | 4 |
| 2026 | LLM-based Embodied Conversational Agent for Reducing Foreign Language Speaking Anxiety in Social VRabstractForeign language speaking anxiety (FLSA) poses a major challenge for English-language learners, suppressing confidence and triggering a cycle of avoidance that hinders language acquisition. To address this, we explored the use of LLM-based embodied conversational agents (ECA) in social virtual reality (VR), which provide personalized support and multimodal interaction in a contextualized environment. We developed three English-language learning scenarios in social VR and conducted a five-day mixed-methods study where participants (N=20) engaged in daily 30-minute role-play practice with an LLM-based ECA to evaluate the efficacy of the system. Quantitative results showed a significant reduction in self-reported FLAS after 3 days, along with subtle gains in speaking proficiency measures. Qualitatively, learners perceived increased confidence, attributing it to the LLM-based ECA’s non-judgmental stance, linguistic scaffolding, affective encouragement, and adaptive feedback. Our findings suggest the potential of LLM-based ECAs in social VR for language learning and offer considerations for future agent design. Mengxu Pan, Panxin Liu, Jinda Zhang, Raina Cao, Viduni Ariyawansa, Bingsheng Yao, Dakuo Wang, Philippe Pasquier, Alexandra Kitson, Mirjana Prpa |
CHI | 7 |
| 2026 | Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human OversightabstractThe dark patterns, deceptive interface designs manipulating user behaviors, have been extensively studied for their effects on human decision-making and autonomy. Yet, with the rising prominence of LLM-powered GUI agents that automate tasks from high-level intents, understanding how dark patterns affect agents is increasingly important. We present a two-phase empirical study examining how agents, human participants, and human-AI teams respond to 16 types of dark patterns across diverse scenarios. Phase 1 highlights that agents often fail to recognize dark patterns, and even when aware, prioritize task completion over protective action. Phase 2 revealed divergent failure modes: humans succumb due to cognitive shortcuts and habitual compliance, while agents falter from procedural blind spots. Human oversight improved avoidance but introduced costs such as attentional tunneling and cognitive load. Our findings show neither humans nor agents are uniformly resilient, and collaboration introduces new vulnerabilities, suggesting design needs for transparency, adjustable autonomy, and oversight. Bingcan Guo, Ibrahim Khalilov, Simret Araya Gebreegziabher, Bingsheng Yao, Dakuo Wang, Yanfang Ye 0001, Tianshi Li 0001, Ziang Xiao, Yaxing Yao, Toby Jia-Jun Li |
CHI | 8 |
| 2026 | Through the Lens of Human-Human Collaboration: An Configurable Research Platform for Exploring Human-Agent CollaborationabstractIntelligent systems have traditionally been designed as tools rather than collaborators, often lacking critical characteristics that collaboration partnerships require. Recent advances in large language model (LLM) agents open new opportunities for human-LLM-agent collaboration by enabling natural communication and various social and cognitive behaviors. Yet it remains unclear whether principles of computer-mediated collaboration established in HCI and CSCW persist, change, or fail when humans collaborate with LLM agents. To support systematic investigations of these questions, we introduce an open and configurable research platform for HCI researchers1. The platform’s modular design allows seamless adaptation of classic CSCW experiments and manipulation of theory-grounded interaction controls. We demonstrate the platform’s research efficacy and usability through three case studies: (1) two Shape FactoryHidden Profile experiment for information pooling with 16 participants, and (3) a participatory cognitive walkthrough with five HCI researchers to refine workflows of researcher interface for experiment setup and analysis. Bingsheng Yao, Jiaju Chen, April Yi Wang, Toby Jia-Jun Li, Dakuo Wang |
CHI | 1 |
| 2026 | Exploring Collaboration Breakdowns Between Provider Teams and Patients in Post-Surgery CareabstractPost-surgery care involves ongoing collaboration between provider teams and patients, which starts from post-surgery hospitalization through home recovery after discharge. While prior HCI research has primarily examined patients' challenges at home, less is known about how provider teams coordinate discharge preparation and care handoffs, and how breakdowns in communication and care pathways may affect patient recovery. To investigate this gap, we conducted semi-structured interviews with 13 healthcare providers and 4 patients in the context of gastrointestinal (GI) surgery. We found coordination boundaries between in- and out-patient teams, coupled with complex organizational structures within teams, impeded the "invisible work" of preparing patients' home care plans and triaging patient information. For patients, these breakdowns resulted in inadequate preparation for home transition and fragmented self-collected data, both of which undermine timely clinical decision-making. Based on these findings, we outline design opportunities to formalize task ownership and handoffs, contextualize co-temporal signals, and align care plans with home resources. Bingsheng Yao, Menglin Zhao, Zhan Zhang 0008, Pengqi Wang, Emma G. Chester, Changchang Yin, Tianshi Li 0001, Varun Mishra 0001, Lace M. K. Padilla, Odysseas Chatzipanagiotou, Timothy Pawlik, Ping Zhang 0016, Weidan Cao, Dakuo Wang |
CHI | 1 |
| 2026 | Balancing Efficiency and Empathy: Healthcare Providers' Perspectives on AI-Supported Workflows for Serious Illness Conversations in the Emergency DepartmentabstractSerious Illness Conversations (SICs)—discussions about values and care preferences for patients with life-threatening illness—rarely occur in Emergency Departments (EDs), despite evidence that early conversations improve care alignment and reduce unnecessary interventions. We interviewed 11 ED providers to identify challenges in SICs and opportunities for technology support, with a focus on AI. Our analysis revealed a four-stage SIC workflow (identification, preparation, conduction, documentation) and barriers at each stage, including fragmented patient information, limited time and space, lack of conversational guidance, and burdensome documentation. Providers expressed interest in AI systems for synthesizing information, supporting real-time conversations, and automating documentation, but emphasized concerns about preserving human connection and clinical autonomy. This tension highlights the need for technologies that enhance efficiency without undermining the interpersonal nature of SICs. We propose design guidelines for ambient and peripheral AI systems to support providers while preserving the essential humanity of these conversations. Menglin Zhao, Zhuorui Yong, Ruijian Hannah Guan, Kai-Wei Chang 0001, Adrian Haimovich, Kei Ouchi, Timothy W. Bickmore, Zhan Zhang 0008, Bingsheng Yao, Dakuo Wang, Smit Desai |
CHI | 9 |
| 2026 | Striking a Balance: Evaluating How Aggregations of Multiple Forecasts Impact Judgment Under Uncertainty
Ruishi Zou, Racquel Fygenson, Bingsheng Yao, Dakuo Wang, Lace M. K. Padilla |
PacificVis | 4 |
| 2026 | RECOVER: Designing a Large Language Model-based Remote Patient Monitoring System for Postoperative Gastrointestinal Cancer Care CSCW032abstractCancer surgery is a key treatment for gastrointestinal (GI) cancers, a group of cancers that account for more than 35% of cancer-related deaths worldwide, but postoperative complications are unpredictable and can be life-threatening. In this paper, we investigate how recent advancements in large language models (LLMs) can benefit remote patient monitoring (RPM) systems through clinical integration by designing RECOVER, an LLM-powered RPM system for postoperative GI cancer care. To closely engage stakeholders in the design process, we first conducted seven participatory design sessions with five clinical staff and interviewed five cancer patients to derive six major design strategies for integrating clinical guidelines and information needs into LLM-based RPM systems. We then designed and implemented RECOVER, which features an LLM-powered conversational agent for cancer patients and an interactive dashboard for clinical staff to enable efficient postoperative RPM. Finally, we used RECOVER as a pilot system to assess the implementation of our design strategies with four clinical staff and five patients, providing design implications by identifying crucial design elements, offering insights on responsible AI, and outlining opportunities for future LLM-powered RPM systems. Yuxuan Lu 0003, Jennifer Bagdasarian, Vedant Das Swain, Collin Campbell, Waddah Al-Refaie, Jehan El-Bayoumi, Guodong Gordon Gao, Dakuo Wang, Bingsheng Yao, Nawar Shara |
Proc. ACM Hum. Comput. Interact. | 11 |
| 2025 | Examining Student and Teacher Perspectives on Undisclosed Use of Generative AI in Academic Work
Rudaiba Adnin, Atharva Pandkar, Bingsheng Yao, Dakuo Wang, Maitraye Das |
CHI | 3 |
| 2025 | Characterizing LLM-Empowered Personalized Story Reading and Interaction for Children: Insights From Multi-Stakeholder PerspectivesabstractPeer Reviewed Jiaju Chen, Minglong Tang, Yuxuan Lu 0003, Bingsheng Yao, Elissa Fan, Xiaojuan Ma, Dakuo Wang, Yuling Sun, Liang He 0001 |
CHI | 4 |
| 2025 | Promoting Prosociality via Micro-acts of Joy: A Large-Scale Well-Being Intervention StudyabstractProsociality has been well-documented to positively impact mental, social, and physical well-being.However, existing studies of interventions for promoting prosociality have limitations such as Hitesh Goel, Yoobin Park, Jin Liou, Darwin A. Guevarra, Peggy Callahan, Jolene Smith, Bingsheng Yao, Dakuo Wang, Xin Liu 0034, Daniel McDuff, Noémie Elhadad, Emiliana Simon-Thomas, Elissa Epel, Xuhai Xu |
CHI | 7 |
| 2025 | Live-Streaming-Based Dual-Teacher Classes for Equitable Education: Insights and Challenges From Local Teachers' Perspective in Disadvantaged AreasabstractEducational inequalities in disadvantaged areas have long been a global concern. While Information and Communication Technologies (ICTs) have shown great potential in addressing this issue, the unique challenges in disadvantaged areas often hinder the practical effectiveness of such technologies. This paper examines live-streaming-based dual-teacher classes (LSDC) through a qualitative study in disadvantaged regions of China. Our findings indicate that, although LSDC offers students in these regions access to high-quality educational resources, its practical implementation is fraught with challenges. Specifically, we foreground the pivotal role of local teachers in mitigating these challenges. Through a series of situated efforts, local teachers contextualize high-quality lectures to the local classroom environment, ensuring the expected educational outcomes. Based on our findings, we argue that greater recognition and support for the situational practices of local teachers is essential for fostering a more equitable, sustainable, and scalable technology-driven educational model in disadvantaged areas. Yuling Sun, Jiaju Chen, Xiaomu Zhou, Xiaojuan Ma, Bingsheng Yao, Liang He 0001, Dakuo Wang |
CHI | 5 |
| 2025 | CardioAI: A Multimodal AI-based System to Support Symptom Monitoring and Risk Prediction of Cancer Treatment-Induced CardiotoxicityabstractDespite recent advances in cancer treatments that prolong patients' lives, treatment-induced cardiotoxicity (i.e., the various heart damages caused by cancer treatments) emerges as one major side effect. The clinical decision-making process of cardiotoxicity is challenging, as early symptoms may happen in non-clinical settings and are too subtle to be noticed until life-threatening events occur at a later stage; clinicians already have a high workload focusing on the cancer treatment, no additional effort to spare on the cardiotoxicity side effect. Our project starts with a participatory design study with 11 clinicians to understand their decision-making practices and their feedback on an initial design of an AI-based decision-support system. Based on their feedback, we then propose a multimodal AI system, CardioAI, that can integrate wearables data and voice assistant data to model a patient's cardiotoxicity risk to support clinicians' decision-making. We conclude our paper with a small-scale heuristic evaluation with four experts and the discussion of future design considerations. Weidan Cao, Shihan Fu, Bingsheng Yao, Changchang Yin, Varun Mishra 0001, Daniel Addison, Ping Zhang 0016, Dakuo Wang |
CHI | 4 |
| 2025 | SepsisCalc: Integrating Clinical Calculators into Early Sepsis Prediction via Dynamic Temporal Graph Constructionabstract., the six-organ dysfunction assessment of SOFA in Figure 1) play a vital role in sepsis identification within clinicians' workflow, providing evidence-based risk assessments essential for sepsis diagnosis. However, artificial intelligence (AI) sepsis prediction models typically generate a single sepsis risk score without incorporating clinical calculators for assessing organ dysfunctions, making the models less convincing and transparent to clinicians. To bridge the gap, we propose to mimic clinicians' workflow with a novel framework SepsisCalc to integrate clinical calculators into the predictive model, yielding a clinically transparent and precise model for utilization in clinical settings. Practically, clinical calculators usually combine information from multiple component variables in Electronic Health Records (EHR), and might not be applicable when the variables are (partially) missing. We mitigate this issue by representing EHRs as temporal graphs and integrating a learning module to dynamically add the accurately estimated calculator to the graphs. Experimental results on real-world datasets show that the proposed model outperforms state-of-the-art methods on sepsis prediction tasks. Moreover, we developed a system to identify organ dysfunctions and potential sepsis risks, providing a human-AI interaction tool for deployment, which can help clinicians understand the prediction outputs and prepare timely interventions for the corresponding dysfunctions, paving the way for actionable clinical decision-making support for early intervention. Changchang Yin, Shihan Fu, Bingsheng Yao, Thai-Hoang Pham, Weidan Cao, Dakuo Wang, Jeffrey M. Caterino, Ping Zhang 0016 |
KDD (1) | 3 |
| 2025 | User Interaction Patterns and Breakdowns in Conversing with LLM-Powered Voice Assistants
Amama Mahmood, Bingsheng Yao, Dakuo Wang, Chien-Ming Huang 0001 |
Int. J. Hum. Comput. Stud. | 3 |
| 2025 | "Mango Mango, How to Let The Lettuce Dry Without A Spinner?": Exploring User Perceptions of Using An LLM-Based Conversational Assistant Toward Cooking PartnerabstractThe rapid advancement of Large Language Models (LLMs) has created numerous potentials for integration with conversational assistants (CAs) assisting people in their daily tasks, particularly due to their extensive flexibility. However, users' real-world experiences interacting with these assistants remain unexplored. In this research, we chose cooking, a complex daily task, as a scenario to explore people's successful and unsatisfactory experiences while receiving assistance from an LLM-based CA, Mango Mango . We discovered that participants value the system's ability to offer customized instructions based on context, provide extensive information beyond the recipe, and assist them in dynamic task planning. However, users expect the system to be more adaptive to oral conversation and provide more suggestive responses to keep them actively involved. Recognizing that users began treating our LLM-CA as a personal assistant or even a partner rather than just a recipe-reading tool, we propose five design considerations for future development. Szeyi Chan, Bingsheng Yao, Amama Mahmood, Chien-Ming Huang 0001, Holly Jimison, Elizabeth D. Mynatt, Dakuo Wang |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2025 | Secret Use of Large Language Model (LLM)abstractThe advancements of Large Language Models (LLMs) have decentralized the responsibility for the transparency of AI usage. Specifically, LLM users are now encouraged or required to disclose the use of LLM-generated content for varied types of real-world tasks. However, an emerging phenomenon, users' secret use of LLM , raises challenges in ensuring end users adhere to the transparency requirement. Our study used mixed-methods with an exploratory survey (125 real-world secret use cases reported) and a controlled experiment among 300 users to investigate the contexts and causes behind the secret use of LLMs. We found that such secretive behavior is often triggered by certain tasks, transcending demographic and personality differences among users. Task types were found to affect users' intentions to use secretive behavior, primarily through influencing perceived external judgment regarding LLM usage. Our results yield important insights for future work on designing interventions to encourage more transparent disclosure of the use of LLMs or other AI technologies. Chenxinran Shen, Bingsheng Yao, Dakuo Wang, Tianshi Li 0001 |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2024 | "It's a Fair Game", or Is It? Examining How Users Navigate Disclosure Risks and Benefits When Using LLM-Based Conversational AgentsabstractThe widespread use of Large Language Model (LLM)-based conversational agents (CAs), especially in high-stakes domains, raises many privacy concerns. Building ethical LLM-based CAs that respect user privacy requires an in-depth understanding of the privacy risks that concern users the most. However, existing research, primarily model-centered, does not provide insight into users’ perspectives. To bridge this gap, we analyzed sensitive disclosures in real-world ChatGPT conversations and conducted semi-structured interviews with 19 LLM-based CA users. We found that users are constantly faced with trade-offs between privacy, utility, and convenience when using LLM-based CAs. However, users’ erroneous mental models and the dark patterns in system design limited their awareness and comprehension of the privacy risks. Additionally, the human-like interactions encouraged more sensitive disclosures, which complicated users’ ability to navigate the trade-offs. We discuss practical design guidelines and the needs for paradigm shifts to protect the privacy of LLM-based CA users. Michelle Jia, Hao-Ping Lee, Bingsheng Yao, Sauvik Das, Ada Lerner, Dakuo Wang, Tianshi Li 0001 |
CHI | 4 |
| 2024 | Rethinking Human-AI Collaboration in Complex Medical Decision Making: A Case Study in Sepsis DiagnosisabstractToday's AI systems for medical decision support often succeed on benchmark datasets in research papers but fail in real-world deployment. This work focuses on the decision making of sepsis, an acute life-threatening systematic infection that requires an early diagnosis with high uncertainty from the clinician. Our aim is to explore the design requirements for AI systems that can support clinical experts in making better decisions for the early diagnosis of sepsis. The study begins with a formative study investigating why clinical experts abandon an existing AI-powered Sepsis predictive module in their electrical health record (EHR) system. We argue that a human-centered AI system needs to support human experts in the intermediate stages of a medical decision-making process (e.g., generating hypotheses or gathering data), instead of focusing only on the final decision. Therefore, we build SepsisLab based on a state-of-the-art AI algorithm and extend it to predict the future projection of sepsis development, visualize the prediction uncertainty, and propose actionable suggestions (i.e., which additional laboratory tests can be collected) to reduce such uncertainty. Through heuristic evaluation with six clinicians using our prototype system, we demonstrate that SepsisLab enables a promising human-AI collaboration paradigm for the future of AI-assisted sepsis diagnosis and other high-stakes medical decision making. Shao Zhang, Xuhai Xu, Changchang Yin, Yuxuan Lu 0003, Bingsheng Yao, Melanie Tory, Lace M. K. Padilla, Jeffrey M. Caterino, Ping Zhang 0016, Dakuo Wang |
CHI | 6 |
| 2024 | StorySparkQA: Expert-Annotated QA Pairs with Real-World Knowledge for Children's Story-Based LearningabstractJiaju Chen, Yuxuan Lu, Shao Zhang, Bingsheng Yao, Yuanzhe Dong, Ying Xu, Yunyao Li, Qianwen Wang, Dakuo Wang, Yuling Sun. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Jiaju Chen, Yuxuan Lu 0003, Shao Zhang, Bingsheng Yao, Yuanzhe Dong, Yunyao Li 0001, Dakuo Wang, Yuling Sun |
EMNLP | 4 |
| 2024 | SepsisLab: Early Sepsis Prediction with Uncertainty Quantification and Active SensingabstractSepsis is the leading cause of in-hospital mortality in the USA. Early sepsis onset prediction and diagnosis could significantly improve the survival of sepsis patients. Existing predictive models are usually trained on high-quality data with few missing information, while missing values widely exist in real-world clinical scenarios (especially in the first hours of admissions to the hospital), which causes a significant decrease in accuracy and an increase in uncertainty for the predictive models. The common method to handle missing values is imputation, which replaces the unavailable variables with estimates from the observed data. The uncertainty of imputation results can be propagated to the sepsis prediction outputs, which have not been studied in existing works on either sepsis prediction or uncertainty quantification. In this study, we first define such propagated uncertainty as the variance of prediction output and then introduce uncertainty propagation methods to quantify the propagated uncertainty. Moreover, for the potential high-risk patients with low confidence due to limited observations, we propose a robust active sensing algorithm to increase confidence by actively recommending clinicians to observe the most informative variables. We validate the proposed models in both publicly available data (i.e., MIMIC-III and AmsterdamUMCdb) and proprietary data in The Ohio State University Wexner Medical Center (OSUWMC). The experimental results show that the propagated uncertainty is dominant at the beginning of admissions to hospitals and the proposed algorithm outperforms state-of-the-art active sensing methods. Finally, we implement a SepsisLab system for early sepsis prediction and active sensing based on our pre-trained models. Clinicians and potential sepsis patients can benefit from the system in early prediction and diagnosis of sepsis. Changchang Yin, Bingsheng Yao, Dakuo Wang, Jeffrey M. Caterino, Ping Zhang 0016 |
KDD | 3 |
| 2024 | Exploring Parent's Needs for Children-Centered AI to Support Preschoolers' Interactive Storytelling and Reading ActivitiesabstractInteractive storytelling is vital for preschooler development. While children's interactive partners have traditionally been their parents and teachers, recent advances in artificial intelligence (AI) have sparked a surge of AI-based storytelling and reading technologies. As these technologies become increasingly ubiquitous in preschoolers' lives, questions arise regarding how they function in practical storytelling and reading scenarios and, how parents, the most critical stakeholders, experience and perceive these technologies. This paper investigates these questions through a qualitative study with 17 parents of children aged 3-6. Our findings suggest that even though AI-based storytelling and reading technologies provide more immersive and engaging interaction, they still cannot meet parents' expectations due to a series of interactive and algorithmic challenges. We elaborate on these challenges and discuss the possible implications of future AI-based interactive storytelling technologies for preschoolers. Yuling Sun, Jiaju Chen, Bingsheng Yao, Dakuo Wang, Xiaojuan Ma, Yuxuan Lu 0003, Liang He 0001 |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2023 | Are Human Explanations Always Helpful? Towards Objective Evaluation of Human Natural Language ExplanationsabstractHuman-annotated labels and explanations are critical for training explainable NLP models.However, unlike human-annotated labels whose quality is easier to calibrate (e.g., with a majority vote), human-crafted free-form explanations can be quite subjective.Before blindly using them as ground truth to train ML models, a vital question needs to be asked: How do we evaluate a human-annotated explanation's quality?In this paper, we build on the view that the quality of a human-annotated explanation can be measured based on its helpfulness (or impairment) to the ML models' performance for the desired NLP tasks for which the annotations were collected.In comparison to the commonly used Simulatability score, we define a new metric that can take into consideration of the helpfulness of an explanation for model performance at both fine-tuning and inference.With the help of a unified dataset format, we evaluated the proposed metric on five datasets (e.g., e-SNLI) against two model architectures (T5 and BART), and the results show that our proposed metric can objectively evaluate the quality of human-annotated explanations, while Simulatability falls short. Bingsheng Yao, Prithviraj Sen, Lucian Popa 0001, James A. Hendler, Dakuo Wang |
ACL (1) | 1 |
| 2023 | 'Don't Get Too Technical with Me': A Discourse Structure-Based Framework for Automatic Science JournalismabstractScience journalism refers to the task of reporting technical findings of a scientific paper as a less technical news article to the general public audience.We aim to design an automated system to support this real-world task (i.e., automatic science journalism) by 1) introducing a newly-constructed and real-world dataset (SCITECHNEWS), with tuples of a publiclyavailable scientific paper, its corresponding news article, and an expert-written short summary snippet; 2) proposing a novel technical framework that integrates a paper's discourse structure with its metadata to guide generation; and, 3) demonstrating with extensive automatic and human experiments that our framework outperforms other baseline methods (e.g.Alpaca and ChatGPT) in elaborating a content plan meaningful for the target audience, simplifying the information selected, and producing a coherent final report in a layman's style. Ronald Cardenas, Bingsheng Yao, Dakuo Wang, Yufang Hou 0001 |
EMNLP | 2 |
| 2022 | Fantastic Questions and Where to Find Them: FairytaleQA - An Authentic Dataset for Narrative ComprehensionabstractYing Xu, Dakuo Wang, Mo Yu, Daniel Ritchie, Bingsheng Yao, Tongshuang Wu, Zheng Zhang, Toby Li, Nora Bradford, Branda Sun, Tran Hoang, Yisi Sang, Yufang Hou, Xiaojuan Ma, Diyi Yang, Nanyun Peng, Zhou Yu, Mark Warschauer. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Dakuo Wang, Mo Yu, Daniel Ritchie 0002, Bingsheng Yao, Sherry Tongshuang Wu, Zheng Zhang 0043, Toby Jia-Jun Li, Nora Bradford, Branda Sun, Tran Bao Hoang, Yisi Sang, Yufang Hou 0001, Xiaojuan Ma, Diyi Yang, Nanyun Peng 0001, Zhou Yu 0005, Mark Warschauer |
ACL (1) | 5 |
| 2022 | It is AI's Turn to Ask Humans a Question: Question-Answer Pair Generation for Children's Story BooksabstractExisting question answering (QA) techniques are created mainly to answer questions asked by humans.But in educational applications, teachers often need to decide what questions they should ask, in order to help students to improve their narrative understanding capabilities.We design an automated question-answer generation (QAG) system for this education scenario: given a story book at the kindergarten to eighth-grade level as input, our system can automatically generate QA pairs that are capable of testing a variety of dimensions of a student's comprehension skills.Our proposed QAG model architecture is demonstrated using a new expert-annotated FairytaleQA dataset, which has 278 child-friendly storybooks with 10,580 QA pairs.Automatic and human evaluations show that our model outperforms stateof-the-art QAG baseline systems.On top of our QAG system, we also start to build an interactive story-telling application for the future real-world deployment in this educational scenario. Bingsheng Yao, Dakuo Wang, Sherry Tongshuang Wu, Zheng Zhang 0043, Toby Jia-Jun Li, Mo Yu |
ACL (1) | 1 |
| 2022 | StoryBuddy: A Human-AI Collaborative Chatbot for Parent-Child Interactive Storytelling with Flexible Parental InvolvementabstractDespite its benefits for children’s skill development and parent-child bonding, many parents do not often engage in interactive storytelling by having story-related dialogues with their child due to limited availability or challenges in coming up with appropriate questions. While recent advances made AI generation of questions from stories possible, the fully-automated approach excludes parent involvement, disregards educational goals, and underoptimizes for child engagement. Informed by need-finding interviews and participatory design (PD) results, we developed StoryBuddy, an AI-enabled system for parents to create interactive storytelling experiences. StoryBuddy’s design highlighted the need for accommodating dynamic user needs between the desire for parent involvement and parent-child bonding and the goal of minimizing parent intervention when busy. The PD revealed varied assessment and educational goals of parents, which StoryBuddy addressed by supporting configuring question types and tracking child progress. A user study validated StoryBuddy’s usability and suggested design insights for future parent-AI collaboration systems. Zheng Zhang 0043, Bingsheng Yao, Daniel Ritchie 0002, Sherry Tongshuang Wu, Mo Yu, Dakuo Wang, Toby Jia-Jun Li |
CHI | 4 |
| 2022 | A Corpus for Commonsense Inference in Story Cloze TestabstractThe Story Cloze Test (SCT) is designed for training and evaluating machine learning algorithms for narrative understanding and inferences. The SOTA models can achieve over 90% accuracy on predicting the last sentence. However, it has been shown that high accuracy can be achieved by merely using surface-level features. We suspect these models may not truly understand the story. Based on the SCT dataset, we constructed a human-labeled and human-verified commonsense knowledge inference dataset. Given the first four sentences of a story, we asked crowd-source workers to choose from four types of narrative inference for deciding the ending sentence and which sentence contributes most to the inference. We accumulated data on 1871 stories, and three human workers labeled each story. Analysis of the intra-category and inter-category agreements show a high level of consensus. We present two new tasks for predicting the narrative inference categories and contributing sentences. Our results show that transformer-based models can reach SOTA performance on the original SCT task using transfer learning but don’t perform well on these new and more challenging tasks. Bingsheng Yao, Ethan Joseph, Julian Lioanag |
LREC | 1 |
| 2021 | Narrative Question Answering with Cutting-Edge Open-Domain QA Techniques: A Comprehensive StudyabstractAbstract Recent advancements in open-domain question answering (ODQA), that is, finding answers from large open-domain corpus like Wikipedia, have led to human-level performance on many datasets. However, progress in QA over book stories (Book QA) lags despite its similar task formulation to ODQA. This work provides a comprehensive and quantitative analysis about the difficulty of Book QA: (1) We benchmark the research on the NarrativeQA dataset with extensive experiments with cutting-edge ODQA techniques. This quantifies the challenges Book QA poses, as well as advances the published state-of-the-art with a ∼7% absolute improvement on ROUGE-L. (2) We further analyze the detailed challenges in Book QA through human studies.1 Our findings indicate that the event-centric questions dominate this task, which exemplifies the inability of existing QA models to handle event-oriented scenarios. Xiangyang Mou, Chenghao Yang 0001, Mo Yu, Bingsheng Yao, Saloni Potdar, Hui Su |
Trans. Assoc. Comput. Linguistics | 4 |
| 2020 | Trust in AutoML: exploring information needs for establishing trust in automated machine learning systemsabstractWe explore trust in a relatively new area of data science: Automated Machine Learning (AutoML). In AutoML, AI methods are used to generate and optimize machine learning models by automatically engineering features, selecting models, and optimizing hyperparameters. In this paper, we seek to understand what kinds of information influence data scientists' trust in the models produced by AutoML? We operationalize trust as a willingness to deploy a model produced using automated methods. We report results from three studies - qualitative interviews, a controlled experiment, and a card-sorting task - to understand the information needs of data scientists for establishing trust in AutoML systems. We find that including transparency features in an AutoML tool increased user trust and understandability in the tool; and out of all proposed features, model performance metrics and visualizations are the most important information to data scientists when establishing their trust with an AutoML tool. Jaimie Drozdal, Justin D. Weisz, Dakuo Wang, Gaurav Dass, Bingsheng Yao, Changruo Zhao, Michael J. Muller, Lin Ju, Hui Su |
IUI | 5 |