VLDB 2026 Research / reviewers in the wild / expert
Lydia B. Chilton
dblp:73/7444
· DBLP profile ↗
55ranked-venue papers
5as first author
39since 2021 · last 2026
0000-0002-1737-1276ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 41 · 4 first-author · 28 since 2021Artificial intelligence and machine learning · 9 · 7 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior SimulationabstractZiyi Wang, Yuxuan Lu, Wenbo Li, Amirali Amini, Bo Sun, Yakov Bart, Weimin Lyu, Jiri Gesi, Tian Wang, Jing Huang, Yu Su, Upol Ehsan, Malihe Alikhani, Toby Jia-Jun Li, Lydia Chilton, Dakuo Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuxuan Lu 0003, Amirali Amini, Yakov Bart, Weimin Lyu, Jiri Gesi, Upol Ehsan, Malihe Alikhani, Toby Jia-Jun Li, Lydia B. Chilton, Dakuo Wang |
ACL (1) | 15 |
| 2026 | Identifying, Explaining, and Correcting Ableist Language with AIabstractAbleist language perpetuates harmful stereotypes and exclusion, yet its nuanced nature makes it difficult to recognize and address. Artificial intelligence could serve as a powerful ally in the fight against ableist language, offering tools that detect and suggest alternatives to biased terms. This two-part study investigates the potential of large language models (LLMs), specifically ChatGPT, to rectify ableist language and educate users about inclusive communication. We compared GPT-4o generations with crowdsourced annotations from trained disability community members, then invited disabled participants to evaluate both. Participants reported equal agreement with human and AI annotations but significantly preferred the AI, citing its narrative consistency and accessible style. At the same time, they valued the emotional depth and cultural grounding of human annotations. These findings highlight the promise and limits of LLMs in handling culturally sensitive content. Our contributions include a dataset of nuanced ableism annotations and design considerations for inclusive writing tools. Kynnedy Simone Smith, Lydia B. Chilton, Danielle Bragg |
CHI | 2 |
| 2026 | Rewriting Video: Text-Driven Reauthoring of Video FootageabstractVideo is a powerful medium for communication and storytelling, yet reauthoring existing footage remains challenging. Even simple edits often demand expertise, time, and careful planning, constraining how creators envision and shape their narratives. Recent advances in generative AI suggest a new paradigm: what if editing a video were as straightforward as rewriting text? To investigate this, we present a tech probe and a study on text-driven video reauthoring. Our approach involves two technical contributions: (1) a generative reconstruction algorithm that reverse-engineers video into an editable text prompt, and (2) an interactive probe, Rewrite Kit, that allows creators to manipulate these prompts. A technical evaluation of the algorithm reveals a critical human-AI perceptual gap. A probe study with 12 creators surfaced novel use cases such as virtual reshooting, synthetic continuity, and aesthetic restyling. It also highlighted key tensions around coherence, control, and creative alignment in this new paradigm. Our work contributes empirical insights into the opportunities and challenges of text-driven video reauthoring, offering design implications for future co-creative video tools. Sitong Wang 0001, Anh Truong, Lydia B. Chilton, Dingzeyu Li |
IUI | 3 |
| 2026 | PERSONAJUDGE: Simulating Individual Human Preference Judgments with Evaluator-Specific Demonstration DataabstractLarge language models increasingly serve as judges in AI evaluation, but current approaches rely on consensus preferences that ignore individual evaluator variation. We propose a novel simulation approach that combines categorical judgments with evaluator-specific auxiliary data—retrospective reasoning traces and interface telemetry—to enable LLM-based simulation of individual evaluators via in-context learning. We conduct a systematic empirical study of this approach using multi-facet data from 32 trained annotators across 4,200 preference judgments in a 4 × 4 × 4 factorial design. Our key findings: (1) The simulation approach achieves up to 9.9 percentage point improvements over the Base Judge; (2) Reasoning traces provide the largest gains with higher collection efforts, while interface telemetry often hurts rather than helps performance despite being cheaper to collect. (3) Simulation difficulty is systematic, predicted by an evaluator’s neutral usage (most clearly on Helpfulness) and divergence from consensus; the neutral-usage tendency—rather than simulatability itself—is the cross-task-stable property (r = 0.728). These results establish both the potential and limits of evaluator-specific auxiliary data for personalized evaluation, offering methodological insights for scaling individual aware AI assessment. Xuan Qi, Subramanian Chidambaram, Zhichao Xu 0001, Vinayak Arannil, Lydia B. Chilton, Alex C. Williams |
SIGDIAL | 6 |
| 2025 | LogoMotion: Visually-Grounded Code Synthesis for Creating and Editing Animation
Vivian Liu, Rubaiat Habib Kazi, Li-Yi Wei, Matthew Fisher, Timothy R. Langlois, Seth Walker, Lydia B. Chilton |
CHI | 7 |
| 2025 | DynEx: Dynamic Code Synthesis with Structured Design Exploration for Accelerated Exploratory Programming
Jenny Ma, Karthik Sreedhar, Vivian Liu, Pedro Alejandro Perez, Sitong Wang 0001, Riya Sahni, Lydia B. Chilton |
CHI | 7 |
| 2025 | Copying style, Extracting value: Illustrators' Perception of AI Style Transfer and its Impact on Creative Labor
Julien Porquet, Sitong Wang 0001, Lydia B. Chilton |
CHI | 3 |
| 2025 | Simulating Cooperative Prosocial Behavior with Multi-Agent LLMs: Evidence and Mechanisms for AI Agents to Inform Policy Decisions
Karthik Sreedhar, Alice Cai, Jenny Ma, Jeffrey V. Nickerson, Lydia B. Chilton |
IUI | 5 |
| 2025 | Answering Developer Questions with Annotated Agent-Discovered Program Traces
Litao Yan, Jeffrey Tao, Lydia B. Chilton, Andrew Head |
UIST | 3 |
| 2024 | Not Just Novelty: A Longitudinal Study on Utility and Customization of an AI WorkflowabstractGenerative AI brings novel and impressive abilities to help people in everyday tasks. There are many AI workflows that solve real and complex problems by chaining AI outputs together with human interaction. Although there is an undeniable lure of AI, it is uncertain how useful generative AI workflows are after the novelty wears off. Additionally, workflows built with generative AI have the potential to be easily customized to fit users’ individual needs, but do users take advantage of this? We conducted a three-week longitudinal study with 12 users to understand the familiarization and customization of generative AI tools for science communication. Our study revealed that there exists a familiarization phase, during which users were exploring the novel capabilities of the workflow and discovering which aspects they found useful. After this phase, users understood the workflow and were able to anticipate the outputs. Surprisingly, after familiarization the perceived utility of the system was rated higher than before, indicating that the perceived utility of AI is not just a novelty effect. The increase in benefits mainly comes from end-users’ ability to customize prompts, and thus potentially appropriate the system to their own needs. This points to a future where generative AI systems can allow us to design for appropriation. Tao Long 0003, Katy Ilonka Gero, Lydia B. Chilton |
Conference on Designing Interactive Systems | 3 |
| 2024 | PodReels: Human-AI Co-Creation of Video Podcast TeasersabstractVideo podcast teasers are short videos that can be shared on social media platforms to capture interest in full episodes of a video podcast. These teasers enable long-form podcasters to reach new audiences and gain more followers. However, creating a compelling teaser from an hour-long episode can be challenging. Selecting interesting clips requires significant mental effort; editing the chosen clips into a cohesive, well-produced teaser is time-consuming. To support the creation of video podcast teasers, we first investigated what makes a good teaser. We combined insights from audience comments and creator interviews to identify key ingredients. We also identified a common workflow used by creators during this process. Based on these findings, we developed a human-AI co-creative tool called PodReels to assist video podcasters in crafting teasers. Our user study demonstrated that PodReels significantly reduces creators’ mental demand and improves their efficiency in producing video podcast teasers. Sitong Wang 0001, Zheng Ning, Anh Truong, Mira Dontcheva, Dingzeyu Li, Lydia B. Chilton |
Conference on Designing Interactive Systems | 6 |
| 2024 | Seeing Science: Inquiry-Based Learning at Home Through Mobile Messaging SystemabstractThis work-in-progress proposes an approach that uses a low-cost, smartphone-based system for at-home, inquiry-driven science learning called STEM-Messaging System (SMS). SMS supports real-time, interactive, message-based science activities and is part of a broader project aimed at integrating science into children's daily lives by uncovering the science behind everyday objects via computer vision overlays. We discuss how three pedagogical principles–inquiry-based learning, culturally relevant pedagogies, and modeling-based learning–inform key design features of the system and its curricular activities. We identify tensions that surfaced from pilot studies involving students, parents, and teachers, providing examples of how pedagogical principles and practical applications influence design decisions. Tamar Fuhrmann, Marina A. Lemee, Jonathan Pang, Je Seung You, Lydia B. Chilton, Carl Vondrick, Paulo Blikstein |
IDC | 5 |
| 2024 | Writing out the Storm: Designing and Evaluating Tools for Weather Risk MessagingabstractCommunicating risk to the public in the lead-up to and during severe weather events has the potential to reduce the impacts of these events on lives and property. Globally, these events are anticipated to increase due to climate change, rendering effective risk communication an integral component of climate adaptation policies. Research in risk communications literature has developed substantial knowledge and best practices for the design of risk messaging. This study considers the potential for quantifying the compliance of severe weather risk messages with these best practices, individually and at scale, and developing tools to improve risk communication messaging. The current work makes two contributions. First, we develop a string-matching approach to evaluate whether messaging complies with best practices and suggest areas for improvement. Second, we conduct an interview study with risk communication professionals to inform the design space of authoring tools and other technologies to support severe weather risk communicators. Sophia Jit, Jennifer Spinney, Priyank Chandra, Lydia B. Chilton, Robert Soden |
CHI | 4 |
| 2024 | ReelFramer: Human-AI Co-Creation for News-to-Video TranslationabstractShort videos on social media are the dominant way young people consume content. News outlets aim to reach audiences through news reels—short videos conveying news—but struggle to translate traditional journalistic formats into short, entertaining videos. To translate news into social media reels, we support journalists in reframing the narrative. In literature, narrative framing is a high-level structure that shapes the overall presentation of a story. We identified three narrative framings for reels that adapt social media norms but preserve news value, each with a different balance of information and entertainment. We introduce ReelFramer, a human-AI co-creative system that helps journalists translate print articles into scripts and storyboards. ReelFramer supports exploring multiple narrative framings to find one appropriate to the story. AI suggests foundational narrative details, including characters, plot, setting, and key information. ReelFramer also supports visual framing; AI suggests character and visual detail designs before generating a full storyboard. Our studies show that narrative framing introduces the necessary diversity to translate various articles into reels, and establishing foundational details helps generate scripts that are more relevant and coherent. We also discuss the benefits of using narrative framing and foundational details in content retargeting. Sitong Wang 0001, Samia Menon, Tao Long 0003, Keren Henderson, Dingzeyu Li, Kevin Crowston, Mark Hansen, Jeffrey V. Nickerson, Lydia B. Chilton |
CHI | 9 |
| 2024 | STORYSUMM: Evaluating Faithfulness in Story SummarizationabstractHuman evaluation has been the gold standard for checking faithfulness in abstractive summarization.However, with a challenging source domain like narrative, multiple annotators can agree a summary is faithful, while missing details that are obvious errors only once pointed out.We therefore introduce a new dataset, STORYSUMM, comprising LLM summaries of short stories with localized faithfulness labels and error explanations.This benchmark is for evaluation methods, testing whether a given method can detect challenging inconsistencies.Using this dataset, we first show that any one human annotation protocol is likely to miss inconsistencies, and we advocate for pursuing a range of methods when establishing ground truth for a summarization dataset.We finally test recent automatic metrics and find that none of them achieve more than 70% balanced accuracy on this task, demonstrating that it is a challenging benchmark for future work in faithfulness evaluation. Melanie Subbiah, Faisal Ladhak, Akankshya Mishra, Griffin Adams, Lydia B. Chilton, Kathy McKeown |
EMNLP | 5 |
| 2024 | MoodSmith: Enabling Mood-Consistent Multimedia for AI-Generated Advocacy Campaigns
Samia Menon, Sitong Wang 0001, Lydia B. Chilton |
ICCC | 3 |
| 2024 | MIRACLE: An Online, Explainable Multimodal Interactive Concept Learning SystemabstractWe present MIRACLE, a system for online, interpretable visual concept and video action recognition. Through a chat interface, users query the recognition system with an uploaded image or video. For images, MIRACLE returns concept predictions from its structured knowledge base, justifying its predictions with heatmaps and natural language-based attribute detections. For videos, MIRACLE predicts an action and justifies its prediction with time varying entity-entity relations. With its ability to learn new concepts in an online, few-shot manner and its support of dynamic changes to its knowledge base, MIRACLE represents a step forward in interpretable multimodal learning systems. Ansel Blume, Khanh Duy Nguyen, Zhenhailong Wang, Yangyi Chen, Michal Shlapentokh-Rothman, Xiaomeng Jin, Zhen Zhu 0006, Jiateng Liu, Kuan-Hao Huang, Mankeerat Sidhu, Xuanming Zhang, Vivian Liu, Raunak Sinha, Te-Lin Wu, Abhaysinh Zala, Elias Stengel-Eskin, Da Yin, Utkarsh Mall, Zhou Yu 0005, Kai-Wei Chang 0001, Camille Cobb, Karrie Karahalios, Lydia B. Chilton, Mohit Bansal, Nanyun Peng 0001, Carl Vondrick, Derek Hoiem, Heng Ji 0001 |
ACM Multimedia | 25 |
| 2024 | Reading Subtext: Evaluating Large Language Models on Short Story Summarization with WritersabstractAbstract We evaluate recent Large Language Models (LLMs) on the challenging task of summarizing short stories, which can be lengthy, and include nuanced subtext or scrambled timelines. Importantly, we work directly with authors to ensure that the stories have not been shared online (and therefore are unseen by the models), and to obtain informed evaluations of summary quality using judgments from the authors themselves. Through quantitative and qualitative analysis grounded in narrative theory, we compare GPT-4, Claude-2.1, and LLama-2-70B. We find that all three models make faithfulness mistakes in over 50% of summaries and struggle with specificity and interpretation of difficult subtext. We additionally demonstrate that LLM ratings and other automatic metrics for summary quality do not correlate well with the quality ratings from the writers. Melanie Subbiah, Sean Zhang, Lydia B. Chilton, Kathy McKeown |
Trans. Assoc. Comput. Linguistics | 3 |
| 2023 | Metaphorian: Leveraging Large Language Models to Support Extended Metaphor Creation for Science WritingabstractScience writers commonly use extended metaphors to communicate unfamiliar concepts in a more accessible way to a wider audience. However, creating metaphors for science writing is challenging even for professional writers; according to our formative study (n=6), finding inspiration and extending metaphors with coherent structures were critical yet significantly challenging tasks for them. We contribute Metaphorian, a system that supports science writers with the creation of scientific metaphors by facilitating the search, extension, and iterative revision of metaphors. Metaphorian uses a large language model-based workflow inspired by the heuristic rules revealed from a study with six professional writers. A user study (n=16) revealed that Metaphorian significantly enhances satisfaction, confidence, and inspiration in metaphor writing without decreasing writers’ sense of agency. We discuss design implications for creativity support for figurative writing in science. Jeongyeon Kim, Sangho Suh, Lydia B. Chilton, Haijun Xia |
Conference on Designing Interactive Systems | 3 |
| 2023 | StoryWars: A Dataset and Instruction Tuning Baselines for Collaborative Story Understanding and GenerationabstractCollaborative stories, which are texts created through the collaborative efforts of multiple authors with different writing styles and intentions, pose unique challenges for NLP models.Understanding and generating such stories remains an underexplored area due to the lack of open-domain corpora.To address this, we introduce STORYWARS, a new dataset of over 40,000 collaborative stories written by 9,400 different authors from an online platform.We design 12 task types, comprising 7 understanding and 5 generation task types, on STORY-WARS, deriving 101 diverse story-related tasks in total as a multi-task benchmark covering all fully-supervised, few-shot, and zero-shot scenarios.Furthermore, we present our instructiontuned model, INSTRUCTSTORY, for the story tasks showing that instruction tuning, in addition to achieving superior results in zero-shot and few-shot scenarios, can also obtain the best performance on the fully-supervised tasks in STORYWARS, establishing strong multi-task benchmark performances on STORYWARS. 1 Yulun Du, Lydia B. Chilton |
ACL (1) | 2 |
| 2023 | How Much is Performance Worth to Users?abstractIn computer systems design, computer architects can evaluate features and techniques in terms of traditional design metrics like power, performance, and area but not in terms of net benefit or cost to the end user. For example, a security feature may come at a 10% cost to performance, but is this worth the tradeoff? The problem is that user-level features like security, privacy, and usability can be converted into a monetary amount but traditional architecture metrics (which lack the notion of user value) cannot. In this paper, we make the first known attempt to bridge this gap: We conduct two studies (one of which is incentive compatible) that elicit the value of performance in terms of US$ to end users. Thus in this work, we make the first known quantitative measurement of the tradeoff between performance and user value, providing architects with a novel design metric and filling a crucial gap in the end-to-end quantitative evaluation of systems. Adam Hastings, Lydia B. Chilton, Simha Sethumadhavan |
CF | 2 |
| 2023 | Social Dynamics of AI Support in Creative WritingabstractRecently, large language models have made huge advances in generating coherent, creative text. While much research focuses on how users can interact with language models, less work considers the social-technical gap that this technology poses. What are the social nuances that underlie receiving support from a generative AI? In this work we ask when and why a creative writer might turn to a computer versus a peer or mentor for support. We interview 20 creative writers about their writing practice and their attitudes towards both human and computer support. We discover three elements that govern a writer’s interaction with support actors: 1) what writers desire help with, 2) how writers perceive potential support actors, and 3) the values writers hold. We align our results with existing frameworks of writing cognition and creativity support, uncovering the social dynamics which modulate user responses to generative technologies. Katy Ilonka Gero, Tao Long 0003, Lydia B. Chilton |
CHI | 3 |
| 2023 | Improving Automatic Summarization for Browsing Longform Spoken DialogabstractLongform spoken dialog delivers rich streams of informative content through podcasts, interviews, debates, and meetings. While production of this medium has grown tremendously, spoken dialog remains challenging to consume as listening is slower than reading and difficult to skim or navigate relative to text. Recent systems leveraging automatic speech recognition (ASR) and automatic summarization allow users to better browse speech data and forage for information of interest. However, these systems intake disfluent speech which causes automatic summarization to yield readability, adequacy, and accuracy problems. To improve navigability and browsability of speech, we present three training agnostic post-processing techniques that address dialog concerns of readability, coherence, and adequacy. We integrate these improvements with user interfaces which communicate estimated summary metrics to aid user browsing heuristics. Quantitative evaluation metrics show a 19% improvement in summary quality. We discuss how summarization technologies can help people browse longform audio in trustworthy and readable ways. Daniel Li 0002, Alec Zadikian, Albert Tung, Lydia B. Chilton |
CHI | 5 |
| 2023 | AngleKindling: Supporting Journalistic Angle Ideation with Large Language ModelsabstractNews media often leverage documents to find ideas for stories, while being critical of the frames and narratives present. Developing angles from a document such as a press release is a cognitively taxing process, in which journalists critically examine the implicit meaning of its claims. Informed by interviews with journalists, we developed AngleKindling, an interactive tool which employs the common sense reasoning of large language models to help journalists explore angles for reporting on a press release. In a study with 12 professional journalists, we show that participants found AngleKindling significantly more helpful and less mentally demanding to use for brainstorming ideas, compared to a prior journalistic angle ideation tool. AngleKindling helped journalists deeply engage with the press release and recognize angles that were useful for multiple types of stories. From our findings, we discuss how to help journalists customize and identify promising angles, and extending AngleKindling to other knowledge-work domains. Savvas Petridis, Nicholas Diakopoulos, Kevin Crowston, Mark Hansen, Keren Henderson, Stan Jastrzebski, Jeffrey V. Nickerson, Lydia B. Chilton |
CHI | 8 |
| 2023 | PopBlends: Strategies for Conceptual Blending with Large Language ModelsabstractPop culture is an important aspect of communication. On social media people often post pop culture reference images that connect an event, product or other entity to a pop culture domain. Creating these images is a creative challenge that requires finding a conceptual connection between the users’ topic and a pop culture domain. In cognitive theory, this task is called conceptual blending. We present a system called PopBlends that automatically suggests conceptual blends. The system explores three approaches that involve both traditional knowledge extraction methods and large language models. Our annotation study shows that all three methods provide connections with similar accuracy, but with very different characteristics. Our user study shows that people found twice as many blend suggestions as they did without the system, and with half the mental demand. We discuss the advantages of combining large language models with knowledge bases for supporting divergent and convergent thinking. Sitong Wang 0001, Savvas Petridis, Taeahn Kwon, Xiaojuan Ma, Lydia B. Chilton |
CHI | 5 |
| 2023 | Tweetorial Hooks: Generative AI Tools to Motivate Science on Social Media
Tao Long 0003, Dorothy Zhang, Grace Li, Batool Taraif, Samia Menon, Kynnedy Simone Smith, Sitong Wang 0001, Katy Ilonka Gero, Lydia B. Chilton |
ICCC | 9 |
| 2022 | Eliciting Gestures for Novel Note-taking InteractionsabstractHandwriting recognition is improving in leaps and bounds, and this opens up new opportunities for stylus-based interactions. In particular, note-taking applications can become a more intelligent user interface, incorporating new features like autocomplete and integrated search. In this work we ran a gesture elicitation study, asking 21 participants to imagine how they would interact with an imaginary, intelligent note-taking application. Participants were prompted to produce gestures for common actions such as select and delete, as well as less common actions (for gesture interaction) such as autocomplete accept/reject, ‘hide’, and search. We report agreement on the elicited gestures, finding that while existing interactions are prevalent (like double taps and long presses) a number of more novel interactions (like dragging selected items to hotspots or using annotations) were also well-represented. We discuss the mental models participants drew on when explaining their gestures and what kind of feedback users might need to move to more stylus-centric interactions. Katy Ilonka Gero, Lydia B. Chilton, Chris Melancon, Mike Cleron |
Conference on Designing Interactive Systems | 2 |
| 2022 | Sparks: Inspiration for Science Writing using Language ModelsabstractLarge-scale language models are rapidly improving, performing well on a wide variety of tasks with little to no customization. In this work we investigate how language models can support science writing, a challenging writing task that is both open-ended and highly constrained. We present a system for generating “sparks”, sentences related to a scientific concept intended to inspire writers. We find that our sparks are more coherent and diverse than a competitive language model baseline, and approach a human-written gold standard. We run a user study with 13 STEM graduate students writing on topics of their own selection and find three main use cases of sparks—inspiration, translation, and perspective—each of which correlates with a unique interaction pattern. We also find that while participants were more likely to select higher quality sparks, the average quality of sparks seen by a given participant did not correlate with their satisfaction with the tool. We end with a discussion about what impacts human satisfaction with AI support tools, considering participant attitudes towards influence, their openness to technology, as well as issues of plagiarism, trustworthiness, and bias in AI. Katy Ilonka Gero, Vivian Liu, Lydia B. Chilton |
Conference on Designing Interactive Systems | 3 |
| 2022 | Initial Images: Using Image Prompts to Improve Subject Representation in Multimodal AI Generated ArtabstractAdvances in text-to-image generative models have made it easier for people to create art by just prompting models with text. However, creating through text leaves users with limited control over the final composition or the way the subject is represented. A potential solution is to use image prompts alongside text prompts to condition the model. To better understand how and when image prompts can improve subject representation in generations, we conduct an annotation experiment to quantify their effect on generations of abstract, concrete plural, and concrete singular subjects. We find that initial images improved subject representation across all subject types, with the most noticeable improvement in concrete singular subjects. In an analysis of different types of initial images, we find that icons and photos produced high quality generations of different aesthetics. We conclude with design guidelines for how initial images can improve subject representation in AI art. Han Qiao, Vivian Liu, Lydia B. Chilton |
Creativity & Cognition | 3 |
| 2022 | Design Guidelines for Prompt Engineering Text-to-Image Generative ModelsabstractText-to-image generative models are a new and powerful way to generate visual artwork. However, the open-ended nature of text as interaction is double-edged; while users can input anything and have access to an infinite range of generations, they also must engage in brute-force trial and error with the text prompt when the result quality is poor. We conduct a study exploring what prompt keywords and model hyperparameters can help produce coherent outputs. In particular, we study prompts structured to include subject and style keywords and investigate success and failure modes of these prompts. Our evaluation of 5493 generations over the course of five experiments spans 51 abstract and concrete subjects as well as 51 abstract and figurative styles. From this evaluation, we present design guidelines that can help people produce better outcomes from text-to-image generative models. Vivian Liu, Lydia B. Chilton |
CHI | 2 |
| 2022 | Insights and Opportunities for HCI Research into Hurricane Risk CommunicationabstractCommunicating risk to the public in the lead-up to tropical storms has the potential to significantly reduce the impacts on both livelihood and property. While significant research has been conducted in the storm risk community on how people receive, seek, and utilize risk information, given the importance of computing technologies and social media in these activities, human-centered design stands to make important contributions to this area. Drawing on an extensive literature review and 48 interviews with hurricane experts and members of the public, this paper makes three contributions. First, we provide a broad overview of hurricane risk communication. We then offer a set of guiding insights to inform HCI research work in this domain. Finally, we identify 6 opportunities that future human centered design work might pursue. In sum, this paper offers an invitation and a starting point for HCI to take up the problem of hurricane risk communication. Robert Soden, Lydia B. Chilton, Scott B. Miles, Rebecca Bicksler, Kaira Ray Villanueva, Melissa Bica |
CHI | 2 |
| 2022 | SafeText: A Benchmark for Exploring Physical Safety in Language ModelsabstractSharon Levy, Emily Allaway, Melanie Subbiah, Lydia Chilton, Desmond Patton, Kathleen McKeown, William Yang Wang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Sharon Levy, Emily Allaway, Melanie Subbiah, Lydia B. Chilton, Desmond Upton Patton, Kathy McKeown, William Yang Wang |
EMNLP | 4 |
| 2022 | First Workshop on Content Understanding and Generation for E-commerceabstractShopping experience on any e-commerce website is largely driven by the content customers interact with. The large volume of diverse content on e-commerce platforms, and the advances in machine learning, pose unique opportunities for gathering insights through content understanding and applying these insights to generate content better shopper experience. The purpose of the first edition of this workshop was to bring together researchers from industry and academia on questions surrounding e-commerce content understanding and generation. Sumit Negi, Manisha Verma, Rajdeep H. Banerjee, Pooja A, Lydia B. Chilton, Mithun Das Gupta, Vinay P. Namboodiri, Dinesh Garg |
KDD | 5 |
| 2022 | Opal: Multimodal Image Generation for News IllustrationabstractAdvances in multimodal AI have presented people with powerful ways to create images from text. Recent work has shown that text-to-image generations are able to represent a broad range of subjects and artistic styles. However, finding the right visual language for text prompts is difficult. In this paper, we address this challenge with Opal, a system that produces text-to-image generations for news illustration. Given an article, Opal guides users through a structured search for visual concepts and provides a pipeline allowing users to generate illustrations based on an article’s tone, keywords, and related artistic styles. Our evaluation shows that Opal efficiently generates diverse sets of news illustrations, visual assets, and concept ideas. Users with Opal generated two times more usable results than users without. We discuss how structured exploration can help users better understand the capabilities of human AI co-creative systems. Vivian Liu, Han Qiao, Lydia B. Chilton |
UIST | 3 |
| 2021 | VisiFit: Structuring Iterative Improvement for Novice DesignersabstractVisual blends are an advanced graphic design technique to seamlessly integrate two objects into one. Existing tools help novices create prototypes of blends, but it is unclear how they would improve them to be higher fidelity. To help novices, we aim to add structure to the iterative improvement process. We introduce a method for improving prototypes that uses secondary design dimensions to explore a structured design space. This method is grounded in the cognitive principles of human visual object recognition. We present VisiFit – a computational design system that uses this method to enable novice graphic designers to improve blends with computationally generated options they can select, adjust, and chain together. Our evaluation shows novices can substantially improve 76% of blends in under 4 minutes. We discuss how the method can be generalized to other blending problems, and how computational tools can support novices by enabling them to explore a structured design space quickly and efficiently. Lydia B. Chilton, Ecenaz Jen Ozmen, Sam H. Ross, Vivian Liu |
CHI | 1 |
| 2021 | AmbiTeam: Providing Team Awareness Through Ambient DisplaysabstractDue to the COVID-19 pandemic, research is increasingly conducted remotely without the benefit of informal interactions that help maintain awareness of each collaborator's work progress. We developed AmbiTeam, an ambient display that shows activity related to the files of a team project, to help collaborations preserve a sense of the team's involvement while working remotely. We found that using AmbiTeam did have a quantifiable effect on researchers' perceptions of their collaborators' project prioritization. We also found that the use of the system motivated researchers to work on their collaborative projects. This effect is known as "the motivational presences of others," one of the key challenges that make distance work difficult. We discuss how ambient displays can support remote collaborative work by recreating the motivational presence of others. Sarah Morrison-Smith, Lydia B. Chilton, Jaime Ruiz 0002 |
Graphics Interface | 2 |
| 2021 | Hierarchical Summarization for Longform Spoken DialogabstractEvery day we are surrounded by spoken dialog. This medium delivers rich diverse streams of information auditorily; however, systematically understanding dialog can often be non-trivial. Despite the pervasiveness of spoken dialog, automated speech understanding and quality information extraction remains markedly poor, especially when compared to written prose. Furthermore, compared to understanding text, auditory communication poses many additional challenges such as speaker disfluencies, informal prose styles, and lack of structure. These concerns all demonstrate the need for a distinctly speech tailored interactive system to help users understand and navigate the spoken language domain. While individual automatic speech recognition (ASR) and text summarization methods already exist, they are imperfect technologies; neither consider user purpose and intent nor address spoken language induced complications. Consequently, we design a two stage ASR and text summarization pipeline and propose a set of semantic segmentation and merging algorithms to resolve these speech modeling challenges. Our system enables users to easily browse and navigate content as well as recover from errors in these underlying technologies. Finally, we present an evaluation of the system which highlights user preference for hierarchical summarization as a tool to quickly skim audio and identify content of interest to the user. Daniel Li 0002, Albert Tung, Lydia B. Chilton |
UIST | 4 |
| 2021 | SymbolFinder: Brainstorming Diverse Symbols Using Local Semantic NetworksabstractVisual symbols are the building blocks for visual communication. They convey abstract concepts like reform and participation quickly and effectively. When creating graphics with symbols, novice designers often struggle to brainstorm multiple, diverse symbols because they fixate on a few associations instead of broadly exploring different aspects of the concept. We present SymbolFinder, an interactive tool for finding visual symbols for abstract concepts. SymbolFinder molds symbol-finding into a recognition rather than recall task by introducing the user to diverse clusters of words associated with the concept. Users can dive into these clusters to find related, concrete objects that symbolize the concept. We evaluate SymbolFinder with two studies: a comparative user study, demonstrating that SymbolFinder helps novices find more unique symbols for abstract concepts with significantly less effort than a popular image database and a case study demonstrating how SymbolFinder helped design students create visual metaphors for three cover illustrations of news articles. Savvas Petridis, Hijung Shin, Lydia B. Chilton |
UIST | 3 |
| 2021 | What Makes Tweetorials Tick: How Experts Communicate Complex Topics on TwitterabstractPeople are increasingly getting information and news from social media. On Twitter we are seeing the emergence of "tweetorials" -- long, explanatory Twitter threads written by experts. In this work we study tweetorials as a form of science writing. While scientists have begun to champion the importance of Twitter as a science communication medium, few have studied how people are successfully using this medium to communicate complex and nuanced ideas. To understand how tweetorials work, we curated a collection of 46 clear and engaging tweetorials from multiple domains. We analyzed these tweetorials for the writing techniques that they employ, and found that while tweetorials use many traditional science writing techniques, they also use more subjective language, actively build credibility, and incorporate media in unique ways. In addition, we report on a workshop we ran to aid science PhD students in writing tweetorials, and find that while providing common tweetorial techniques improves their writing, the students still struggle to balance their scientific sensibilities with the informal tone associated with tweetorials. We discuss the implications of using informal and subjective language in science communication, as well as how technology can support scientists in writing tweetorials. Katy Ilonka Gero, Vivian Liu, Sarah Huang, Jennifer Lee, Lydia B. Chilton |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2020 | WordBlender: principles and tools for generating word blendsabstractCombining text and images is a powerful strategy in graphic design because images convey meaning faster, but text conveys more precise meaning. Word blends are a technique to combine both elements in a succinct yet expressive image. In a word blend, a letter is replaced by a symbol relevant to the message. This is difficult because the replacement must look blended enough to be readable, yet different enough to recognize the symbol. Currently, there are no known design principles to find aesthetically pleasing word blends. To establish these principles, we run two experiments and find that to be readable, the object should have a similar shape as the letter. However, to be aesthetically pleasing, the font should match some of the secondary features of the image: color, style, and thickness. We present WordBlender, an AI-powered design tool to quickly and easily create word blends based on these visual design principles. WordBlender automatically generates shape-based matches and allows users to explore combinations of color, style, and font that improve the design of blends. Sam H. Ross, Ecenaz Jen Ozmen, Maria V. Kogan, Lydia B. Chilton |
IUI | 4 |
| 2019 | Making Memes AccessibleabstractImages on social media platforms are inaccessible to people with vision impairments due to a lack of descriptions that can be read by screen readers. Providing accurate alternative text for all visual content on social media is not yet feasible, but certain subsets of images, such as internet memes, offer affordances for automatic or semi-automatic generation of alternative text. We present two methods for making memes accessible semi-automatically through (1) the generation of rich alternative text descriptions and (2) the creation of audio macro memes. Meme authors create alternative text templates or audio meme templates, and insert placeholders instead of the meme text. When a meme with the same image is encountered again, it is automatically recognized from a database of meme templates. Text is then extracted and either inserted into the alternative text template or rendered in the audio template using text-to-speech. In our evaluation of meme formats with 10 Twitter users with vision impairments, we found that most users preferred alternative text memes because the description of the visual content conveys the emotional tone of the character. As the preexisting templates can be automatically matched to memes using the same visual image, this combined approach can make a large subset of images on the web accessible, while preserving the emotion and tone inherent in the image memes. Cole Gleason, Amy Pavel, Xingyu Liu 0002, Patrick Carrington, Lydia B. Chilton, Jeffrey P. Bigham |
ASSETS | 5 |
| 2019 | How a Stylistic, Machine-Generated Thesaurus Impacts a Writer's ProcessabstractWriters regularly use a thesaurus to help them write well; the thesaurus is one of the few widespread writing support tools and many writers find it integral to their writing practice. A normal thesaurus is hand-crafted and structured around strict synonymy for a given word sense. However, writers rarely look for a perfectly synonymous word -- instead they have additional ideas or constraints, such as words that are less cliche, more specific, or less gendered. Poets describe their usage as searching for words that "hold more interesting connotations." We present a machine learning approach to thesaurus generation, using word embeddings, that leverages stylistically distinct corpora -- such as naturalist writing, novels by a particular author, or writing from a technical discipline. We show examples of how stylistic thesauruses differ from each other and from a regular thesaurus, as well as preliminary responses from two writers who are given multiple stylistic thesauruses. Writers describe these thesauruses as reflective of style, unique from each other, and more exploratory and associative than a regular thesaurus. They also describe an increased attention to connotation. We outline plans for quantitative evaluation of stylistic thesauruses, and user studies to understand their impact on specific tasks. Katy Ilonka Gero, Lydia B. Chilton |
Creativity & Cognition | 2 |
| 2019 | Human Errors in Interpreting Visual MetaphorabstractVisual metaphors are a creative technique used in print media to convey a message through images. This message is not said directly, but implied through symbols and how those symbols are juxtaposed in the image. The messages we see affect our thoughts and lives, and it is an open research challenge to get machines to automatically understand the implied messages in images. However, it is unclear how people process these images or to what degree they understand the meaning. We test several theories about how people interpret visual metaphors and find people can interpret the visual metaphor correctly without explanatory text with 41.3% accuracy. We provide evidence for four distinct types of errors people make in their interpretation, which speaks to the cognitive processes people use to infer the meaning. We also show that people's ability to interpret a visual message is not simply a function of image content but also of message familiarity. This implies that efforts to automatically understand visual images should take into account message familiarity. Savvas Petridis, Lydia B. Chilton |
Creativity & Cognition | 2 |
| 2019 | Cicero: Multi-Turn, Contextual Argumentation for Accurate CrowdsourcingabstractTraditional approaches for ensuring high quality crowdwork have failed to achieve high-accuracy on difficult problems. Aggregating redundant answers often fails on the hardest problems when the majority is confused. Argumentation has been shown to be effective in mitigating these drawbacks. However, existing argumentation systems only support limited interactions and show workers general justifications, not context-specific arguments targeted to their reasoning. This paper presents Cicero, a new workflow that improves crowd accuracy on difficult tasks by engaging workers in multi-turn, contextual discussions through real-time, synchronous argumentation. Our experiments show that compared to previous argumentation systems which only improve the average individual worker accuracy by 6.8 percentage points on the Relation Extraction domain, our workflow achieves 16.7 percentage point improvement. Furthermore, previous argumentation approaches don't apply to tasks with many possible answers; in contrast, Cicero works well in these cases, raising accuracy from 66.7% to 98.8% on the Codenames domain. Quanze Chen, Jonathan Bragg, Lydia B. Chilton, Daniel S. Weld |
CHI | 3 |
| 2019 | VisiBlends: A Flexible Workflow for Visual BlendsabstractVisual blends are an advanced graphic design technique to draw attention to a message. They combine two objects in a way that is novel and useful in conveying a message symbolically. This paper presents VisiBlends, a flexible workflow for creating visual blends that follows the iterative design process. We introduce a design pattern for blending symbols based on principles of human visual object recognition. Our workflow decomposes the process into both computational techniques and human microtasks. It allows users to collaboratively generate visual blends with steps involving brainstorming, synthesis, and iteration. An evaluation of the workflow shows that decentralized groups can generate blends in independent microtasks, co-located groups can collaboratively make visual blends for their own messages, and VisiBlends improves novices' ability to make visual blends. Lydia B. Chilton, Savvas Petridis, Maneesh Agrawala |
CHI | 1 |
| 2019 | Metaphoria: An Algorithmic Companion for Metaphor CreationabstractCreative writing, from poetry to journalism, is at the crux of human ingenuity and social interaction. Existing creative writing support tools produce entire passages or fully formed sentences, but these approaches fail to adapt to the writer's own ideas and intentions. Instead we posit to build tools that generate ideas coherent with the writer's context and encourage writers to produce divergent outcomes. To explore this, we focus on supporting metaphor creation. We present Metaphoria, an interactive system that generates metaphorical connections based on an input word from the writer. Our studies show that Metaphoria provides more coherent suggestions than existing systems, and supports the expression of writers' unique intentions. We discuss the complex issue of ownership in human-machine collaboration and how to build adaptive creativity support tools in other domains. Katy Ilonka Gero, Lydia B. Chilton |
CHI | 2 |
| 2019 | Low Level Linguistic Controls for Style Transfer and Content PreservationabstractDespite the success of style transfer in image processing, it has seen limited progress in natural language generation. Part of the problem is that content is not as easily decoupled from style in the text domain. Curiously, in the field of stylometry, content does not figure prominently in practical methods of discriminating stylistic elements, such as authorship and genre. Rather, syntax and function words are the most salient features. Drawing on this work, we model style as a suite of low-level linguistic controls, such as frequency of pronouns, prepositions, and subordinate clause constructions. We train a neural encoder-decoder model to reconstruct reference sentences given only content words and the setting of the controls. We perform style transfer by keeping the content words fixed while adjusting the controls to be indicative of another style. In experiments, we show that the model reliably responds to the linguistic controls and perform both automatic and manual evaluations on style transfer. We find we can fool a style classifier 84% of the time, and that our model produces highly diverse and stylistically distinctive outputs. This work introduces a formal, extendable model of style that can add control to any neural text generation system. Katy Ilonka Gero, Chris Kedzie, Jonathan Reeve, Lydia B. Chilton |
INLG | 4 |
| 2016 | MicroTalk: Using Argumentation to Improve Crowdsourcing AccuracyabstractCrowd workers are human and thus sometimes make mistakes. In order to ensure the highest quality output, requesters often issue redundant jobs with gold test questions and sophisticated aggregation mechanisms based on expectation maximization (EM). While these methods yield accurate results in many cases, they fail on extremely difficult problems with local minima, such as situations where the majority of workers get the answer wrong. Indeed, this has caused some researchers to conclude that on some tasks crowdsourcing can never achieve high accuracies, no matter how many workers are involved. This paper presents a new quality-control workflow, called MicroTalk, that requires some workers to Justify their reasoning and asks others to Reconsider their decisions after reading counter-arguments from workers with opposing views. Experiments on a challenging NLP annotation task with workers from Amazon Mechanical Turk show that (1) argumentation improves the accuracy of individual workers by 20%, (2) restricting consideration to workers with complex explanations improves accuracy even more, and (3) our complete MicroTalk aggregation workflow produces much higher accuracy than simpler voting approaches for a range of budgets. Ryan Drapeau, Lydia B. Chilton, Jonathan Bragg, Daniel S. Weld |
HCOMP | 2 |
| 2014 | Frenzy: collaborative data organization for creating conference sessionsabstractOrganizing conference sessions around themes improves the experience for attendees. However, the session creation process can be difficult and time-consuming due to the amount of expertise and effort required to consider alternative paper groupings. We present a collaborative web application called Frenzy to draw on the efforts and knowledge of an entire program committee. Frenzy comprises (a) interfaces to support large numbers of experts working collectively to create sessions, and (b) a two-stage process that decomposes the session-creation problem into meta-data elicitation and global constraint satisfaction. Meta-data elicitation involves a large group of experts working simultaneously, while global constraint satisfaction involves a smaller group that uses the meta-data to form sessions. Lydia B. Chilton, Juho Kim 0001, Paul André, Felicia Cordeiro, James A. Landay, Daniel S. Weld, Steven Dow, Rob Miller 0001 |
CHI | 1 |
| 2013 | Cascade: crowdsourcing taxonomy creationabstractTaxonomies are a useful and ubiquitous way of organizing information. However, creating organizational hierarchies is difficult because the process requires a global understanding of the objects to be categorized. Usually one is created by an individual or a small group of people working together for hours or even days. Unfortunately, this centralized approach does not work well for the large, quickly changing datasets found on the web. Cascade is an automated workflow that allows crowd workers to spend as little at 20 seconds each while collectively making a taxonomy. We evaluate Cascade and show that on three datasets its quality is 80-90% of that of experts. Cascade has a competitive cost to expert information architects, despite taking six times more human labor. Fortunately, this labor can be parallelized such that Cascade will run in as fast as four minutes instead of hours or days. Lydia B. Chilton, Greg Little, Darren Edge, Daniel S. Weld, James A. Landay |
CHI | 1 |
| 2013 | Community Clustering: Leveraging an Academic Crowd to Form Coherent Conference SessionsabstractCreating sessions of related papers for a large conference is a complex and time-consuming task. Traditionally, a few conference organizers group papers into sessions manually. Organizers often fail to capture the affinities between papers beyond created sessions, making incoherent sessions difficult to fix and alternative groupings hard to discover. This paper proposes committeesourcing and authorsourcing approaches to session creation (a specific instance of clustering and constraint satisfaction) that tap into the expertise and interest of committee members and authors for identifying paper affinities. During the planning of ACM CHI'13, a large conference on human-computer interaction, we recruited committee members to group papers using two online distributed clustering methods. To refine these paper affinities — and to evaluate the committeesourcing methods against existing manual and automated approaches — we recruited authors to identify papers that fit well in a session with their own. Results show that authors found papers grouped by the distributed clustering methods to be as relevant as, or more relevant than, papers suggested through the existing in-person meeting. Results also demonstrate that communitysourced results capture affinities beyond sessions and provide flexibility during scheduling. Paul André, Juho Kim 0001, Lydia B. Chilton, Steven Dow, Rob Miller 0001 |
HCOMP | 4 |
| 2013 | Cobi: a community-informed conference scheduling toolabstractEffectively planning a large multi-track conference requires an understanding of the preferences and constraints of organizers, authors, and attendees. Traditionally, the onus of scheduling the program falls on a few dedicated organizers. Resolving conflicts becomes difficult due to the size and complexity of the schedule and the lack of insight into community members' needs and desires. Cobi presents an alternative approach to conference scheduling that engages the entire community in the planning process. Cobi comprises (a) communitysourcing applications that collect preferences, constraints, and affinity data from community members, and (b) a visual scheduling interface that combines communitysourced data and constraint-solving to enable organizers to make informed improvements to the schedule. This paper describes Cobi's scheduling tool and reports on a live deployment for planning CHI 2013, where organizers considered input from 645 authors and resolved 168 scheduling conflicts. Results show the value of integrating community input with an intelligent user interface to solve complex planning tasks. Juho Kim 0001, Paul André, Lydia B. Chilton, Wendy E. Mackay, Michel Beaudouin-Lafon, Rob Miller 0001, Steven Dow |
UIST | 4 |
| 2011 | Addressing people's information needs directly in a web search result pageabstractWeb search engines have historically focused on connecting people with information resources. For example, if a person wanted to know when their flight to Hyderabad was leaving, a search engine might connect them with the airline where they could find flight status information. However, search engines have recently begun to try to meet people's search needs directly, providing, for example, flight status information in response to queries that include an airline and a flight number. In this paper, we use large scale query log analysis to explore the challenges a search engine faces when trying to meet an information need directly in the search result page. We look at how people's interaction behavior changes when inline content is returned, finding that such content can cannibalize clicks from the algorithmic results. We see that in the absence of interaction behavior, an individual's repeat search behavior can be useful in understanding the content's value. We also discuss some of the ways user behavior can be used to provide insight into when inline answers might better trigger and what types of additional information might be included in the results. Lydia B. Chilton, Jaime Teevan |
WWW | 1 |
| 2010 | The labor economics of paid crowdsourcingabstractWe present a model of workers supplying labor to paid crowdsourcing projects. We also introduce a novel method for estimating a worker's reservation wage - the key parameter in our labor supply model. We tested our model by presenting experimental subjects with real-effort work scenarios that varied in the offered payment and difficulty. As predicted, subjects worked less when the pay was lower. However, they did not work less when the task was more time-consuming. Interestingly, at least some subjects appear to be "target earners," contrary to the assumptions of the rational model. The strongest evidence for target earning is an observed preference for earning total amounts evenly divisible by 5, presumably because these amounts make good targets. Despite its predictive failures, we calibrate our model with data pooled from both experiments. We find that the reservation wages of our sample are approximately log normally distributed, with a median wage of $1.38/hour. We discuss how to use our calibrated model in applications. John Joseph Horton, Lydia B. Chilton |
EC | 2 |
| 2010 | TurKit: human computation algorithms on mechanical turkabstractMechanical Turk (MTurk) provides an on-demand source of human computation. This provides a tremendous opportunity to explore algorithms which incorporate human computation as a function call. However, various systems challenges make this difficult in practice, and most uses of MTurk post large numbers of independent tasks. TurKit is a toolkit for prototyping and exploring algorithmic human computation, while maintaining a straight-forward imperative programming style. We present the crash-and-rerun programming model that makes TurKit possible, along with a variety of applications for human computation algorithms. We also present case studies of TurKit used for real experiments across different fields. Greg Little, Lydia B. Chilton, Max Goldman, Rob Miller 0001 |
UIST | 2 |