Vivian Liu

dblp:264/7245 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0001-5328-0120ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 10 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 LogoMotion: Visually-Grounded Code Synthesis for Creating and Editing Animation
Vivian Liu, Rubaiat Habib Kazi, Li-Yi Wei, Matthew Fisher, Timothy R. Langlois, Seth Walker, Lydia B. Chilton
CHI1
2025 DynEx: Dynamic Code Synthesis with Structured Design Exploration for Accelerated Exploratory Programming
Jenny Ma, Karthik Sreedhar, Vivian Liu, Pedro Alejandro Perez, Sitong Wang 0001, Riya Sahni, Lydia B. Chilton
CHI3
2024 MIRACLE: An Online, Explainable Multimodal Interactive Concept Learning System
abstract
We present MIRACLE, a system for online, interpretable visual concept and video action recognition. Through a chat interface, users query the recognition system with an uploaded image or video. For images, MIRACLE returns concept predictions from its structured knowledge base, justifying its predictions with heatmaps and natural language-based attribute detections. For videos, MIRACLE predicts an action and justifies its prediction with time varying entity-entity relations. With its ability to learn new concepts in an online, few-shot manner and its support of dynamic changes to its knowledge base, MIRACLE represents a step forward in interpretable multimodal learning systems.
Ansel Blume, Khanh Duy Nguyen, Zhenhailong Wang, Yangyi Chen, Michal Shlapentokh-Rothman, Xiaomeng Jin, Zhen Zhu 0006, Jiateng Liu, Kuan-Hao Huang, Mankeerat Sidhu, Xuanming Zhang, Vivian Liu, Raunak Sinha, Te-Lin Wu, Abhaysinh Zala, Elias Stengel-Eskin, Da Yin, Utkarsh Mall, Zhou Yu 0005, Kai-Wei Chang 0001, Camille Cobb, Karrie Karahalios, Lydia B. Chilton, Mohit Bansal, Nanyun Peng 0001, Carl Vondrick, Derek Hoiem, Heng Ji 0001
ACM Multimedia13
2023 3DALL-E: Integrating Text-to-Image AI in 3D Design Workflows
abstract
Text-to-image AI are capable of generating novel images for inspiration, but their applications for 3D design workflows and how designers can build 3D models using AI-provided inspiration have not yet been explored. To investigate this, we integrated DALL-E, GPT-3, and CLIP within a CAD software in 3DALL-E, a plugin that generates 2D image inspiration for 3D design. 3DALL-E allows users to construct text and image prompts based on what they are modeling. In a study with 13 designers, we found that designers saw great potential in 3DALL-E within their workflows and could use text-to-image AI to produce reference images, prevent design fixation, and inspire design considerations. We elaborate on prompting patterns observed across 3D modeling tasks and provide measures of prompt complexity observed across participants. From our findings, we discuss how 3DALL-E can merge with existing generative design workflows and propose prompt bibliographies as a form of human-AI design history.
Vivian Liu, Jo Vermeulen, George W. Fitzmaurice, Justin Matejka
Conference on Designing Interactive Systems1
2023 CLIP-Sculptor: Zero-Shot Generation of High-Fidelity and Diverse Shapes from Natural Language
abstract
Recent works have demonstrated that natural language can be used to generate and edit 3D shapes. However, these methods generate shapes with limited fidelity and diversity. We introduce CLIP-Sculptor, a method to address these constraints by producing high-fidelity and diverse 3D shapes without the need for (text, shape) pairs during training. CLIP-Sculptor achieves this in a multi-resolution approach that first generates in a low-dimensional latent space and then upscales to a higher resolution for improved shape fidelity. For improved shape diversity, we use a discrete latent space which is modeled using a transformer conditioned on CLIP's image-text embedding space. We also present a novel variant of classifier-free guidance, which improves the accuracy-diversity trade-off. Finally, we perform extensive experiments demonstrating that CLIP-Sculptor outperforms state-of-the-art baselines.
Aditya Sanghi, Rao Fu 0003, Vivian Liu, Karl D. D. Willis, Hooman Shayani, Amir Khasahmadi, Srinath Sridhar 0002, Daniel Ritchie 0001
CVPR3
2022 Sparks: Inspiration for Science Writing using Language Models
abstract
Large-scale language models are rapidly improving, performing well on a wide variety of tasks with little to no customization. In this work we investigate how language models can support science writing, a challenging writing task that is both open-ended and highly constrained. We present a system for generating “sparks”, sentences related to a scientific concept intended to inspire writers. We find that our sparks are more coherent and diverse than a competitive language model baseline, and approach a human-written gold standard. We run a user study with 13 STEM graduate students writing on topics of their own selection and find three main use cases of sparks—inspiration, translation, and perspective—each of which correlates with a unique interaction pattern. We also find that while participants were more likely to select higher quality sparks, the average quality of sparks seen by a given participant did not correlate with their satisfaction with the tool. We end with a discussion about what impacts human satisfaction with AI support tools, considering participant attitudes towards influence, their openness to technology, as well as issues of plagiarism, trustworthiness, and bias in AI.
Katy Ilonka Gero, Vivian Liu, Lydia B. Chilton
Conference on Designing Interactive Systems2
2022 A Deep Convolutional Neural Network For Diagnosis of Diabetic Retinopathy
abstract
Diabetic Retinopathy (DR) diagnosis is a time consuming and complex task, in addition to requiring experienced doctors. In this study, an advanced Convolution Neural Network (CNN) and data augmentation were successfully used to predict the stages of diabetic retinopathy. Specifically, we modified and trained this network with the publicly available Kaggle dataset and obtained promising results. On the validation of 1452 images using this classifier, an accuracy of 82%, average sensitivity of 80% over the three categories and average specificity of 91% over the three categories could be achieved, indicating the feasibility of our CCN model for the identification of different stages of the diabetic retinopathy. It implies the great potential of this artificial intelligent method for the diagnosis of diabetic retinopathy, especially for the early diagnosis for patients in the future.
Vivian Liu, Mikhail Y. Shalaginov, Rory Liao, Tingying Helen Zeng
BIBM1
2022 Initial Images: Using Image Prompts to Improve Subject Representation in Multimodal AI Generated Art
abstract
Advances in text-to-image generative models have made it easier for people to create art by just prompting models with text. However, creating through text leaves users with limited control over the final composition or the way the subject is represented. A potential solution is to use image prompts alongside text prompts to condition the model. To better understand how and when image prompts can improve subject representation in generations, we conduct an annotation experiment to quantify their effect on generations of abstract, concrete plural, and concrete singular subjects. We find that initial images improved subject representation across all subject types, with the most noticeable improvement in concrete singular subjects. In an analysis of different types of initial images, we find that icons and photos produced high quality generations of different aesthetics. We conclude with design guidelines for how initial images can improve subject representation in AI art.
Han Qiao, Vivian Liu, Lydia B. Chilton
Creativity & Cognition2
2022 Design Guidelines for Prompt Engineering Text-to-Image Generative Models
abstract
Text-to-image generative models are a new and powerful way to generate visual artwork. However, the open-ended nature of text as interaction is double-edged; while users can input anything and have access to an infinite range of generations, they also must engage in brute-force trial and error with the text prompt when the result quality is poor. We conduct a study exploring what prompt keywords and model hyperparameters can help produce coherent outputs. In particular, we study prompts structured to include subject and style keywords and investigate success and failure modes of these prompts. Our evaluation of 5493 generations over the course of five experiments spans 51 abstract and concrete subjects as well as 51 abstract and figurative styles. From this evaluation, we present design guidelines that can help people produce better outcomes from text-to-image generative models.
Vivian Liu, Lydia B. Chilton
CHI1
2022 Opal: Multimodal Image Generation for News Illustration
abstract
Advances in multimodal AI have presented people with powerful ways to create images from text. Recent work has shown that text-to-image generations are able to represent a broad range of subjects and artistic styles. However, finding the right visual language for text prompts is difficult. In this paper, we address this challenge with Opal, a system that produces text-to-image generations for news illustration. Given an article, Opal guides users through a structured search for visual concepts and provides a pipeline allowing users to generate illustrations based on an article’s tone, keywords, and related artistic styles. Our evaluation shows that Opal efficiently generates diverse sets of news illustrations, visual assets, and concept ideas. Users with Opal generated two times more usable results than users without. We discuss how structured exploration can help users better understand the capabilities of human AI co-creative systems.
Vivian Liu, Han Qiao, Lydia B. Chilton
UIST1
2021 VisiFit: Structuring Iterative Improvement for Novice Designers
abstract
Visual blends are an advanced graphic design technique to seamlessly integrate two objects into one. Existing tools help novices create prototypes of blends, but it is unclear how they would improve them to be higher fidelity. To help novices, we aim to add structure to the iterative improvement process. We introduce a method for improving prototypes that uses secondary design dimensions to explore a structured design space. This method is grounded in the cognitive principles of human visual object recognition. We present VisiFit – a computational design system that uses this method to enable novice graphic designers to improve blends with computationally generated options they can select, adjust, and chain together. Our evaluation shows novices can substantially improve 76% of blends in under 4 minutes. We discuss how the method can be generalized to other blending problems, and how computational tools can support novices by enabling them to explore a structured design space quickly and efficiently.
Lydia B. Chilton, Ecenaz Jen Ozmen, Sam H. Ross, Vivian Liu
CHI4
2021 What Makes Tweetorials Tick: How Experts Communicate Complex Topics on Twitter
abstract
People are increasingly getting information and news from social media. On Twitter we are seeing the emergence of "tweetorials" -- long, explanatory Twitter threads written by experts. In this work we study tweetorials as a form of science writing. While scientists have begun to champion the importance of Twitter as a science communication medium, few have studied how people are successfully using this medium to communicate complex and nuanced ideas. To understand how tweetorials work, we curated a collection of 46 clear and engaging tweetorials from multiple domains. We analyzed these tweetorials for the writing techniques that they employ, and found that while tweetorials use many traditional science writing techniques, they also use more subjective language, actively build credibility, and incorporate media in unique ways. In addition, we report on a workshop we ran to aid science PhD students in writing tweetorials, and find that while providing common tweetorial techniques improves their writing, the students still struggle to balance their scientific sensibilities with the informal tone associated with tweetorials. We discuss the implications of using informal and subjective language in science communication, as well as how technology can support scientists in writing tweetorials.
Katy Ilonka Gero, Vivian Liu, Sarah Huang, Jennifer Lee, Lydia B. Chilton
Proc. ACM Hum. Comput. Interact.2
2020 Interacting with Literary Style through Computational Tools
abstract
Style is an important aspect of writing, shaping how audiences interpret and engage with literary works. However, for most people style is difficult to articulate precisely. While users frequently interact with computational word processing tools with well-defined metrics, such as spelling and grammar checkers, style is a significantly more nuanced concept. In this paper, we present a computational technique to help surface style in written text. We collect a dataset of crowdsourced human judgments of style, derive a model of style by training a neural net on this data, and present novel applications for visualizing and browsing style across broad bodies of literature, as well as an interactive text editor with real-time style feedback. We study these interactive style applications with users and discuss implications for enabling this novel approach to style.
Sarah Sterman, Evey Huang, Vivian Liu, Eric Paulos
CHI3