Jiaju Ma

dblp:186/6610 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0003-2880-8506ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 10 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Artistic Practice Opportunities in CST Evaluations: A Longitudinal Group Deployment of ArtKrit
abstract
Creativity support tools (CSTs) aim to elevate the quality of artists’ creative processes and artifacts. Yet most current CST evaluations overlook temporal and social aspects of tool use. To address this gap, we present a longitudinal, group-based CST evaluation through a three-week deployment of ArtKrit, a computational drawing tool that supports disciplined drawing. Nine digital artists, organized into three communities of practice, completed weekly “master studies” alongside a researcher-artist. Our results show users’ evolving relationships with ArtKrit over time—from early experimentation to selective incorporation or misuse—alongside changes in their ways of artistic seeing. These changes unfolded within artist support networks that fostered confidence and creative safety, and validated individual expression. Overall, our findings suggest that CST evaluations can—and should—be designed as opportunities for meaningful artistic engagement rather than purely extractive measurement exercises. We contribute this longitudinal, group-based approach as one CST evaluation method.
Catherine Liu, Tao Long 0003, Asya Lyubavina, Chau Vu, Jiaju Ma
DIS5
2026 Collaposer: Transforming Photo Collections into Visual Assets for Storytelling with Collages
abstract
Digital collage is an artistic practice that combines image cutouts to tell stories. However, preparing cutouts from a set of photos remains a tedious and time-consuming task. A formative study identified three main challenges: 1) inefficient search for relevant photos, 2) manual image cutout, and 3) difficulty in organizing large sets of cutouts. To meet these challenges and facilitate asset preparation for collage, we propose Collaposer, a tool that transforms a collection of photos into organized, ready-to-use visual cutouts based on user-provided story descriptions. Collaposer tags, detects, and segments photos, and then uses an LLM to select central and related labels based on the user-provided story description. Collaposer presents the resulting visuals in varying sizes, clustered according to semantic hierarchy. Our evaluation shows that Collaposer effectively automates the preparation process to produce diverse sets of visual cutouts adhering to the storyline, allowing users to focus on collaging these assets for storytelling.
Liwenhan Xie, Jiaju Ma, Zheng Wei 0003, Huamin Qu, Anyi Rao
CHI3
2026 Script2Screen: Supporting Dialogue-Centric Scriptwriting with Interactive Audiovisual Generation
abstract
Scriptwriting has traditionally been text-centric, a modality that only partially conveys the produced audiovisual experience. A formative study with professional writers informed us that connecting textual and audiovisual modalities can aid ideation and iteration, especially for writing dialogues. In this work, we present Script2Screen, an AI-assisted tool that integrates scriptwriting with audiovisual scene creation in a unified, synchronized workflow. Focusing on dialogues in scripts, Script2Screen generates expressive scenes with emotional speeches and animated characters through a novel text-to-audiovisual-scene pipeline. The user interface provides fine-grained controls, allowing writers to fine-tune audiovisual elements such as character gestures, speech emotions, and camera angles. A user study with both novice and professional writers from various domains demonstrated that Script2Screen’s interactive audiovisual generation enhances the scriptwriting process, facilitating iterative refinement while complementing - rather than replacing - their creative efforts.
Zhecheng Wang 0001, Jiaju Ma, Eitan Grinspun, Tovi Grossman, Bryan Wang
IUI2
2025 OmniQuery: Contextually Augmenting Captured Multimodal Memories to Enable Personal Question Answering
Jiahao Nick Li, Zhuohao (Jerry) Zhang, Jiaju Ma
CHI3
2025 AmbigChat: Interactive Hierarchical Clarification for Ambiguous Open-Domain Question Answering
Jiaju Ma, Kenneth Aleksander Robertsen, Peggy Chi
UIST1
2025 Computational Scaffolding of Composition, Value, and Color for Disciplined Drawing
abstract
One way illustrators engage in disciplined drawing - the process of drawing to improve technical skills - is through studying and replicating reference images. However, for many novice and intermediate digital artists, knowing how to approach studying a reference image can be challenging. It can also be difficult to receive immediate feedback on their works-in-progress. To help these users develop their professional vision, we propose ArtKrit, a tool that scaffolds the process of replicating a reference image into three main steps: composition, value, and color. At each step, our tool offers computational guidance, such as adaptive composition line generation, and automatic feedback, such as value and color accuracy. Evaluating this tool with intermediate digital artists revealed that ArtKrit could flexibly accommodate their unique workflows. Our code and supplemental materials are available at https://majiaju.io/artkrit .
Jiaju Ma, Chau Vu, Asya Lyubavina, Catherine Liu
UIST1
2025 MoVer: Motion Verification for Motion Graphics Animations
abstract
While large vision-language models can generate motion graphics animations from text prompts, they regularly fail to include all spatio-temporal properties described in the prompt. We introduce MoVer, a motion verification DSL based on first-order logic that can check spatio-temporal properties of a motion graphics animation. We identify a general set of such properties that people commonly use to describe animations (e.g., the direction and timing of motions, the relative positioning of objects, etc.). We implement these properties as predicates in MoVer and provide an execution engine that can apply a MoVer program to any input SVG-based motion graphics animation. We then demonstrate how MoVer can be used in an LLM-based synthesis and verification pipeline for iteratively refining motion graphics animations. Given a text prompt, our pipeline synthesizes a motion graphics animation and a corresponding MoVer program. Executing the verification program on the animation yields a report of the predicates that failed and the report can be automatically fed back to LLM to iteratively correct the animation. To evaluate our pipeline, we build a synthetic dataset of 5600 text prompts paired with ground truth MoVer verification programs. We find that while our LLM-based pipeline is able to automatically generate a correct motion graphics animation for 58.8% of the test prompts without any iteration, this number raises to 93.6% with up to 50 correction iterations. Our code and dataset are at https://mover-dsl.github.io.
Jiaju Ma, Maneesh Agrawala
ACM Trans. Graph.1
2023 Automated Conversion of Music Videos into Lyric Videos
abstract
Musicians and fans often produce lyric videos, a form of music videos that showcase the song’s lyrics, for their favorite songs. However, making such videos can be challenging and time-consuming as the lyrics need to be added in synchrony and visual harmony with the video. Informed by prior work and close examination of existing lyric videos, we propose a set of design guidelines to help creators make such videos. Our guidelines ensure the readability of the lyric text while maintaining a unified focus of attention. We instantiate these guidelines in a fully automated pipeline that converts an input music video into a lyric video. We demonstrate the robustness of our pipeline by generating lyric videos from a diverse range of input sources. A user study shows that lyric videos generated by our pipeline are effective in maintaining text readability and unifying the focus of attention.
Jiaju Ma, Anyi Rao, Li-Yi Wei, Rubaiat Habib Kazi, Hijung Shin, Maneesh Agrawala
UIST1
2023 Editing Motion Graphics Video via Motion Vectorization and Transformation
abstract
Motion graphics videos are widely used in Web design, digital advertising, animated logos and film title sequences, to capture a viewer's attention. But editing such video is challenging because the video provides a low-level sequence of pixels and frames rather than higher-level structure such as the objects in the video with their corresponding motions and occlusions. We present a motion vectorization pipeline for converting motion graphics video into an SVG motion program that provides such structure. The resulting SVG program can be rendered using any SVG renderer (e.g. most Web browsers) and edited using any SVG editor. We also introduce a program transformation API that facilitates editing of a SVG motion program to create variations that adjust the timing, motions and/or appearances of objects. We show how the API can be used to create a variety of effects including retiming object motion to match a music beat, adding motion textures to objects, and collision preserving appearance changes.
Sharon Zhang, Jiaju Ma, Jiajun Wu 0001, Daniel Ritchie 0001, Maneesh Agrawala
ACM Trans. Graph.2
2022 A Layered Authoring Tool for Stylized 3D animations
abstract
Guided by the 12 principles of animation, stylization is a core 2D animation feature but has been utilized mainly by experienced animators. Although there are tools for stylizing 2D animations, creating stylized 3D animations remains a challenging problem due to the additional spatial dimension and the need for responsive actions like contact and collision. We propose a system that helps users create stylized casual 3D animations. A layered authoring interface is employed to balance between ease of use and expressiveness. Our surface level UI is a timeline sequencer that lets users add preset stylization effects such as squash and stretch and follow through to plain motions. Users can adjust spatial and temporal parameters to fine-tune these stylizations. These edits are propagated to our node-graph-based second level UI, in which the users can create custom stylizations after they are comfortable with the surface level UI. Our system also enables the stylization of interactions among multiple objects like force, energy, and collision. A pilot user study has shown that our fluid layered UI design allows for both ease of use and expressiveness better than existing tools.
Jiaju Ma, Li-Yi Wei, Rubaiat Habib Kazi
CHI1
2021 Portalware: Exploring Free-Hand AR Drawing with a Dual-Display Smartphone-Wearable Paradigm
abstract
Free-hand interaction enables users to directly create artistic augmented reality content using a smartphone, but lacks natural spatial depth information due to the small 2D display’s limited visual feedback. Through an autobiographical design process, three authors explored free-hand drawing over a total of 14 weeks. During this process, they expanded the design space from a single-display smartphone format to a dual-display smartphone-wearable format (Portalware). This new configuration extends the virtual content from a smartphone to a wearable display and enables multi-display free-hand interactions. The authors documented experiences where 1) the display extends the smartphone’s canvas perceptually, allowing the authors to work beyond the smartphone screen view; 2) the additional perspective mitigates the difficulties of depth perception and improves the usability of direct free-hand manipulation; 3) the wearable use cases depend on the nature of the drawing, such as: replicating physical objects, “in-situ” mixed reality pieces, and multi-planar drawings.
Tongyu Zhou, Meredith Young-Ng, Jiaju Ma, Angel Cheung, Ian Gonsher, Jeff Huang 0002
Conference on Designing Interactive Systems4
2019 Portal-ble: Intuitive Free-hand Manipulation in Unbounded Smartphone-based Augmented Reality
abstract
Smartphone augmented reality (AR) lets users interact with physical and virtual spaces simultaneously. With 3D hand tracking, smartphones become apparatus to grab and move virtual objects directly. Based on design considerations for interaction, mobility, and object appearance and physics, we implemented a prototype for portable 3D hand tracking using a smartphone, a Leap Motion controller, and a computation unit. Following an experience prototyping procedure, 12 researchers used the prototype to help explore usability issues and define the design space. We identified issues in perception (moving to the object, reaching for the object), manipulation (successfully grabbing and orienting the object), and behavioral understanding (knowing how to use the smartphone as a viewport). To overcome these issues, we designed object-based feedback and accommodation mechanisms and studied their perceptual and behavioral effects via two tasks: picking up distant objects, and assembling a virtual house from blocks. Our mechanisms enabled significantly faster and more successful user interaction than the initial prototype in picking up and manipulating stationary and moving objects, with a lower cognitive load and greater user preference. The resulting system---Portal-ble---improves user intuition and aids free-hand interactions in mobile situations.
Jiaju Ma, Benjamin Attal, Haoming Lai, James Tompkin 0001, John F. Hughes, Jeff Huang 0002
UIST2